Conversation
Upstream (benfred/implicit) builds via scikit-build-core + CMake + Cython, with an optional CUDA extension gated behind find_package(CUDAToolkit) - verified locally that the CPU-only build succeeds cleanly with no CUDA toolkit present, and the full test suite (155 passed, 55 skipped) runs against the built wheel. CMake links OpenMP automatically when found, so a gpl_sources job collects the image's gcc sources for the vendored libgomp. CUDA install is x86_64-only in upstream's own CI, but upstream also disables auditwheel repair globally purely to keep CUDA libraries out of the wheel; since riscv64 has no CUDA toolkit to vendor, this build re-enables the normal repair-wheel-command so libgomp is properly vendored instead.
luhenry
added a commit
that referenced
this pull request
Sep 28, 2026
Contributor
|
cp314-manylinux_riscv64 segfaulted (exit 139) in tests/recommender_base_test.py::test_fit_non_csr_matrix, inside implicit/cpu/als.py:164 (AlternatingLeastSquares.fit), with the fault frame inside implicit.libs/libgomp-*.so.1.0.0 called from implicit/cpu/_als.so. This is gotcha 166: libgomp's dynamic and guided work-share schedules fault on this fleet's riscv64 runners, static is unaffected (the lightgbm case, #617). implicit's own ALS, all-pairs KNN and top-k Cython cores ask for schedule='dynamic'/ 'guided'; its BPR and LMF solvers already use schedule='static' and are unaffected, matching the established pattern. Patch the 6 affected prange() loops to schedule='static', mirroring patches/lightgbm/4.7.0/0002-*, and apply it in the workflow the same way.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
implicit0.7.3Cython/C++ collaborative-filtering models (ALS, BPR, Logistic MF, item-item kNN) built via scikit-build-core + CMake, with OpenMP-parallel CPU kernels and an optional CUDA extension. Upstream publishes no riscv64 wheel; PyPI records no project URL for it either, so this PR also fills that in.
Mirrors upstream's
build.yml.Differs from upstream
CIBW_REPAIR_WHEEL_COMMAND(auditwheel repair) - upstream'spyproject.tomlsetsrepair-wheel-command = ""globally, only to keep auditwheel from vendoring CUDA libraries onmanylinux_x86_64; riscv64 has no CUDA toolkit to vendor, so the normal repair runs and correctly bundleslibgomp(CMake'sfind_package(OpenMP)links it automatically), with agpl_sourcesjob collecting the image's gcc sources for it.ci/install_cuda_13.sh) is upstream's ownmanylinux_x86_64-only override; riscv64 needs no equivalent sincefind_package(CUDAToolkit)inimplicit/gpu/CMakeLists.txtsimply finds nothing and skips the GPU extension - verified locally that the CPU-only build succeeds cleanly with no CUDA toolkit present andimplicit.gpu.HAS_CUDAisFalse.Matrix: cp312/cp313/cp314, no cp314t - upstream's own cibuildwheel
skiplist excludes it ("*t-*") and itsbuild.ymlmatrix never builds it.Testing
test-wheelsjob (pytest tests/), minus installingannoy/nmslibfor the optional approximate-nearest-neighbour backends - neither has a riscv64 wheel anywhere, and their tests self-skip via a plaintry/importwhen absent (nopytest.skipneeded).implicit.cpu.*extensions linklibgomp, and the full suite passes against the built wheel (155 passed, 55 skipped, all GPU/ANN-backend tests).License: OK - MIT, nothing third-party vendored beyond the toolchain's own
libgomp(covered bygpl_sources).Patches
0001-use-schedule-static-for-the-dynamic-and-guided-OpenM.patch- Inappropriate [works around a libgomp defect on the riscv64 runners, libgomp dynamic and guided OpenMP schedules segfault on the riscv64 runners #617]. cp314 segfaulted inside libgomp's dynamic/guided iterator duringAlternatingLeastSquares.fit; implicit's ALS, all-pairs KNN and top-k Cython cores move toschedule='static'(its BPR/LMF solvers already use it). Reproduces on riscv64 only.