Skip to content

implicit: add riscv64 wheel build for 0.7.3 - #2410

Merged
luhenry merged 2 commits into
mainfrom
implicit
Sep 28, 2026
Merged

luhenry merged 2 commits into
mainfrom
implicit

Conversation

@luhenry

@luhenry luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Cython/C++ collaborative-filtering models (ALS, BPR, Logistic MF, item-item kNN) built via scikit-build-core + CMake, with OpenMP-parallel CPU kernels and an optional CUDA extension. Upstream publishes no riscv64 wheel; PyPI records no project URL for it either, so this PR also fills that in.

Mirrors upstream's build.yml.

Differs from upstream

  • Re-enables CIBW_REPAIR_WHEEL_COMMAND (auditwheel repair) - upstream's pyproject.toml sets repair-wheel-command = "" globally, only to keep auditwheel from vendoring CUDA libraries on manylinux_x86_64; riscv64 has no CUDA toolkit to vendor, so the normal repair runs and correctly bundles libgomp (CMake's find_package(OpenMP) links it automatically), with a gpl_sources job collecting the image's gcc sources for it.
  • CUDA install (ci/install_cuda_13.sh) is upstream's own manylinux_x86_64-only override; riscv64 needs no equivalent since find_package(CUDAToolkit) in implicit/gpu/CMakeLists.txt simply finds nothing and skips the GPU extension - verified locally that the CPU-only build succeeds cleanly with no CUDA toolkit present and implicit.gpu.HAS_CUDA is False.

Matrix: cp312/cp313/cp314, no cp314t - upstream's own cibuildwheel skip list excludes it ("*t-*") and its build.yml matrix never builds it.

Testing

  • same as upstream's test-wheels job (pytest tests/), minus installing annoy/nmslib for the optional approximate-nearest-neighbour backends - neither has a riscv64 wheel anywhere, and their tests self-skip via a plain try/import when absent (no pytest.skip needed).
  • Locally validated end-to-end: CPU-only build succeeds with no CUDA toolkit present, implicit.cpu.* extensions link libgomp, and the full suite passes against the built wheel (155 passed, 55 skipped, all GPU/ANN-backend tests).

License: OK - MIT, nothing third-party vendored beyond the toolchain's own libgomp (covered by gpl_sources).

Patches

  • 0001-use-schedule-static-for-the-dynamic-and-guided-OpenM.patch - Inappropriate [works around a libgomp defect on the riscv64 runners, libgomp dynamic and guided OpenMP schedules segfault on the riscv64 runners #617]. cp314 segfaulted inside libgomp's dynamic/guided iterator during AlternatingLeastSquares.fit; implicit's ALS, all-pairs KNN and top-k Cython cores move to schedule='static' (its BPR/LMF solvers already use it). Reproduces on riscv64 only.

Upstream (benfred/implicit) builds via scikit-build-core + CMake + Cython,
with an optional CUDA extension gated behind find_package(CUDAToolkit) -
verified locally that the CPU-only build succeeds cleanly with no CUDA
toolkit present, and the full test suite (155 passed, 55 skipped) runs
against the built wheel. CMake links OpenMP automatically when found, so a
gpl_sources job collects the image's gcc sources for the vendored libgomp.

CUDA install is x86_64-only in upstream's own CI, but upstream also disables
auditwheel repair globally purely to keep CUDA libraries out of the wheel;
since riscv64 has no CUDA toolkit to vendor, this build re-enables the
normal repair-wheel-command so libgomp is properly vendored instead.
luhenry added a commit that referenced this pull request Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-28 13:07 UTC

cp314-manylinux_riscv64 segfaulted (exit 139) in
tests/recommender_base_test.py::test_fit_non_csr_matrix, inside
implicit/cpu/als.py:164 (AlternatingLeastSquares.fit), with the fault
frame inside implicit.libs/libgomp-*.so.1.0.0 called from
implicit/cpu/_als.so.

This is gotcha 166: libgomp's dynamic and guided work-share schedules
fault on this fleet's riscv64 runners, static is unaffected (the
lightgbm case, #617). implicit's own ALS,
all-pairs KNN and top-k Cython cores ask for schedule='dynamic'/
'guided'; its BPR and LMF solvers already use schedule='static' and are
unaffected, matching the established pattern.

Patch the 6 affected prange() loops to schedule='static', mirroring
patches/lightgbm/4.7.0/0002-*, and apply it in the workflow the same
way.
@luhenry
luhenry marked this pull request as ready for review September 28, 2026 12:08
@luhenry
luhenry merged commit 4b85158 into main Sep 28, 2026
15 checks passed
@luhenry
luhenry deleted the implicit branch September 28, 2026 12:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant