Conversation
Mirrors upstream's own default pybind wheel build (setup.py + CMake --preset pybind): XNNPACK CPU delegate, portable/optimized/quantized kernels, CoreML's portable util+inmemoryfs+pybind pieces (no Apple delegate on Linux), OpenVINO backend (dlopen-based, no build-time SDK). QNN, CUDA and Vulkan stay off, same as upstream's own CI on riscv64/ARM hosts (QNN auto-skips under GITHUB_ACTIONS; CUDA/Vulkan are opt-in and neither is requested). torch, pytorch-tokenizers and coremltools (all hard runtime deps) are already on pypi.riseproject.dev for riscv64 cp312/cp313/cp314; torchao>=0.18.0 resolves to its py3-none-any wheel there (no riscv64 tag exists on PyPI, but its own x86_64 wheel doesn't match this platform anyway, so pip falls back to the pure-Python one). cp310/cp311 are dropped: torch and pytorch-tokenizers ship neither for riscv64. The smoke test replaces upstream's torchvision-based MobileNetV3 example (torchvision has no riscv64 wheel) with a small hand-built model, keeping upstream's own registered-backend assertions and exercising both the portable (non-delegated) and XNNPACK-delegated export+execute paths for real. A small packaging patch widens license-files so the wheel also carries the licences of the third-party libraries statically linked into its extensions (XNNPACK, cpuinfo, pthreadpool, FP16, FXdiv, flatbuffers, flatcc, Eigen, pocketfft, nlohmann/json, pybind11) -- upstream's own wheel ships only its own LICENSE.
|
- Fold the packaging patch's Upstream-Status: Inappropriate reason onto one line: check_patch.py's regex has no re.DOTALL, so it only reads the tag's first physical line, and a bracket closed two lines down never reaches the format check. - Add CIBW_BEFORE_ALL_LINUX: dnf -y install openssl-devel. No riscv64 cmake wheel is published yet for the pinned version, so pip builds it from source; its own bootstrap's Utilities/cmcurl needs OpenSSL's dev headers, which the manylinux_2_39_riscv64 image doesn't ship by default.
…us bracket must stay on one line Found while triaging PR #2434 (executorch): cibuildwheel building cmake from source (no riscv64 wheel for the pinned version) fails configure with "Could not find OpenSSL" until openssl-devel is installed; and check_patch.py's Upstream-Status regex has no re.DOTALL, so a bracketed Inappropriate reason wrapped across multiple lines fails the format check even though it reads as correctly bracketed.
cibuildwheel always mounts the source into its manylinux container at the fixed path /project, which trips CMakeLists.txt's hard requirement that the checkout be named exactly `executorch` (upstream tracks lifting this at pytorch/executorch#6475). Add a second patch that provides the same `<executorch/...>` include path through a configure-time symlink instead, so the build no longer depends on the checkout directory's name. Validated by applying the patch to a real v1.4.1 checkout and by exercising the isolated CMake snippet under both the original and patched forms: unpatched fails with the reported FATAL_ERROR when the directory isn't named `executorch`, and patched configures cleanly with the symlink resolving to the real source tree.
|
Root cause: Fix: squash this branch into a single commit (or Unrelated and already fixed on this same branch: the |
CMake's find_package_torch_headers/get_torch_base_path resolves torch by running `import torch` in the same interpreter running the build backend. torch is a runtime dependency (setup.py), not declared in [build-system].requires, so pip's isolated build venv never has it and CMake configure fails with AttributeError: 'NoneType' object has no attribute 'submodule_search_locations'. Preinstall build-system.requires plus torch in CIBW_BEFORE_BUILD and disable isolation via CIBW_BUILD_FRONTEND, mirroring the same pattern already used in build-torchaudio.yml and build-torchcodec.yml.
The riscv64 build now fails during CMake configure, not link:
CMake Error at .../FindPython/Support.cmake:4243 (message):
Python_ADD_LIBRARY: dependent target 'Python::Python' is not defined.
Did you miss to request COMPONENT 'Development.Embed'?
CMake Error at third-party/pybind11/tools/pybind11NewTools.cmake:276
(target_link_libraries):
Cannot specify link libraries for target "executorchcoreml" which is
not built by this project.
Every pybind11_add_module(<name> SHARED ...) call in upstream's CMakeLists
(portable_lib, data_loader, selective_build, _training_lib, _llm_runner,
executorchcoreml) makes CMake's Python_add_library() require the
Python::Python target (COMPONENT Development.Embed) to exist at configure
time. quay.io/pypa/manylinux_2_39_riscv64, like every vanilla pypa
manylinux image, ships only a static libpythonX.Y.a and never satisfies
Development.Embed, so find_package(Python) never defines Python::Python
and configure aborts on the first such call it reaches (coreml's, by
subdirectory order) -- the second CMake error above is just fallout from
the first one leaving that target half-defined.
These are all ordinary dlopen()'d Python extension modules -- upstream's
own strip_python_lib() helper immediately strips the resulting
Python::Python/pybind11::embed link dependency right back out after
creating them, precisely because they must resolve Python symbols from
the running interpreter rather than linking a build-time libpython.
Dropping SHARED (pybind11's own default, MODULE) only requires
Development.Module, which manylinux images do provide, and produces the
same importable extension without ever needing Development.Embed.
Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since this is a real portability bug on any manylinux-like
environment lacking an embeddable libpython, not riscv64-specific.
The bracketed comment on the To upstream line spanned 5 lines, so check_patch.py's single-line regex only captured the first line's text (an unclosed '[') and rejected the format. Reflow it onto a single line per this repo's Upstream-Status convention, keeping the same substance: not yet submitted, a real portability bug on any vanilla manylinux image lacking Development.Embed, independent of riscv64, and likely unseen upstream because pytorch/executorch's own CI runs a custom manylinux-builder image that does provide an embeddable libpython.
Configure now fails past everything 0003 fixed, on a problem specific to
one of the six pybind targets it converted to MODULE:
CMake Error at extension/llm/runner/CMakeLists.txt:122
(target_link_libraries):
Target "portable_lib" of type MODULE_LIBRARY may not be linked into
another target. One may link only to INTERFACE, OBJECT, STATIC or
SHARED libraries, or to executables with the ENABLE_EXPORTS property
set.
extension/llm/runner/CMakeLists.txt does
target_link_libraries(_llm_runner PRIVATE ... portable_lib ...) --
portable_lib is both a Python extension module and a library another
CMake target links against, unlike the other five 0003 touched. CMake
refuses to target_link_libraries() a MODULE_LIBRARY into anything else no
matter how it was created; there is no pybind11_add_module() argument that
opts out of this, so portable_lib has to go back to SHARED and the
Development.Embed gap has to be closed a different way.
FindPython's __Python_ADD_LIBRARY() (Modules/FindPython/Support.cmake)
requires a Python::Python target to already exist before it will
add_library(... SHARED ...), then links the new target against it --
but nothing requires that target to carry a real embeddable libpython.
Defining Python::Python ourselves as a link-free INTERFACE IMPORTED
target when Development.Embed didn't find one satisfies the existence
check; strip_python_lib() (already applied to portable_lib) then strips
it back out of portable_lib's real link line exactly as it already does
for the other five, so no embeddable libpython is ever actually linked.
This reproduces upstream's own working configuration, since SHARED is
what pybind11_add_module(portable_lib ...) already used before 0003, on
their manylinux-builder images that do have Development.Embed.
Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since this is the same class of portability bug as 0003,
not riscv64-specific.
The second independent configure failure past 0003 is unrelated to
pybind/MODULE-vs-SHARED entirely:
CMake Error at extension/llm/custom_ops/CMakeLists.txt:61 (message):
Unsupported CMAKE_SYSTEM_PROCESSOR riscv64. (If 32-bit x86, try using
fht_avx.c and send a PR if it works!)
extension/llm/custom_ops/CMakeLists.txt picks a vendored FFHT (Fast
Hadamard Transform) SIMD kernel source file based on
CMAKE_SYSTEM_PROCESSOR -- fht_neon.c for aarch64/arm64/armv7, fht_avx.c
for x86_64/AMD64 -- and hard-errors for anything else.
But nothing on riscv64 needs an FFHT SIMD kernel at all.
extension/llm/custom_ops/spinquant/fast_hadamard_transform.h's
fast_hadamard_transform_ffht_impl() -- the only caller of fht_float()/
fht_double() anywhere outside the vendored FFHT sources themselves --
branches on compiler-defined architecture macros, not
CMAKE_SYSTEM_PROCESSOR:
#if defined(__aarch64__) || defined(__x86_64__)
fht_float(vec, log2_vec_size);
#else
fast_hadamard_transform_simple_impl(vec, log2_vec_size);
#endif
so on riscv64 the portable, scalar fast_hadamard_transform_simple_impl()
already runs instead, and fht_float()/fht_double() are never referenced.
The CMakeLists.txt FATAL_ERROR was refusing to configure over a kernel
source file that would never have been needed. Replace it with a STATUS
message and add no FFHT source for these architectures -- the vendored
fht_avx.c/fht_neon.c/fht_sse.c/dumb_fht.c files use a different, narrower
API regardless (dumb_fht.c is void dumb_fht(float*, int), not
fht_float()'s int fht_float(float*, int), with no fht_double()/_oop
equivalents at all), so none of them was ever a drop-in match here.
Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since the FATAL_ERROR blocks every architecture outside
x86_64/aarch64/armv7 even though a portable fallback needing none of them
already exists, not riscv64-specific.
The cp314 leg's build step got past CMake configure (0003/0004/0005) and failed at actual compilation instead, in two clusters: no matching constructor for c10::complex<c10::BFloat16>::complex(c10::complex<float>), and "expected ',' or '...' before C10_LIFETIMEBOUND" across ArrayRef.h/ TensorAccessor.h/Tensor.h. torch_pin.py pins TORCH_VERSION="2.13.0", the release executorch's vendored runtime/core/portable_type/c10 header snapshot is frozen against. CIBW_BEFORE_BUILD copied setup.py's own unbounded "torch>=2.13.0a0" floor, so once pypi.riseproject.dev started serving torch 2.14.0 for every interpreter, pip resolved that instead of 2.13.0 (confirmed from the job log: "Successfully installed ... torch-2.14.0+cpu"). The real 2.14.0 headers and the vendored 2.13.0-era copy then collide in the same translation unit: 2.14.0's c10::complex<T> gained constructors the vendored, 2.13.0-era class doesn't have, and C10_LIFETIMEBOUND (added between 2.13.0 and 2.14.0, confirmed by diffing torch/headeronly/macros/ Macros.h between the two tags) resolves to the vendored, pre-macro copy and is left undefined. torch 2.13.0 has a riscv64 wheel on our registry for every interpreter in this matrix, so this is a pin problem, not an unsupported-interpreter one (gotcha 615) - pin the build-time install to exactly what torch_pin.py names instead of dropping cp314, and pin CIBW_TEST_REQUIRES the same way so the test install doesn't resolve a newer torch than the extension was actually compiled against.
Compiling now succeeds (torch pinned to 2.13.0), but the default CIBW_REPAIR_WHEEL_COMMAND then fails: auditwheel can't locate libc10.so because it isn't on any path it searches. libtorch/libc10/libgomp etc. are provided by the torch wheel at runtime (torch is imported before executorch and loads them RTLD_GLOBAL), so exclude them the same way build-torchaudio.yml and build-vllm.yml do (gotcha 17).
Restrict the matrix to cp312 only, keep debug symbols (CFLAGS/CXXFLAGS=-g and CMAKE_ARGS=-DCMAKE_BUILD_TYPE=RelWithDebInfo, matching build-tesserocr.yml's precedent), install gdb, and run the smoke test under gdb to capture a native backtrace of the deterministic cpuinfo_get_uarch abort. To be reset once the backtrace is read.
executorch1.4.1PyTorch's on-device inference runtime: a setuptools + CMake pybind11 extension with the XNNPACK CPU delegate, portable/optimized/quantized kernels, the OpenVINO backend and CoreML's portable pieces. Upstream publishes no riscv64 wheel.
Mirrors upstream's
build-wheels-linux.yml, narrowed to the default--preset pybindCMake config it builds with.Differs from upstream
Matrix: cp312/cp313/cp314 only - torch and pytorch-tokenizers (hard runtime deps) ship neither cp310 nor cp311 for riscv64; upstream itself ships no cp314t
Testing
License: Wheel bundles XNNPACK/cpuinfo/pthreadpool/FP16/FXdiv/flatbuffers/flatcc/Eigen/pocketfft/nlohmann-json/pybind11 (BSD/MIT/Apache-2.0/MPL-2.0); upstream ships no licence text for them, so the build adds it.
Patches
0001-packaging-Include-vendored-third-party-licences-in-.patch- Inappropriate, only relevant to a distributor bundling vendored licences. Reproduces on any arch.