Skip to content

executorch: Add version 1.4.1 - #2434

Draft
luhenry wants to merge 11 commits into
mainfrom
executorch
Draft

luhenry wants to merge 11 commits into
mainfrom
executorch

Conversation

@luhenry

@luhenry luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

PyTorch's on-device inference runtime: a setuptools + CMake pybind11 extension with the XNNPACK CPU delegate, portable/optimized/quantized kernels, the OpenVINO backend and CoreML's portable pieces. Upstream publishes no riscv64 wheel.

Mirrors upstream's build-wheels-linux.yml, narrowed to the default --preset pybind CMake config it builds with.

Differs from upstream

  • Nothing beyond the riscv64 image.

Matrix: cp312/cp313/cp314 only - torch and pytorch-tokenizers (hard runtime deps) ship neither cp310 nor cp311 for riscv64; upstream itself ships no cp314t

Testing

  • Replaced upstream's torchvision-based MobileNetV3 example with a hand-built model - torchvision has no riscv64 wheel

License: Wheel bundles XNNPACK/cpuinfo/pthreadpool/FP16/FXdiv/flatbuffers/flatcc/Eigen/pocketfft/nlohmann-json/pybind11 (BSD/MIT/Apache-2.0/MPL-2.0); upstream ships no licence text for them, so the build adds it.

Patches

  • 0001-packaging-Include-vendored-third-party-licences-in-.patch - Inappropriate, only relevant to a distributor bundling vendored licences. Reproduces on any arch.

Mirrors upstream's own default pybind wheel build (setup.py + CMake
--preset pybind): XNNPACK CPU delegate, portable/optimized/quantized
kernels, CoreML's portable util+inmemoryfs+pybind pieces (no Apple
delegate on Linux), OpenVINO backend (dlopen-based, no build-time SDK).
QNN, CUDA and Vulkan stay off, same as upstream's own CI on riscv64/ARM
hosts (QNN auto-skips under GITHUB_ACTIONS; CUDA/Vulkan are opt-in and
neither is requested).

torch, pytorch-tokenizers and coremltools (all hard runtime deps) are
already on pypi.riseproject.dev for riscv64 cp312/cp313/cp314;
torchao>=0.18.0 resolves to its py3-none-any wheel there (no riscv64
tag exists on PyPI, but its own x86_64 wheel doesn't match this
platform anyway, so pip falls back to the pure-Python one). cp310/cp311
are dropped: torch and pytorch-tokenizers ship neither for riscv64.

The smoke test replaces upstream's torchvision-based MobileNetV3
example (torchvision has no riscv64 wheel) with a small hand-built
model, keeping upstream's own registered-backend assertions and
exercising both the portable (non-delegated) and XNNPACK-delegated
export+execute paths for real.

A small packaging patch widens license-files so the wheel also carries
the licences of the third-party libraries statically linked into its
extensions (XNNPACK, cpuinfo, pthreadpool, FP16, FXdiv, flatbuffers,
flatcc, Eigen, pocketfft, nlohmann/json, pybind11) -- upstream's own
wheel ships only its own LICENSE.
luhenry added a commit that referenced this pull request Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

QR code for preview link

🚀 View preview at
https://riseproject-dev.github.io/python-wheels/pr-preview/pr-2434/

Built to branch gh-pages at 2026-09-30 10:02 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

- Fold the packaging patch's Upstream-Status: Inappropriate reason onto
  one line: check_patch.py's regex has no re.DOTALL, so it only reads
  the tag's first physical line, and a bracket closed two lines down
  never reaches the format check.
- Add CIBW_BEFORE_ALL_LINUX: dnf -y install openssl-devel. No riscv64
  cmake wheel is published yet for the pinned version, so pip builds
  it from source; its own bootstrap's Utilities/cmcurl needs OpenSSL's
  dev headers, which the manylinux_2_39_riscv64 image doesn't ship by
  default.
luhenry added a commit that referenced this pull request Sep 28, 2026
…us bracket must stay on one line

Found while triaging PR #2434 (executorch): cibuildwheel building cmake
from source (no riscv64 wheel for the pinned version) fails configure
with "Could not find OpenSSL" until openssl-devel is installed; and
check_patch.py's Upstream-Status regex has no re.DOTALL, so a bracketed
Inappropriate reason wrapped across multiple lines fails the format
check even though it reads as correctly bracketed.
cibuildwheel always mounts the source into its manylinux container at
the fixed path /project, which trips CMakeLists.txt's hard requirement
that the checkout be named exactly `executorch` (upstream tracks
lifting this at pytorch/executorch#6475). Add a second patch that
provides the same `<executorch/...>` include path through a
configure-time symlink instead, so the build no longer depends on the
checkout directory's name.

Validated by applying the patch to a real v1.4.1 checkout and by
exercising the isolated CMake snippet under both the original and
patched forms: unpatched fails with the reported FATAL_ERROR when the
directory isn't named `executorch`, and patched configures cleanly
with the symlink resolving to the real source tree.

luhenry commented Sep 28, 2026 •

Copy link
Copy Markdown
Member Author

check_patches is red and will stay red until this branch's history is rewritten — no fix-forward commit can clear it.

Root cause: ci_scripts/check_patch.py validates every commit in the PR range independently, checking each commit's own snapshot of any patch file it touched (for commit in commits: ... check_upstream_status(...)). Commit dddfb9e83d62 introduced patches/executorch/1.4.1/0001-packaging-Include-vendored-third-party-licences-in-.patch with its Upstream-Status: Inappropriate reason wrapped onto a second physical line. The script's regex has no re.DOTALL, so it only reads the tag's first line, sees no closing bracket, and fails that commit's snapshot. Commit 161c3ab77cc later folded the header onto one line — which fixes the current file content and would pass check_patches on its own — but dddfb9e83d62's original snapshot is still in history, so the job still fails on it.

Fix: squash this branch into a single commit (or rebase -i to amend dddfb9e83d62 so the patch never carries the malformed header at any point in history), then force-push. A plain follow-up commit cannot fix this.

Unrelated and already fixed on this same branch: the cmake source-directory-name build failure (patch 0002-cmake-relax-source-directory-name-check-for-cibuild.patch, commit 53206a23d4dc) — that one only needed a normal fix-forward commit and is unaffected by the above.

CMake's find_package_torch_headers/get_torch_base_path resolves torch by
running `import torch` in the same interpreter running the build backend.
torch is a runtime dependency (setup.py), not declared in
[build-system].requires, so pip's isolated build venv never has it and
CMake configure fails with AttributeError: 'NoneType' object has no
attribute 'submodule_search_locations'.

Preinstall build-system.requires plus torch in CIBW_BEFORE_BUILD and
disable isolation via CIBW_BUILD_FRONTEND, mirroring the same pattern
already used in build-torchaudio.yml and build-torchcodec.yml.
The riscv64 build now fails during CMake configure, not link:

  CMake Error at .../FindPython/Support.cmake:4243 (message):
    Python_ADD_LIBRARY: dependent target 'Python::Python' is not defined.
      Did you miss to request COMPONENT 'Development.Embed'?
  CMake Error at third-party/pybind11/tools/pybind11NewTools.cmake:276
  (target_link_libraries):
    Cannot specify link libraries for target "executorchcoreml" which is
    not built by this project.

Every pybind11_add_module(<name> SHARED ...) call in upstream's CMakeLists
(portable_lib, data_loader, selective_build, _training_lib, _llm_runner,
executorchcoreml) makes CMake's Python_add_library() require the
Python::Python target (COMPONENT Development.Embed) to exist at configure
time. quay.io/pypa/manylinux_2_39_riscv64, like every vanilla pypa
manylinux image, ships only a static libpythonX.Y.a and never satisfies
Development.Embed, so find_package(Python) never defines Python::Python
and configure aborts on the first such call it reaches (coreml's, by
subdirectory order) -- the second CMake error above is just fallout from
the first one leaving that target half-defined.

These are all ordinary dlopen()'d Python extension modules -- upstream's
own strip_python_lib() helper immediately strips the resulting
Python::Python/pybind11::embed link dependency right back out after
creating them, precisely because they must resolve Python symbols from
the running interpreter rather than linking a build-time libpython.
Dropping SHARED (pybind11's own default, MODULE) only requires
Development.Module, which manylinux images do provide, and produces the
same importable extension without ever needing Development.Embed.

Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since this is a real portability bug on any manylinux-like
environment lacking an embeddable libpython, not riscv64-specific.
The bracketed comment on the To upstream line spanned 5 lines, so
check_patch.py's single-line regex only captured the first line's text
(an unclosed '[') and rejected the format. Reflow it onto a single
line per this repo's Upstream-Status convention, keeping the same
substance: not yet submitted, a real portability bug on any vanilla
manylinux image lacking Development.Embed, independent of riscv64, and
likely unseen upstream because pytorch/executorch's own CI runs a
custom manylinux-builder image that does provide an embeddable
libpython.
Configure now fails past everything 0003 fixed, on a problem specific to
one of the six pybind targets it converted to MODULE:

  CMake Error at extension/llm/runner/CMakeLists.txt:122
  (target_link_libraries):
    Target "portable_lib" of type MODULE_LIBRARY may not be linked into
    another target. One may link only to INTERFACE, OBJECT, STATIC or
    SHARED libraries, or to executables with the ENABLE_EXPORTS property
    set.

extension/llm/runner/CMakeLists.txt does
target_link_libraries(_llm_runner PRIVATE ... portable_lib ...) --
portable_lib is both a Python extension module and a library another
CMake target links against, unlike the other five 0003 touched. CMake
refuses to target_link_libraries() a MODULE_LIBRARY into anything else no
matter how it was created; there is no pybind11_add_module() argument that
opts out of this, so portable_lib has to go back to SHARED and the
Development.Embed gap has to be closed a different way.

FindPython's __Python_ADD_LIBRARY() (Modules/FindPython/Support.cmake)
requires a Python::Python target to already exist before it will
add_library(... SHARED ...), then links the new target against it --
but nothing requires that target to carry a real embeddable libpython.
Defining Python::Python ourselves as a link-free INTERFACE IMPORTED
target when Development.Embed didn't find one satisfies the existence
check; strip_python_lib() (already applied to portable_lib) then strips
it back out of portable_lib's real link line exactly as it already does
for the other five, so no embeddable libpython is ever actually linked.
This reproduces upstream's own working configuration, since SHARED is
what pybind11_add_module(portable_lib ...) already used before 0003, on
their manylinux-builder images that do have Development.Embed.

Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since this is the same class of portability bug as 0003,
not riscv64-specific.
The second independent configure failure past 0003 is unrelated to
pybind/MODULE-vs-SHARED entirely:

  CMake Error at extension/llm/custom_ops/CMakeLists.txt:61 (message):
    Unsupported CMAKE_SYSTEM_PROCESSOR riscv64. (If 32-bit x86, try using
    fht_avx.c and send a PR if it works!)

extension/llm/custom_ops/CMakeLists.txt picks a vendored FFHT (Fast
Hadamard Transform) SIMD kernel source file based on
CMAKE_SYSTEM_PROCESSOR -- fht_neon.c for aarch64/arm64/armv7, fht_avx.c
for x86_64/AMD64 -- and hard-errors for anything else.

But nothing on riscv64 needs an FFHT SIMD kernel at all.
extension/llm/custom_ops/spinquant/fast_hadamard_transform.h's
fast_hadamard_transform_ffht_impl() -- the only caller of fht_float()/
fht_double() anywhere outside the vendored FFHT sources themselves --
branches on compiler-defined architecture macros, not
CMAKE_SYSTEM_PROCESSOR:

  #if defined(__aarch64__) || defined(__x86_64__)
    fht_float(vec, log2_vec_size);
  #else
    fast_hadamard_transform_simple_impl(vec, log2_vec_size);
  #endif

so on riscv64 the portable, scalar fast_hadamard_transform_simple_impl()
already runs instead, and fht_float()/fht_double() are never referenced.
The CMakeLists.txt FATAL_ERROR was refusing to configure over a kernel
source file that would never have been needed. Replace it with a STATUS
message and add no FFHT source for these architectures -- the vendored
fht_avx.c/fht_neon.c/fht_sse.c/dumb_fht.c files use a different, narrower
API regardless (dumb_fht.c is void dumb_fht(float*, int), not
fht_float()'s int fht_float(float*, int), with no fht_double()/_oop
equivalents at all), so none of them was ever a drop-in match here.

Patch under patches/executorch/1.4.1/ per the porting skill, marked
"To upstream" since the FATAL_ERROR blocks every architecture outside
x86_64/aarch64/armv7 even though a portable fallback needing none of them
already exists, not riscv64-specific.
The cp314 leg's build step got past CMake configure (0003/0004/0005) and
failed at actual compilation instead, in two clusters: no matching
constructor for c10::complex<c10::BFloat16>::complex(c10::complex<float>),
and "expected ',' or '...' before C10_LIFETIMEBOUND" across ArrayRef.h/
TensorAccessor.h/Tensor.h.

torch_pin.py pins TORCH_VERSION="2.13.0", the release executorch's
vendored runtime/core/portable_type/c10 header snapshot is frozen against.
CIBW_BEFORE_BUILD copied setup.py's own unbounded "torch>=2.13.0a0" floor,
so once pypi.riseproject.dev started serving torch 2.14.0 for every
interpreter, pip resolved that instead of 2.13.0 (confirmed from the job
log: "Successfully installed ... torch-2.14.0+cpu"). The real 2.14.0
headers and the vendored 2.13.0-era copy then collide in the same
translation unit: 2.14.0's c10::complex<T> gained constructors the
vendored, 2.13.0-era class doesn't have, and C10_LIFETIMEBOUND (added
between 2.13.0 and 2.14.0, confirmed by diffing torch/headeronly/macros/
Macros.h between the two tags) resolves to the vendored, pre-macro copy
and is left undefined.

torch 2.13.0 has a riscv64 wheel on our registry for every interpreter in
this matrix, so this is a pin problem, not an unsupported-interpreter one
(gotcha 615) - pin the build-time install to exactly what torch_pin.py
names instead of dropping cp314, and pin CIBW_TEST_REQUIRES the same way
so the test install doesn't resolve a newer torch than the extension was
actually compiled against.
Compiling now succeeds (torch pinned to 2.13.0), but the default
CIBW_REPAIR_WHEEL_COMMAND then fails: auditwheel can't locate libc10.so
because it isn't on any path it searches. libtorch/libc10/libgomp etc.
are provided by the torch wheel at runtime (torch is imported before
executorch and loads them RTLD_GLOBAL), so exclude them the same way
build-torchaudio.yml and build-vllm.yml do (gotcha 17).
Restrict the matrix to cp312 only, keep debug symbols (CFLAGS/CXXFLAGS=-g
and CMAKE_ARGS=-DCMAKE_BUILD_TYPE=RelWithDebInfo, matching build-tesserocr.yml's
precedent), install gdb, and run the smoke test under gdb to capture a native
backtrace of the deterministic cpuinfo_get_uarch abort. To be reset once the
backtrace is read.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant