You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 2994d28
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: .queue.yml
+4-4Lines changed: 4 additions & 4 deletions
Original file line number
Diff line number
Diff line change
@@ -145,7 +145,7 @@ packages:
145
145
repo: https://www.tensorflow.org/
146
146
status: parked
147
147
pr: null
148
-
notes: '8 Linux wheels upstream (abi: cp310,cp311,cp312,cp313); no riscv64 on PyPI or pypi.riseproject.dev.'
148
+
notes: '8 Linux wheels upstream (abi: cp310,cp311,cp312,cp313); no riscv64 on PyPI or pypi.riseproject.dev. The original park reason was never recorded here; re-derived 2026-09-20 while triaging the tensorflow-cpu sibling, whose entry carries the full evidence (and gotcha 426). Summary, stated as a re-derivation rather than as the original reasoning: this is the enormous Bazel-driven C++/XLA/MLIR build TensorFlow always is, and it is disproportionate rather than hard-blocked. The wheel is per-interpreter (cp{v}-cp{v} from HERMETIC_PYTHON_VERSION), so cp310-cp313 is 4 full Bazel builds with no abi3/py3-none collapse; PR #2104 (mediapipe) compiles only the framework/TFLite subset of this very TF 2.21.0 source, with no XLA/MLIR/LLVM, and still spent 5h23m in its build step plus 2h24m-3h21m of queue wait per job on the shared riscv64 pool, against a repo record of ~10h (libclang) and a repo-wide job cap of 1440min (build-torch.yml). Upstream maps riscv64 nowhere in its own build: zero riscv occurrences in .bazelrc, release_cpu_linux pulls in -mavx, WORKSPACE registers only linux_x86_64/linux_aarch64 CC toolchains, and the platform_tag select in tensorflow/tools/pip_package/BUILD has no riscv64 entry and no //conditions:default, so the wheel target fails at analysis time. There is also no sdist and no upstream riscv64 CI to narrow (goal 2). Note for anyone reopening this: riscv64 has no CUDA, so a riscv64 tensorflow wheel would be the CPU-only build - which is exactly what upstream already ships under this plain name on aarch64 (268MB, matching x86_64 tensorflow_cpu at 261MB rather than x86_64 tensorflow at 545MB). That makes `tensorflow` the correct distribution name for a riscv64 port, and tensorflow-cpu the wrong one; if this is ever unparked, port this entry and leave tensorflow-cpu parked. Not the jaxlib/PR #526 blocker: at the XLA revision this tag vendors, xla/backends/cpu/codegen/BUILD does wire if_llvm_riscv_available to llvm:RISCVCodeGen, and the bazel 7.7.0 bootstrap is de-risked by gotchas 47/133.'
notes: 'Upstream located: mitmproxy/mitmproxy_rs (same Cargo workspace as the already-published mitmproxy-rs), member crate mitmproxy-linux/. Not a closed-binary redistribution -- it is a maturin `bindings = "bin"` project (gotcha 117): the wheel is one compiled executable, mitmproxy-linux-redirector, plus a 40-line mitmproxy_linux/__init__.py that only returns its path, which is why PyPI shows py3-none-<platform> and no project URL. Checkout layout follows gotcha 239: mitmproxy-linux/pyproject.toml exists in the git tree (no manifest-path key) but the PyPI sdist hoists it to the root, so the workflow builds with --manifest-path mitmproxy-linux/Cargo.toml. The interesting part is the eBPF half: build.rs uses aya-build to cross-compile mitmproxy-linux-ebpf to bpfel-unknown-none, which needs a nightly toolchain with rust-src plus the bpf-linker binary -- installed in-container in before-script-linux, mirroring upstream .github/actions/setup. No riscv64 wall there: bpf-linker 0.9.15 defaults to rust-llvm-21/aya-rustc-llvm-proxy (no system LLVM, it dlopens rustc''s own), rustup ships a nightly riscv64gc-unknown-linux-gnu host toolchain, and aya-build explicitly lists riscv64 as a bpf_target_arch (with a riscv64gc -> riscv64 fixup). Validated locally before pushing: the exact recipe (stable+nightly+rust-src, cargo install --locked bpf-linker@0.9.15, maturin build --release --locked) built a correct mitmproxy_linux-0.12.11-py3-none-manylinux_2_34_x86_64.whl natively in manylinux_2_34_x86_64 in ~20 min; the embedded eBPF object and its cgroup/sock_create program were confirmed present and byte-shape-identical to upstream''s published x86_64 wheel. A native riscv64 rehearsal under qemu (manylinux_2_39_riscv64, gotcha 188 LTO override) was also run. LICENSE staged into mitmproxy-linux/ so maturin bundles it -- without it the wheel ships no licence text (upstream''s own aarch64 wheel does not). Test job cannot run the redirector end to end (needs root + tun + cgroup eBPF attach; upstream gates that behind its own root-tests feature), so it asserts the riscv64 ELF carries an EM_BPF object with the cgroup/sock_create section and INTERCEPT_CONF map, then that the binary''s usage path works. Follow-up once published: build-mitmproxy-rs.yml currently passes --features docs to skip lib.rs''s eager `import mitmproxy_linux`; with a riscv64 mitmproxy_linux on the registry that workaround (and the PIP_NO_DEPS=1 test env) can be dropped. PR #2119 opened, CI running.'
1996
+
notes: 'Upstream located: mitmproxy/mitmproxy_rs (same Cargo workspace as the already-published mitmproxy-rs), member crate mitmproxy-linux/. Not a closed-binary redistribution -- it is a maturin `bindings = "bin"` project (gotcha 117): the wheel is one compiled executable, mitmproxy-linux-redirector, plus a 40-line mitmproxy_linux/__init__.py that only returns its path, which is why PyPI shows py3-none-<platform> and no project URL. Checkout layout follows gotcha 239: mitmproxy-linux/pyproject.toml exists in the git tree (no manifest-path key) but the PyPI sdist hoists it to the root, so the workflow builds with --manifest-path mitmproxy-linux/Cargo.toml. The interesting part is the eBPF half: build.rs uses aya-build to cross-compile mitmproxy-linux-ebpf to bpfel-unknown-none, which needs a nightly toolchain with rust-src plus the bpf-linker binary -- installed in-container in before-script-linux, mirroring upstream .github/actions/setup. No riscv64 wall there: bpf-linker 0.9.15 defaults to rust-llvm-21/aya-rustc-llvm-proxy (no system LLVM, it dlopens rustc''s own), rustup ships a nightly riscv64gc-unknown-linux-gnu host toolchain, and aya-build explicitly lists riscv64 as a bpf_target_arch (with a riscv64gc -> riscv64 fixup). Validated locally before pushing: the exact recipe (stable+nightly+rust-src, cargo install --locked bpf-linker@0.9.15, maturin build --release --locked) built a correct mitmproxy_linux-0.12.11-py3-none-manylinux_2_34_x86_64.whl natively in manylinux_2_34_x86_64 in ~20 min; the embedded eBPF object and its cgroup/sock_create program were confirmed present and byte-shape-identical to upstream''s published x86_64 wheel. A native riscv64 rehearsal under qemu (manylinux_2_39_riscv64, gotcha 188 LTO override) was also run. LICENSE staged into mitmproxy-linux/ so maturin bundles it -- without it the wheel ships no licence text (upstream''s own aarch64 wheel does not). Test job cannot run the redirector end to end (needs root + tun + cgroup eBPF attach; upstream gates that behind its own root-tests feature), so it asserts the riscv64 ELF carries an EM_BPF object with the cgroup/sock_create section and INTERCEPT_CONF map, then that the binary''s usage path works. Follow-up once published: build-mitmproxy-rs.yml currently passes --features docs to skip lib.rs''s eager `import mitmproxy_linux`; with a riscv64 mitmproxy_linux on the registry that workaround (and the PIP_NO_DEPS=1 test env) can be dropped. PR #2119 opened, CI running. CI disproved the eBPF assumption above: the first riscv64 run failed in build.rs with bpf-linker aborting (SIGABRT) in aya-rustc-llvm-proxy, "unable to find LLVM shared lib" -- it dlopens a libLLVM* shared object found via LD_LIBRARY_PATH or a PATH entry sibling lib/, and the riscv64 rustc dist ships none (LLVM is static inside librustc_driver with its C API hidden), where x86_64 and aarch64 carry libLLVM-<major>-rust-*.so; that is also why the x86_64 rehearsal passed, and the qemu riscv64 rehearsal never reached the link step (its dist/ stayed empty). Rebuilding bpf-linker against a system LLVM does not substitute: the newest LLVM packaged for riscv64 is 21.1.8 (Rocky 10.2 AppStream) while this workspace MSRV of 1.95 is the Rust release that moved to LLVM 22, and an LLVM 21 bpf-linker rejects newer bitcode with "ERROR llvm: Invalid record". Fixed in the PR by cross-compiling the object on ubuntu-latest with bpf_target_arch="riscv64" and embedding it in the riscv64 build through patch 0001 (gotcha 425); the riscv64 job now installs neither a nightly toolchain nor bpf-linker. Rehearsed end to end on x86_64 under Rocky 10 with bpf-linker removed from PATH: patch applied, object staged, maturin wheel built, and every assertion the test job makes passing. Watching CI.'
notes: '4 Linux wheels upstream (abi: cp310,cp311,cp312,cp313); no riscv64 on PyPI or pypi.riseproject.dev.'
2199
+
notes: '4 Linux wheels upstream (abi: cp310,cp311,cp312,cp313); no riscv64 on PyPI or pypi.riseproject.dev. Parked for two independent reasons - wrong distribution name for this architecture, and disproportionate scope - and explicitly NOT Bazel-blocked the way jaxlib/PR #526 is. Folded into the skill as gotcha 426 (feasibility-and-triage.md). NAMING: tensorflow-cpu is the same upstream tree as the already-parked tensorflow, selected by one repo_env - tensorflow/tools/pip_package/utils/tf_wheel.bzl reads WHEEL_NAME from @python_version_repo and its own docstring says it "Should be set via --repo_env=WHEEL_NAME=tensorflow_cpu" - and the two distributions PyPI metadata is identical down to the same 12 nvidia-* "extra == and-cuda" requirements. Across its entire release history tensorflow-cpu has published win_amd64 and manylinux*_x86_64 only: zero aarch64, zero ppc64le, zero other Linux arch, and no sdist. It exists solely to give x86_64 a CUDA-less wheel, because on x86_64 the default tensorflow wheel is the GPU one (545MB x86_64 tensorflow vs 261MB x86_64 tensorflow_cpu). On aarch64, where there is no CUDA to ship, upstream ships the CPU-only build under the plain name tensorflow - 268MB, matching the cpu wheel size rather than its own arch GPU wheel. riscv64 sits in exactly the aarch64 position, so the riscv64 deliverable for CPU-only TensorFlow is `tensorflow`, and a manylinux_riscv64 wheel named tensorflow-cpu would invent a name upstream uses on exactly one Linux architecture (gotcha 50, divergence with no gap closed, reached from the sibling side per gotcha 79/426). This also settles the plausible-sounding hypothesis that the CPU-only variant sidesteps whatever made full tensorflow impractical: it is backwards. With no CUDA on riscv64 the tensorflow build is already the CPU build, so tensorflow-cpu is the identical compile under a worse name and is never cheaper than the base. SCOPE: the wheel is per-interpreter (_get_full_wheel_name formats cp{v}-cp{v} from HERMETIC_PYTHON_VERSION), so cp310-cp313 means 4 full Bazel builds with none of the abi3/py3-none collapse. Calibrated against PR #2104 (mediapipe, which compiles this same TF 2.21.0 source vendored as org_tensorflow, on this same runner image): mediapipe builds only the framework/TFLite subset its calculators need plus XNNPACK and OpenCV, with no XLA, MLIR or LLVM, and still took 5h23m in the build step on ubuntu-24.04-riscv, with 2h24m-3h21m of queue wait per job on the shared pool and 5 CI runs to clear two arch-specific fixes. Full tensorflow adds all of the TF kernels/ops plus XLA plus MLIR plus a large slice of LLVM on top of that subset, times 4 interpreters. For reference the largest build logged in this repo is libclang at ~10h (LLVM/Clang from scratch) and the highest job cap anywhere in the repo is build-torch.yml at 1440min. NO riscv64 MAPPING IN UPSTREAM BUILD: .bazelrc has zero riscv occurrences across 1023 lines and the official CPU config release_cpu_linux pulls in --config=avx_linux (-mavx), i.e. x86-only by construction; WORKSPACE registers exactly four rules_ml_toolchain CC toolchains (linux_x86_64, linux_x86_64_cuda, linux_aarch64, linux_aarch64_cuda) so there is no hermetic riscv64 toolchain and --config=clang_local plus a local compiler would be mandatory; tensorflow/tools/pip_package/BUILD platform_tag select lists aarch64/arm64/x86_64/ppc with no riscv64 entry and no //conditions:default, so the wheel target fails at analysis time before a single object compiles; third_party/py/python_init_toolchains.bzl filters hermetic interpreter platforms to aarch64/x86_64 with the comment "Avoid obscure platforms for now just in case" (mitigated, since WORKSPACE passes default_python_version = system); and hermetic pip has to resolve hash-pinned requirements_lock_3_*.txt (numpy 2.1.3, scipy 1.14.1, h5py 3.13.0) on riscv64 - the scipy sdist hash IS in the lock so it resolves, but that means compiling scipy and h5py from source inside the build. Upstream publishes no sdist at all (wheels only) and has no riscv64 CI job, so goal 2 has no upstream recipe to narrow - every arch knob would be newly invented here. COUNTERPOINT RECORDED so a later attempt is not misled: at the XLA revision TF 2.21.0 vendors (the full tree at third_party/xla/), xla/backends/cpu/codegen/BUILD does wire if_llvm_riscv_available to @llvm-project//llvm:RISCVCodeGen for the CPU JIT and xla/tsl/BUILD defines linux_riscv64 and riscv64_or_cross, and the embed_bitcode restructure that blocked jaxlib 0.11.1 is not in this tree yet (the rule is cc_ir_header in cc_to_llvm_ir.bzl, whose ir_to_string tool deps are only llvm:Object and llvm:Support, no per-arch CodeGen). The bazel 7.7.0 bootstrap is likewise de-risked by gotchas 47/133 (no riscv64 bazel binary exists, dist-archive bootstrap works, and 7.7.x pins the same rules_python 0.33.2/rules_java 7.6.5 as the 7.5.0 recipe already proven in build-mediapipe.yml). So the wall here is scope and naming, not architecture. REVISIT only if upstream starts publishing a non-x86 Linux wheel under the -cpu name, or if tensorflow itself is unparked - in which case port `tensorflow`, not this entry. No worktree, branch or workflow opened: the call was fully resolvable read-only against upstream BUILD/bzl files at v2.21.0, the PyPI release history, and this repo own PR #526 and #2104 history.'
Copy file name to clipboardExpand all lines: skills/python-project-porting/references/gotchas-index.md
+19-1Lines changed: 19 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,6 +1,6 @@
1
1
# Gotchas index — router for the themed gotcha files
2
2
3
-
The porting gotchas (382 of them) live in [`references/gotchas/`](gotchas/), split by theme so only the relevant slice loads. Every gotcha keeps a **permanent number** cited elsewhere as "gotcha N" (and in workflow comments as "CLAUDE.md gotcha N"). Numbers are stable IDs — **not sequential**, and four are **reused** with different content (two each of 33, 55, 56, 57), disambiguated by theme below.
3
+
The porting gotchas (386 of them) live in [`references/gotchas/`](gotchas/), split by theme so only the relevant slice loads. Every gotcha keeps a **permanent number** cited elsewhere as "gotcha N" (and in workflow comments as "CLAUDE.md gotcha N"). Numbers are stable IDs — **not sequential**, and four are **reused** with different content (two each of 33, 55, 56, 57), disambiguated by theme below.
4
4
5
5
## How to find the gotcha you need
6
6
@@ -158,6 +158,12 @@ The porting gotchas (382 of them) live in [`references/gotchas/`](gotchas/), spl
158
158
is a dead stub — count the sources that branch globs, diff the built `.so` against the
159
159
published CUDA one, and call an op instead of trusting a "did the extension load" flag
160
160
(the xformers case).
161
+
-**426** — A `-cpu` sibling can be an *x86_64-only label* rather than a portable CPU variant:
162
+
where the base package's wheel is already CPU-only on every non-x86 arch, the sibling closes
163
+
no riscv64 gap, is never cheaper than the base, and inherits the base's park — scan the
164
+
sibling's whole release history for platform tags, compare the base's per-arch wheel sizes,
165
+
and re-verify any sibling-family blocker at the revision your target actually pins
0 commit comments