Repository navigation
test_gap_fetch_request_from_node_incoming_message SIGABRTs deterministically on pristine main, and is in no allowlist #7629
Description
Activity
Reproduced, and the cause is not GC — so the suggested
PERRY_GC_ZEAL/ from-space bisect would have been a dead end. Posting the diagnosis rather than a fix; the remaining work is architectural and overlaps #7656.What the abort actually is
The harness classifies it CRASH and discards the message. It is a Rust panic:
thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13: there is no reactor running, must be called from the context of a Tokio 1.x runtimeLine 911 is the
tokio::spawn(async move { … })inside theperry_ffi::spawn_blocking_with_reactorclosure that adopts the listener — i.e. the closure ran on a thread with no tokio reactor, which is precisely the failureperry_ffi_spawn_blocking_with_reactorexists to prevent (its doc names this exact panic and the v0.5.571 net regression batch it closed).It is link-configuration dependent, and here is the A/B
Same compiler binary, same test, same host. The only variable is whether
libperry_ext_*.aare present inPERRY_RUNTIME_DIR:libperry_ext_*.ain the runtime dirresult present PASS 3/3 removed SIGABRT 2/2 That is deterministic in both directions, which is why it reproduces 3/3 on a stock build:
run_parity_tests.shbuilds-p perry -p perry-runtime -p perry-stdlib -p perry-runtime-static -p perry-stdlib-staticand adds the fiveperry-ext-*crates only underPERRY_NO_AUTO_OPTIMIZE(line ~634). So the default run has no prebuilt ext archives, auto-optimize builds its own, and that is the arm that aborts.Why that produces a missing reactor
perry-stdlib'sperry_ffi_spawn_blocking_with_reactoris correct and not feature-gated — it spawns onasync_bridge::runtime(), the multi-thread runtime that has the reactor. The#[cfg(test)]shim inperry-ext-http/src/test_async_shims.rs(which doesinvoke(ctx)synchronously, and would produce exactly this symptom) is not in a release archive, so it is not the culprit either.What is left is the hazard already recorded for this family: an auto-built ext staticlib bundles its own copy of the runtime/stdlib, and that copy wins the link. The ext code then spawns onto a second
async_bridge::runtime()instance that nothing pumps — same statics, different instance — so the closure runs with no reactor in context. Consistent with both directions of the A/B: one prebuilt archive ⇒ one copy ⇒ pass.I did not confirm the duplicate-symbol layout in the emitted binary; that is the next concrete step (
nmthe auto-built archive for a secondasync_bridgeruntime static), and it is the same root as #7656 — nothing per-PR links theperry-ext-*crates, so this configuration is only ever exercised at tag time.On the procedural half
I have not added a
known_failures.jsonentry. The suggestion was conditioned on "if it is going to stay red", and this is a real, located bug in a configuration CI actually ships — suppressing it now would hide a genuine failure behind provenance. Happy to add the entry if you would rather have it tracked-and-skipped while the link issue is scheduled.Incidentally
The documented ad-hoc recipe (
-p perry -p perry-runtime-static -p perry-stdlib-static) cannot compile this test in either mode: with auto-optimize it fails to linkjs_http_client_request_end_full/js_http_client_request_write_full/js_http_incoming_message_set_encoding, and withPERRY_NO_AUTO_OPTIMIZE=1it fails on the zlib pump symbols unless the stdlib feature set matches the ext archives exactly. Worth a line in CLAUDE.md's build section, since "reproduce it ad hoc" is the first thing anyone will try.Root-caused on
main@db44b31b7. It is not a codegen bug and it is no longer aperry-devfailure — it is a release-profile link defect, and the mechanism is fully determined.First, one correction to the report
The issue says 3/3 reproducible under
--profile perry-dev. That is no longer true:build harness direct run perry-devPASS 2/2 len=55 match=true, exit 0 — byte-identical to node--release(thin LTO,codegen-units=1)CRASH 2/2, SIGABRT exit 134 So something between
38ff7ecccand now fixed theperry-devarm, and the release arm still dies. That matters becauserun_parity_tests.shbuilds--release, so theparitygate is the only thing that can see this — andparityis tag-gated, which is the hiding mechanism this issue already identifies.What actually aborts
thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13: there is no reactor running, must be called from the context of a Tokio 1.x runtimeA Rust panic under
panic=abort— hence SIGABRT rather than an output mismatch, hence CRASH rather than FAIL. The message isTcpListener::from_std's: it is running in a tokio context whose I/O reactor is not the one being driven.Why, exactly
libperry_ext_http.ais astaticlib, so it physically bundles perry-runtime/perry-stdlib and their shared deps. The link path already knows this and strips the bundled members —strip_bundled_runtime_from_well_known_lib→strip_bundled_shared_deps_from_well_known_lib→strip_duplicate_objects_from_well_known_lib— andbuild_and_run.rsstates the exact stake:a wrapper's bundled perry-runtime copy would win first-definition over stdlib's and split the runtime's mutable globals in two … spawned async tasks then starve because the event pump's wait-driver slot is registered in one copy and read from the other.
That is precisely what happens here. The strip runs in both profiles and reaches opposite conclusions:
# perry-dev [strip-dedup] libperry_ext_http.a: dropped 1676 bundled member(s) also provided by stdlib (shared transitive deps, fixed-point safe) [strip-dedup] _libperry_ext_http.a_nosharedeps.lib: localized wrapper-only globals in 1 member(s) # --release [strip-dedup] libperry_ext_http.a: keeping bundled alloc-…rcgu.o — needed by a kept sibling and not provided by stdlib [strip-dedup] libperry_ext_http.a: keeping bundled bytes-…rcgu.o — … [strip-dedup] libperry_ext_http.a: keeping bundled core-…rcgu.o — … (no "dropped N member(s)" line at all)The consequence, measured on the same source and the same test:
profile duplicate symbols at link ( libperry_ext_http.avslibperry_stdlib.a)result perry-dev11 passes --release579 SIGABRT The strip's criterion is "is this symbol also provided by stdlib?" Under thin LTO with
codegen-units=1the stdlib archive internalizes those symbols, so the criterion answers no, the ext copies are kept, and both copies load.perry-ext-httpthen binds to a secondasync_bridge::RUNTIME— a current-thread runtime (see the "unified single-thread async model" note incommon/async_bridge.rs) that only the main JS event loop drives, and it drives the other copy. Nothing ever drives the ext copy's reactor, so the firstTcpListener::from_stdon the accept path panics.Two things this rules out, both checked rather than assumed:
- Not the
#[cfg(test)]shims.perry-ext-http/src/test_async_shims.rsandperry-ext-net's define a synchronousperry_ffi_spawn_blocking_with_reactorthat would produce this exact panic, but both modules are#[cfg(test)]and cannot reach a user program. - Not a missing feature on the bundled runtime.
perry-ext-http/Cargo.tomlhasperry-runtime = { workspace = true, features = ["default", "stdlib"] }, so link: perry-ext-* staticlibs bundle a feature-stripped perry-runtime that wins the link — regex/Temporal silently degrade (blocks Express boot) #6303/runtime: ext staticlibs bundle perry-runtime and shadow perry-stdlib's real symbols with its no-op stubs (node:http servers get no tokio reactor) #6314 are satisfied; theext_crates_bundle_a_full_featured_perry_runtimeunit test's premise holds. - Also not the LLVM-tool availability the strip warns about — re-running the release link with
/opt/homebrew/opt/llvm/binonPATHandLLVM_SYS_221_PREFIXset produced the identical 579 duplicates.
Fix direction
The strip needs a criterion that survives LTO internalization — "provided by stdlib" has to mean "provided by the stdlib that will be linked", which under thin LTO is not the archive's exported symbol table. The precedent is already in-tree:
dedup_ui_lib_against_linked_libssolves the same problem for the UI lib by localizing every bundled global the linked stdlib/runtime also define, instead of trying to drop members — "sibling references rebind to the single linked copy (one set of runtime/std mutable state — the #5920 invariant)". That is applied only to the UI lib today; the ext archives take the drop-members path and are exactly where it fails.I have not landed that change: it alters the link path for every
perry-ext-*archive on Apple targets and needs validation well beyond this one test.Worth noting this is the concrete cost of #7656 — no per-PR gate links the
perry-ext-*crates, so a release-only link defect in the ext path is invisible until a tag.- Not the
Verified fixed on current
mainafter rebuilding the exact release-profile auto-optimize configurations, so I am closing this rather than landing speculative linker changes.I checked all six witnesses from the duplicate #7932 individually with the release compiler/linker at
f110261a4and fresh current-workspace thin-LTO archives (codegen-units=1):test_gap_fetch_request_from_node_incoming_message compile=0 run=0 test_gap_http_client_no_redirect_follow compile=0 run=0 test_gap_http_overloads_3226plus compile=0 run=0 test_gap_http_req_async_iterator compile=0 run=0 test_gap_http_res_socket_writable_onfinished compile=0 run=0 test_gap_net_connect_bound_value compile=0 run=0The first witness was also rerun through
run_parity_tests.shfrom a clean isolated release target and passed. The broader HTTP witness and the net-only witness each caused separate cold auto-optimize archive builds, so this was not one cached feature bundle being reused for every result. The incoming-message binary printed the expectedlen=55 match=true/expected len=55; none of the six panicked or aborted.Importantly, the release linker still logged retained bundled members and duplicate-symbol warnings, yet the binaries passed. That disproves the earlier inference that the duplicate count by itself demonstrates a split
async_bridge::RUNTIME; I am not applying the proposed all-Apple archive-localization change without a failing behavioral witness. The checked-in #5920 async-starvation regression remains the guard for the actual two-runtime failure.The report reproduced at
ebefba51a; there was no compile/link or HTTP/net source change between that commit and this passing build. The most likely explanation is stale auto-optimize state in the reporting environment, invalidated by subsequent runtime-source changes, rather than a still-present deterministic release linker defect.Reopening: this defect is live on current
main, independently reproduced.Verified just now on a build of
def2fbfdc(v0.5.1500), separate from the report that prompted this:$ ./test_gap_fetch_request_from_node_incoming_message thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13: there is no reactor running, must be called from the context of a Tokio 1.x runtime exit=134Why this needs reopening rather than a new issue
Both tracking issues are closed and the defect is not fixed. #7932 was filed today for the same abort and closed as a duplicate of this one; this one is also closed. So a live gap-suite failure has had no open owner — which is precisely the shape CLAUDE.md warns about: the gate is red, the red is untracked, and after a while a red gap suite stops looking unusual.
That has a measurable cost. Multiple agents this week burned time A/B-ing these crashes against pristine builds to establish they were not theirs, because nothing open said so.
Scope — six tests, and one discriminating detail
An independent run mapped the full set: six
test_gap_*programs abort with this signature. Five shareperry-ext-http/src/server/server.rs:911;net_connect_bound_valueaborts one frame lower, intokio-1.53.1/.../net/tcp/listener.rs:304.That one-frame difference is what distinguishes this defect from a generic "the http tests crash" and is worth keeping in the issue: any fix should be checked against the
listener.rs:304path too, not justserver.rs:911.Note for anyone triaging a red gap suite
These six are pre-existing on
mainand are not intest-parity/gap_snapshot.json, soscripts/run_gap_tests.shis red onmainbecause of them. Do not accept them into the snapshot — that would launder six tracked crashes through the expected-output channel. A/B against a pristineorigin/mainbuild before attributing them to a branch.Root cause found and fixed: two tokio compilations in one binary. PR #7999.
All six witnesses are one defect, and it is a build-graph defect rather than a runtime one.
net_connect_bound_value'slistener.rs:304frame does not need a separate fix — same
cause, same fix; the differing frame is only where each wrapper first touched the reactor
(perry-ext-http callstokio::spawnat its own call site, perry-ext-net calls
TcpListener::bind, whosePollEvented::newreachesHandle::current()one frame deeper
inside tokio). Both were reproduced here.Mechanism
perry-ext-*wrappers arestaticlibs, so each bundles its own copy of tokio;
libperry_stdlib.abundles one too, and perry-stdlib owns the process's only runtime.
tokio'sruntime::context::CONTEXTis athread_local!whose symbol carries the compiling
crate instance's metadata hash — two tokio compilations are two independent contexts.
perry-stdlib's runtime enters one, the wrapper reads the other, finds it empty, and under
panic = "abort"the process dies.This is a documented invariant that nothing checked.
optimized_libs/driver.rsstates the
failure verbatim (#507) and prevents it on the auto-optimize path by rebuilding every
tokio-using wrapper in the same cargo invocation as perry-stdlib-static.The compilation id is readable straight out of an archive's member names. On
55fd197d5:build libperry_stdlib.alibperry_ext_http.alibperry_ext_net.aresult auto-optimize tokio-692c8788…tokio-692c8788…— PASS 3/3 PERRY_NO_AUTO_OPTIMIZE=1tokio-5aeb6213…tokio-01c4c58f…tokio-59c9ffcf…exit 134, 3/3 one invocation, all packages tokio-5aeb6213…tokio-5aeb6213…tokio-5aeb6213…PASS 3/3 Three different tokios in the middle row — one per
cargo build -p <one crate>, because
cargo resolves feature unification per invocation.The paths that violated it:
no_auto.rs::build_missing_prebuilt_ext_lib(literally
cargo build --release -p perry-ext-http, reached wheneverPERRY_NO_AUTO_OPTIMIZE=1and
the archive is missing — this produced both witnesses);run_parity_tests.sh's node-suite
net step (a secondcargo build -p perry-ext-net -j1); and any hand-run
cargo build -p perry-ext-httpbefore aPERRY_SKIP_BUILD=1run, which is the fast path
agents use — andPERRY_SKIP_BUILD=1exportsPERRY_NO_AUTO_OPTIMIZE=1and then builds
nothing.Why CI stayed green: the 8
conformance-smokegap shards run onubuntu-latestand
build every archive in onecargo build, so the invariant held there by accident. The
failure is reachable only from a build configuration CI does not use.Fix
A link-time check (
crates/perry/src/commands/compile/shared_tokio.rs) that reads each
archive's tokio compilation id and refuses a mismatched pair, naming both ids and the
command that fixes it. Thearcontainer is parsed in-process rather than through
llvm-ar, so the check cannot silently stop gating when a tool is absent, and it reports
what it compared so "compared nothing" is distinguishable from "found no mismatch".
Sabotage-verified: rebuilding perry-ext-http alone diverges the id and the compile is
refused with no binary produced; restoring by rebuilding the full invocation makes it pass
again.Second, independent defect this uncovered
With coherent tokio, the no-auto gap path still failed — now at link, with five undefined
_js_ext_zlib_*. Theexternal-*-pumpfeatures are a property of the ONE prebuilt stdlib
while ext-archive selection is per-import, so no subset of pump features serves a mixed
corpus: the "build the ext packages too" compensationrun_parity_tests.shdocumented had
never actually worked. Fixed by routing the 23 of 554 ext-importing gap tests through
auto-optimize per-test (PERRY_FORCE_WELL_KNOWNalso works but costs 2.2 s → 37.7 s per
compile, measured — 17×).Gap suite
PERRY_SKIP_BUILD=1 ./run_parity_tests.sh --filter test_gap_, macOS arm64:Parity Pass: 538 Parity Fail: 15 Compile Fail: 1 Crashed: 0All six witnesses pass. Against the Linux snapshot: 14 of 15 reproduce,
test_gap_iterator_helpers_2874passes here, and two failures are not in it and are not
caused by the change —test_gap_specabi_reassign(a spec-ABI output mismatch) and
test_gap_zlib_4917_level(perry-ext-zlib has nodeflateRawSync/inflateRawSync; red
under the default runner onmaintoo). Filed separately.
test_gap_fetch_request_from_node_incoming_messagedies with SIGABRT (exit134) after printing 5 lines, while node exits 0 with
len=55 match=true.It is not a flake, and it is not in the allowlist
origin/main(38ff7eccc,--profile perry-dev, macOS arm64), via./run_parity_tests.sh --filter test_gap_fetch_request_from_node_incoming_message.test-parity/known_failures.json, so nothing is suppressing it —it is simply not being run by anything that can block a merge.
arms, so it is unrelated to that change; it is reported here rather than
passed over, because "fails on clean
maintoo" is not the same as "flake".Why it has been able to hide
parityis tag-gated (v0.5.1018), so a gap regression in this family is onlyjudged at release time, not per merge. That is the mechanism CLAUDE.md's hazard
list describes and the one #7580 already paid for in
test_gap_diagchannel_*.This crash may therefore be considerably older than its discovery date.
What the test does
#6432's regression cover:new Request(url, { body: req, duplex: "half" })where
reqis a Nodehttp.IncomingMessage, read back throughrequest.text(). It starts a realnode:httpserver on port 18994, POSTs toitself with
http.request, and compares the body length. Node prints:Perry prints five lines and then aborts. The abort is a signal death rather
than an output mismatch, so the harness classifies it CRASH, not FAIL.
Suggested next step
Run it under
PERRY_GC_ZEAL=1 PERRY_GC_PROTECT_FROMSPACE=1with a raisedPERRY_GC_PROTECT_FROMSPACE_DEPTH, and bisect. A deterministic abort in a paththat mixes a native small-handle body (
IncomingMessage) with the WebRequestreader is the shape CLAUDE.md flags as "a perfectly reproducible GC bug means a
table, not a register" — worth checking the
gc_register_mutable_root_scannerregistry for the http/fetch side tables before assuming a register.
Either way the first fix is procedural: if it is going to stay red, it needs a
known_failures.jsonentry with provenance soparity_known_failures.py --auditcan see it.