Skip to content

test_gap_fetch_request_from_node_incoming_message SIGABRTs deterministically on pristine main, and is in no allowlist #7629

Description

@proggeramlug

test_gap_fetch_request_from_node_incoming_message dies with SIGABRT (exit
134)
after printing 5 lines, while node exits 0 with len=55 match=true.

It is not a flake, and it is not in the allowlist

  • 3/3 reproducible on a build of pristine origin/main (38ff7eccc,
    --profile perry-dev, macOS arm64), via
    ./run_parity_tests.sh --filter test_gap_fetch_request_from_node_incoming_message.
  • Absent from test-parity/known_failures.json, so nothing is suppressing it —
    it is simply not being run by anything that can block a merge.
  • Found while A/B-ing refactor(codegen): migrate instance_misc1 + logical_collections + map_set onto the Layer 1 rooting API (#7615) #7627 (Layer 1 slice 2). It reproduces identically on both
    arms, so it is unrelated to that change; it is reported here rather than
    passed over, because "fails on clean main too" is not the same as "flake".

Why it has been able to hide

parity is tag-gated (v0.5.1018), so a gap regression in this family is only
judged at release time, not per merge. That is the mechanism CLAUDE.md's hazard
list describes and the one #7580 already paid for in test_gap_diagchannel_*.
This crash may therefore be considerably older than its discovery date.

What the test does

#6432's regression cover: new Request(url, { body: req, duplex: "half" })
where req is a Node http.IncomingMessage, read back through
request.text(). It starts a real node:http server on port 18994, POSTs to
itself with http.request, and compares the body length. Node prints:

len=55 match=true
expected len=55

Perry prints five lines and then aborts. The abort is a signal death rather
than an output mismatch, so the harness classifies it CRASH, not FAIL.

Suggested next step

Run it under PERRY_GC_ZEAL=1 PERRY_GC_PROTECT_FROMSPACE=1 with a raised
PERRY_GC_PROTECT_FROMSPACE_DEPTH, and bisect. A deterministic abort in a path
that mixes a native small-handle body (IncomingMessage) with the Web Request
reader is the shape CLAUDE.md flags as "a perfectly reproducible GC bug means a
table, not a register" — worth checking the gc_register_mutable_root_scanner
registry for the http/fetch side tables before assuming a register.

Either way the first fix is procedural: if it is going to stay red, it needs a
known_failures.json entry with provenance so parity_known_failures.py --audit
can see it.

Activity

  1. proggeramlug commented on Aug 10, 2026

    @proggeramlug
    ContributorAuthor

    Reproduced, and the cause is not GC — so the suggested PERRY_GC_ZEAL / from-space bisect would have been a dead end. Posting the diagnosis rather than a fix; the remaining work is architectural and overlaps #7656.

    What the abort actually is

    The harness classifies it CRASH and discards the message. It is a Rust panic:

    thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13:
    there is no reactor running, must be called from the context of a Tokio 1.x runtime
    

    Line 911 is the tokio::spawn(async move { … }) inside the perry_ffi::spawn_blocking_with_reactor closure that adopts the listener — i.e. the closure ran on a thread with no tokio reactor, which is precisely the failure perry_ffi_spawn_blocking_with_reactor exists to prevent (its doc names this exact panic and the v0.5.571 net regression batch it closed).

    It is link-configuration dependent, and here is the A/B

    Same compiler binary, same test, same host. The only variable is whether libperry_ext_*.a are present in PERRY_RUNTIME_DIR:

    libperry_ext_*.a in the runtime dir result
    present PASS 3/3
    removed SIGABRT 2/2

    That is deterministic in both directions, which is why it reproduces 3/3 on a stock build: run_parity_tests.sh builds -p perry -p perry-runtime -p perry-stdlib -p perry-runtime-static -p perry-stdlib-static and adds the five perry-ext-* crates only under PERRY_NO_AUTO_OPTIMIZE (line ~634). So the default run has no prebuilt ext archives, auto-optimize builds its own, and that is the arm that aborts.

    Why that produces a missing reactor

    perry-stdlib's perry_ffi_spawn_blocking_with_reactor is correct and not feature-gated — it spawns on async_bridge::runtime(), the multi-thread runtime that has the reactor. The #[cfg(test)] shim in perry-ext-http/src/test_async_shims.rs (which does invoke(ctx) synchronously, and would produce exactly this symptom) is not in a release archive, so it is not the culprit either.

    What is left is the hazard already recorded for this family: an auto-built ext staticlib bundles its own copy of the runtime/stdlib, and that copy wins the link. The ext code then spawns onto a second async_bridge::runtime() instance that nothing pumps — same statics, different instance — so the closure runs with no reactor in context. Consistent with both directions of the A/B: one prebuilt archive ⇒ one copy ⇒ pass.

    I did not confirm the duplicate-symbol layout in the emitted binary; that is the next concrete step (nm the auto-built archive for a second async_bridge runtime static), and it is the same root as #7656 — nothing per-PR links the perry-ext-* crates, so this configuration is only ever exercised at tag time.

    On the procedural half

    I have not added a known_failures.json entry. The suggestion was conditioned on "if it is going to stay red", and this is a real, located bug in a configuration CI actually ships — suppressing it now would hide a genuine failure behind provenance. Happy to add the entry if you would rather have it tracked-and-skipped while the link issue is scheduled.

    Incidentally

    The documented ad-hoc recipe (-p perry -p perry-runtime-static -p perry-stdlib-static) cannot compile this test in either mode: with auto-optimize it fails to link js_http_client_request_end_full / js_http_client_request_write_full / js_http_incoming_message_set_encoding, and with PERRY_NO_AUTO_OPTIMIZE=1 it fails on the zlib pump symbols unless the stdlib feature set matches the ext archives exactly. Worth a line in CLAUDE.md's build section, since "reproduce it ad hoc" is the first thing anyone will try.

  2. proggeramlug commented on Aug 10, 2026

    @proggeramlug
    ContributorAuthor

    Root-caused on main @ db44b31b7. It is not a codegen bug and it is no longer a perry-dev failure — it is a release-profile link defect, and the mechanism is fully determined.

    First, one correction to the report

    The issue says 3/3 reproducible under --profile perry-dev. That is no longer true:

    build harness direct run
    perry-dev PASS 2/2 len=55 match=true, exit 0 — byte-identical to node
    --release (thin LTO, codegen-units=1) CRASH 2/2, SIGABRT exit 134

    So something between 38ff7eccc and now fixed the perry-dev arm, and the release arm still dies. That matters because run_parity_tests.sh builds --release, so the parity gate is the only thing that can see this — and parity is tag-gated, which is the hiding mechanism this issue already identifies.

    What actually aborts

    thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13:
    there is no reactor running, must be called from the context of a Tokio 1.x runtime
    

    A Rust panic under panic=abort — hence SIGABRT rather than an output mismatch, hence CRASH rather than FAIL. The message is TcpListener::from_std's: it is running in a tokio context whose I/O reactor is not the one being driven.

    Why, exactly

    libperry_ext_http.a is a staticlib, so it physically bundles perry-runtime/perry-stdlib and their shared deps. The link path already knows this and strips the bundled members — strip_bundled_runtime_from_well_known_lib → strip_bundled_shared_deps_from_well_known_lib → strip_duplicate_objects_from_well_known_lib — and build_and_run.rs states the exact stake:

    a wrapper's bundled perry-runtime copy would win first-definition over stdlib's and split the runtime's mutable globals in two … spawned async tasks then starve because the event pump's wait-driver slot is registered in one copy and read from the other.

    That is precisely what happens here. The strip runs in both profiles and reaches opposite conclusions:

    # perry-dev
    [strip-dedup] libperry_ext_http.a: dropped 1676 bundled member(s) also provided by
                  stdlib (shared transitive deps, fixed-point safe)
    [strip-dedup] _libperry_ext_http.a_nosharedeps.lib: localized wrapper-only globals in 1 member(s)
    
    # --release
    [strip-dedup] libperry_ext_http.a: keeping bundled alloc-…rcgu.o — needed by a kept
                  sibling and not provided by stdlib
    [strip-dedup] libperry_ext_http.a: keeping bundled bytes-…rcgu.o — …
    [strip-dedup] libperry_ext_http.a: keeping bundled core-…rcgu.o — …
                  (no "dropped N member(s)" line at all)
    

    The consequence, measured on the same source and the same test:

    profile duplicate symbols at link (libperry_ext_http.a vs libperry_stdlib.a) result
    perry-dev 11 passes
    --release 579 SIGABRT

    The strip's criterion is "is this symbol also provided by stdlib?" Under thin LTO with codegen-units=1 the stdlib archive internalizes those symbols, so the criterion answers no, the ext copies are kept, and both copies load. perry-ext-http then binds to a second async_bridge::RUNTIME — a current-thread runtime (see the "unified single-thread async model" note in common/async_bridge.rs) that only the main JS event loop drives, and it drives the other copy. Nothing ever drives the ext copy's reactor, so the first TcpListener::from_std on the accept path panics.

    Two things this rules out, both checked rather than assumed:

    Fix direction

    The strip needs a criterion that survives LTO internalization — "provided by stdlib" has to mean "provided by the stdlib that will be linked", which under thin LTO is not the archive's exported symbol table. The precedent is already in-tree: dedup_ui_lib_against_linked_libs solves the same problem for the UI lib by localizing every bundled global the linked stdlib/runtime also define, instead of trying to drop members — "sibling references rebind to the single linked copy (one set of runtime/std mutable state — the #5920 invariant)". That is applied only to the UI lib today; the ext archives take the drop-members path and are exactly where it fails.

    I have not landed that change: it alters the link path for every perry-ext-* archive on Apple targets and needs validation well beyond this one test.

    Worth noting this is the concrete cost of #7656 — no per-PR gate links the perry-ext-* crates, so a release-only link defect in the ext path is invisible until a tag.

  3. self-assigned this
    on Aug 12, 2026
  4. proggeramlug commented on Aug 12, 2026

    @proggeramlug
    ContributorAuthor

    Verified fixed on current main after rebuilding the exact release-profile auto-optimize configurations, so I am closing this rather than landing speculative linker changes.

    I checked all six witnesses from the duplicate #7932 individually with the release compiler/linker at f110261a4 and fresh current-workspace thin-LTO archives (codegen-units=1):

    test_gap_fetch_request_from_node_incoming_message  compile=0 run=0
    test_gap_http_client_no_redirect_follow            compile=0 run=0
    test_gap_http_overloads_3226plus                   compile=0 run=0
    test_gap_http_req_async_iterator                   compile=0 run=0
    test_gap_http_res_socket_writable_onfinished       compile=0 run=0
    test_gap_net_connect_bound_value                   compile=0 run=0
    

    The first witness was also rerun through run_parity_tests.sh from a clean isolated release target and passed. The broader HTTP witness and the net-only witness each caused separate cold auto-optimize archive builds, so this was not one cached feature bundle being reused for every result. The incoming-message binary printed the expected len=55 match=true / expected len=55; none of the six panicked or aborted.

    Importantly, the release linker still logged retained bundled members and duplicate-symbol warnings, yet the binaries passed. That disproves the earlier inference that the duplicate count by itself demonstrates a split async_bridge::RUNTIME; I am not applying the proposed all-Apple archive-localization change without a failing behavioral witness. The checked-in #5920 async-starvation regression remains the guard for the actual two-runtime failure.

    The report reproduced at ebefba51a; there was no compile/link or HTTP/net source change between that commit and this passing build. The most likely explanation is stale auto-optimize state in the reporting environment, invalidated by subsequent runtime-source changes, rather than a still-present deterministic release linker defect.

  5. proggeramlug commented on Aug 12, 2026

    @proggeramlug
    ContributorAuthor

    Reopening: this defect is live on current main, independently reproduced.

    Verified just now on a build of def2fbfdc (v0.5.1500), separate from the report that prompted this:

    $ ./test_gap_fetch_request_from_node_incoming_message
    thread '<unnamed>' panicked at crates/perry-ext-http/src/server/server.rs:911:13:
    there is no reactor running, must be called from the context of a Tokio 1.x runtime
    exit=134
    

    Why this needs reopening rather than a new issue

    Both tracking issues are closed and the defect is not fixed. #7932 was filed today for the same abort and closed as a duplicate of this one; this one is also closed. So a live gap-suite failure has had no open owner — which is precisely the shape CLAUDE.md warns about: the gate is red, the red is untracked, and after a while a red gap suite stops looking unusual.

    That has a measurable cost. Multiple agents this week burned time A/B-ing these crashes against pristine builds to establish they were not theirs, because nothing open said so.

    Scope — six tests, and one discriminating detail

    An independent run mapped the full set: six test_gap_* programs abort with this signature. Five share perry-ext-http/src/server/server.rs:911; net_connect_bound_value aborts one frame lower, in tokio-1.53.1/.../net/tcp/listener.rs:304.

    That one-frame difference is what distinguishes this defect from a generic "the http tests crash" and is worth keeping in the issue: any fix should be checked against the listener.rs:304 path too, not just server.rs:911.

    Note for anyone triaging a red gap suite

    These six are pre-existing on main and are not in test-parity/gap_snapshot.json, so scripts/run_gap_tests.sh is red on main because of them. Do not accept them into the snapshot — that would launder six tracked crashes through the expected-output channel. A/B against a pristine origin/main build before attributing them to a branch.

  6. proggeramlug commented on Aug 13, 2026

    @proggeramlug
    ContributorAuthor

    Root cause found and fixed: two tokio compilations in one binary. PR #7999.

    All six witnesses are one defect, and it is a build-graph defect rather than a runtime one.
    net_connect_bound_value's listener.rs:304 frame does not need a separate fix — same
    cause, same fix; the differing frame is only where each wrapper first touched the reactor
    (perry-ext-http calls tokio::spawn at its own call site, perry-ext-net calls
    TcpListener::bind, whose PollEvented::new reaches Handle::current() one frame deeper
    inside tokio). Both were reproduced here.

    Mechanism

    perry-ext-* wrappers are staticlibs, so each bundles its own copy of tokio;
    libperry_stdlib.a bundles one too, and perry-stdlib owns the process's only runtime.
    tokio's runtime::context::CONTEXT is a thread_local! whose symbol carries the compiling
    crate instance's metadata hash — two tokio compilations are two independent contexts.
    perry-stdlib's runtime enters one, the wrapper reads the other, finds it empty, and under
    panic = "abort" the process dies.

    This is a documented invariant that nothing checked. optimized_libs/driver.rs states the
    failure verbatim (#507) and prevents it on the auto-optimize path by rebuilding every
    tokio-using wrapper in the same cargo invocation as perry-stdlib-static.

    The compilation id is readable straight out of an archive's member names. On 55fd197d5:

    build libperry_stdlib.a libperry_ext_http.a libperry_ext_net.a result
    auto-optimize tokio-692c8788… tokio-692c8788… — PASS 3/3
    PERRY_NO_AUTO_OPTIMIZE=1 tokio-5aeb6213… tokio-01c4c58f… tokio-59c9ffcf… exit 134, 3/3
    one invocation, all packages tokio-5aeb6213… tokio-5aeb6213… tokio-5aeb6213… PASS 3/3

    Three different tokios in the middle row — one per cargo build -p <one crate>, because
    cargo resolves feature unification per invocation.

    The paths that violated it: no_auto.rs::build_missing_prebuilt_ext_lib (literally
    cargo build --release -p perry-ext-http, reached whenever PERRY_NO_AUTO_OPTIMIZE=1 and
    the archive is missing — this produced both witnesses); run_parity_tests.sh's node-suite
    net step (a second cargo build -p perry-ext-net -j1); and any hand-run
    cargo build -p perry-ext-http before a PERRY_SKIP_BUILD=1 run, which is the fast path
    agents use — and PERRY_SKIP_BUILD=1 exports PERRY_NO_AUTO_OPTIMIZE=1 and then builds
    nothing.

    Why CI stayed green: the 8 conformance-smoke gap shards run on ubuntu-latest and
    build every archive in one cargo build, so the invariant held there by accident. The
    failure is reachable only from a build configuration CI does not use.

    Fix

    A link-time check (crates/perry/src/commands/compile/shared_tokio.rs) that reads each
    archive's tokio compilation id and refuses a mismatched pair, naming both ids and the
    command that fixes it. The ar container is parsed in-process rather than through
    llvm-ar, so the check cannot silently stop gating when a tool is absent, and it reports
    what it compared so "compared nothing" is distinguishable from "found no mismatch".
    Sabotage-verified: rebuilding perry-ext-http alone diverges the id and the compile is
    refused with no binary produced; restoring by rebuilding the full invocation makes it pass
    again.

    Second, independent defect this uncovered

    With coherent tokio, the no-auto gap path still failed — now at link, with five undefined
    _js_ext_zlib_*. The external-*-pump features are a property of the ONE prebuilt stdlib
    while ext-archive selection is per-import, so no subset of pump features serves a mixed
    corpus
    : the "build the ext packages too" compensation run_parity_tests.sh documented had
    never actually worked. Fixed by routing the 23 of 554 ext-importing gap tests through
    auto-optimize per-test (PERRY_FORCE_WELL_KNOWN also works but costs 2.2 s → 37.7 s per
    compile, measured — 17×).

    Gap suite

    PERRY_SKIP_BUILD=1 ./run_parity_tests.sh --filter test_gap_, macOS arm64:

    Parity Pass:  538      Parity Fail: 15      Compile Fail: 1      Crashed: 0
    

    All six witnesses pass. Against the Linux snapshot: 14 of 15 reproduce,
    test_gap_iterator_helpers_2874 passes here, and two failures are not in it and are not
    caused by the change — test_gap_specabi_reassign (a spec-ABI output mismatch) and
    test_gap_zlib_4917_level (perry-ext-zlib has no deflateRawSync/inflateRawSync; red
    under the default runner on main too). Filed separately.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions