Skip to content

fix(repo): run tests without Sparkplug, whose Node 24 GC bug segfaults them - #219

Merged
nickruigrok merged 3 commits into
mainfrom
fix/node-sparkplug-segfault
Sep 29, 2026
Merged

nickruigrok merged 3 commits into
mainfrom
fix/node-sparkplug-segfault

Conversation

@nickruigrok

@nickruigrok nickruigrok commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What and why

vp test sometimes exited with code 139 (SIGSEGV) partway through a full local run, with no test failing, and passed on a rerun. The cause isn't our code, workerd or vitest-pool-workers. It's a V8 bug in Node 24: nodejs/node#62393.

Sparkplug's BaselineOutOfLinePrologue pushes a register that Node's V8 build never fills (leaptiering on, sandbox off), so its stack slot holds junk such as 0x7. If a major GC runs at that moment (reached through a stack-guard interrupt), ClearStaleLeftTrimmedPointerVisitor reads the junk as a heap pointer and faults at 0x6/0xe. The fix is V8 0b94a9fd23ba, which is in Node 25+. The backport to 24.x, nodejs/node#65753, is open, not landed, and reporters on the issue still hit the crash on 24.21. So no Node release inside devEngines (^24.11.0) fixes it yet, and there's nothing to pin or upgrade.

The mitigation is to switch the baseline tier off (--no-sparkplug) in the two Node processes that hit it:

  • The Vitest process (vite.config.ts). vp test spawns node vitest.mjs itself and gives us no place to pass Node flags, and NODE_OPTIONS rejects --no-sparkplug. So the root config calls v8.setFlagsFromString("--no-sparkplug"), which Vitest loads first. I checked that the runtime flag stops baseline compilation just like the CLI flag (%ActiveTierIsSparkplug stays false).
  • The Wrangler CLI that core's global setup runs to bundle connect (apps/core/test/build-connect.ts). It hit the same crash in an amplified run. Wrangler's launcher turns its child dying from a signal into exit code 0 (code === null ? 0), so the only symptom was a misleading ENOENT … grasp-os-connect-*/index.js from the global setup. The setup now runs node --no-sparkplug wrangler.js, and the launcher passes execArgv on to the CLI.

Both comments say to drop this once a Node 24 release in devEngines carries the backport.

Limits. A runtime flag stops only new baseline compilation. The few hundred functions Node compiles to baseline while starting up, before Vitest loads the config (module loader internals such as createRequire and initializeImportMeta), keep their Sparkplug code. So this makes the crash rare, not impossible: 0 in 15 mitigated runs against about 1 in 10 without it. The flag also covers only vp test from the repo root: a run with an app's own config (vp -C apps/core test) and Vitest's forks pool children don't load the root config. None of those crashed in these runs. If an exit 139 still turns up before the backport ships, the next step is running Vitest ourselves as node --no-sparkplug, not dropping this.

Wrangler's launcher reports any signal death of its CLI as success, not just this segfault (an OOM kill does the same). So bundleConnect now says Wrangler wrote no bundle and points at a possible crash, instead of failing on a bare missing-file error. I checked this by killing the CLI with SIGSEGV from a preload.

Evidence

Crash signature. macOS kept crash reports for every segfault: ~/Library/Logs/DiagnosticReports/node-*.ips (8 since 2026-09-27, 4 of them today). Every one has the same faulting frames:

EXC_BAD_ACCESS (SIGSEGV)  KERN_INVALID_ADDRESS at 0x0000000000000006 / 0x000000000000000e
v8::internal::ClearStaleLeftTrimmedPointerVisitor::VisitRootPointers
v8::internal::InternalFrame::Iterate
v8::internal::Isolate::Iterate
v8::internal::Heap::IterateRoots
v8::internal::MarkCompactCollector::MarkRoots / MarkLiveObjects / CollectGarbage
v8::internal::Heap::CollectGarbage
v8::internal::StackGuard::HandleInterrupts
v8::internal::Runtime_StackGuardWithGap
Builtins_CEntry_Return1_ArgvOnStack_NoBuiltinExit
Builtins_BaselineOutOfLinePrologue

The crashed process is the node …/vitest/vitest.mjs run child of vp test: its parent is node, it runs 31 rolldown-worker threads, and it dies 12–40 s after launch. That window is when the Vitest process is busiest (peak RSS about 4.4 GB, transforms and module graph building). Node 24.15.0, macOS 26.6.2, arm64.

Stress runs. Full vp test, with 8 yes > /dev/null CPU burners, while other sessions ran their own suites on the same machine (load average above 20):

Setup Runs Exit 139 Other failures
main, unmitigated 10 1 (after 19 s, no test failed) 1 console test over its 5 s timeout; 1 run invalid (I ran two suites in one checkout, which race on dist/test-assets)
unmitigated, amplified (--always-sparkplug --no-maglev, the upstream repro's flags, set by a NODE_OPTIONS=--import preload) 5 0 in Vitest; 1 in the Wrangler bundling child, reported only as the ENOENT above none
this branch 10 0 1 agent.test.ts vi.waitFor over its 1 s default
this branch, amplified 5 0 none

A fifth crash report today (12:42) came from another worktree's unmitigated suite that was running at the same time.

Cost. None measurable. Mean of the passing runs: 271 s unmitigated vs 278 s mitigated under load, and 409 s vs 401 s amplified, both within the load noise. Most of the heavy work runs in native code (rolldown) and in workerd, not in the Vitest process's baseline tier.

What it isn't:

  • Memory growth. The Vitest process's RSS peaks at about 4.4 GB early, then falls to 0.5–2 GB, with no upward trend across a run. The workerd runtimes stay flat (about 19 × 200 MB).
  • isolate: false runtime reuse. That reuse happens inside workerd, which has its own V8, while the crash is in Node's V8.
  • The "code had hung … canceled this request" messages (about 550–800 per run, in every run, crashing or not). workerd logs one whenever a test's RPC call into connect rejects as the test expects. apps/connect/test/oauth.test.ts alone prints 26 for 26 passing tests, and connect doesn't use isolate: false. They're log noise inside workerd, not a symptom of this bug.

How it was tested

The stress runs above, plus vp check on the changed files. An independent review then checked that the Wrangler path resolves under pnpm locally and in CI, and that loading vite.config.ts elsewhere (lint, fmt, build, dev) is harmless: it only turns off one of V8's JIT tiers. I also checked that a runtime --no-sparkplug beats --always-sparkplug set before it, and the other way round, so the amplified runs with the fix really ran without Sparkplug.

Checklist

  • vp check and vp test pass
  • Schema changes are expand-only (no drops or renames in the same release)
  • New external calls go through connect; new model calls through the model gateway
  • No secrets, client names or client configuration
  • Risky or user-visible changes ship behind a feature flag (test tooling only)

@nickruigrok
nickruigrok marked this pull request as ready for review September 29, 2026 12:32
@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adjusts test runtime to work around a Node.js bug.

The PR appears safe to merge; the previously reported misleading error has been corrected.

Findings

  1. P2 Missing bundle misdiagnosed ▶
Fix with agent prompt
### Issue 1
apps/core/test/build-connect.ts:undefined-42
If Wrangler exits successfully without producing `index.js` for a reason other than a signal, this error still says its CLI was killed by a signal. The check establishes only that the file is missing, so the message could send someone debugging the failed test setup toward the wrong cause.

```suggestion
        "Wrangler exited without writing connect's bundle."
```

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

Mitigates a Node 24 Sparkplug-related test-process crash by disabling Sparkplug when the root Vite config loads and when core’s test setup invokes Wrangler. It also gives a more useful error if Wrangler exits without writing the connect bundle.

Reviews (2) · Last reviewed commit: "fix(repo): report a missing connect bund..."

Comment thread apps/core/test/build-connect.ts Outdated
@nickruigrok

Copy link
Copy Markdown
Contributor Author

@greptileai review

…s them

Node 24's V8 pushes a stale register in Sparkplug's BaselineOutOfLinePrologue,
which the GC reads as a pointer and segfaults on (nodejs/node#62393; the
backport is open in nodejs/node#65753). It killed the Vitest process mid-run
with exit 139, and can kill the Wrangler child that bundles connect for core's
tests, which Wrangler's launcher reports as success.
…led by a signal

The runtime flag stops only new baseline code, so say it makes the crash
rare rather than impossible. Wrangler's launcher reports its CLI dying of a
signal as success; say so instead of failing on a missing bundle.
@nickruigrok
nickruigrok force-pushed the fix/node-sparkplug-segfault branch from 490fdcf to 9d0faac Compare September 29, 2026 12:49
@nickruigrok
nickruigrok merged commit c88ed96 into main Sep 29, 2026
8 checks passed
@nickruigrok
nickruigrok deleted the fix/node-sparkplug-segfault branch September 29, 2026 13:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant