Repository navigation
fix(repo): run tests without Sparkplug, whose Node 24 GC bug segfaults them - #219
Merged
Merged
Conversation
nickruigrok
marked this pull request as ready for review
September 29, 2026 12:32
|
Contributor
Author
|
@greptileai review |
…s them Node 24's V8 pushes a stale register in Sparkplug's BaselineOutOfLinePrologue, which the GC reads as a pointer and segfaults on (nodejs/node#62393; the backport is open in nodejs/node#65753). It killed the Vitest process mid-run with exit 139, and can kill the Wrangler child that bundles connect for core's tests, which Wrangler's launcher reports as success.
…led by a signal The runtime flag stops only new baseline code, so say it makes the crash rare rather than impossible. Wrangler's launcher reports its CLI dying of a signal as success; say so instead of failing on a missing bundle.
nickruigrok
force-pushed
the
fix/node-sparkplug-segfault
branch
from
September 29, 2026 12:49
490fdcf to
9d0faac
Compare
nickruigrok
enabled auto-merge
September 29, 2026 13:05
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
vp testsometimes exited with code 139 (SIGSEGV) partway through a full local run, with no test failing, and passed on a rerun. The cause isn't our code, workerd or vitest-pool-workers. It's a V8 bug in Node 24: nodejs/node#62393.Sparkplug's
BaselineOutOfLineProloguepushes a register that Node's V8 build never fills (leaptiering on, sandbox off), so its stack slot holds junk such as0x7. If a major GC runs at that moment (reached through a stack-guard interrupt),ClearStaleLeftTrimmedPointerVisitorreads the junk as a heap pointer and faults at0x6/0xe. The fix is V80b94a9fd23ba, which is in Node 25+. The backport to 24.x, nodejs/node#65753, is open, not landed, and reporters on the issue still hit the crash on 24.21. So no Node release insidedevEngines(^24.11.0) fixes it yet, and there's nothing to pin or upgrade.The mitigation is to switch the baseline tier off (
--no-sparkplug) in the two Node processes that hit it:vite.config.ts).vp testspawnsnode vitest.mjsitself and gives us no place to pass Node flags, andNODE_OPTIONSrejects--no-sparkplug. So the root config callsv8.setFlagsFromString("--no-sparkplug"), which Vitest loads first. I checked that the runtime flag stops baseline compilation just like the CLI flag (%ActiveTierIsSparkplugstays false).apps/core/test/build-connect.ts). It hit the same crash in an amplified run. Wrangler's launcher turns its child dying from a signal into exit code 0 (code === null ? 0), so the only symptom was a misleadingENOENT … grasp-os-connect-*/index.jsfrom the global setup. The setup now runsnode --no-sparkplug wrangler.js, and the launcher passesexecArgvon to the CLI.Both comments say to drop this once a Node 24 release in
devEnginescarries the backport.Limits. A runtime flag stops only new baseline compilation. The few hundred functions Node compiles to baseline while starting up, before Vitest loads the config (module loader internals such as
createRequireandinitializeImportMeta), keep their Sparkplug code. So this makes the crash rare, not impossible: 0 in 15 mitigated runs against about 1 in 10 without it. The flag also covers onlyvp testfrom the repo root: a run with an app's own config (vp -C apps/core test) and Vitest'sforkspool children don't load the root config. None of those crashed in these runs. If an exit 139 still turns up before the backport ships, the next step is running Vitest ourselves asnode --no-sparkplug, not dropping this.Wrangler's launcher reports any signal death of its CLI as success, not just this segfault (an OOM kill does the same). So
bundleConnectnow says Wrangler wrote no bundle and points at a possible crash, instead of failing on a bare missing-file error. I checked this by killing the CLI with SIGSEGV from a preload.Evidence
Crash signature. macOS kept crash reports for every segfault:
~/Library/Logs/DiagnosticReports/node-*.ips(8 since 2026-09-27, 4 of them today). Every one has the same faulting frames:The crashed process is the
node …/vitest/vitest.mjs runchild ofvp test: its parent isnode, it runs 31rolldown-workerthreads, and it dies 12–40 s after launch. That window is when the Vitest process is busiest (peak RSS about 4.4 GB, transforms and module graph building). Node 24.15.0, macOS 26.6.2, arm64.Stress runs. Full
vp test, with 8yes > /dev/nullCPU burners, while other sessions ran their own suites on the same machine (load average above 20):main, unmitigateddist/test-assets)--always-sparkplug --no-maglev, the upstream repro's flags, set by aNODE_OPTIONS=--importpreload)ENOENTaboveagent.test.tsvi.waitForover its 1 s defaultA fifth crash report today (12:42) came from another worktree's unmitigated suite that was running at the same time.
Cost. None measurable. Mean of the passing runs: 271 s unmitigated vs 278 s mitigated under load, and 409 s vs 401 s amplified, both within the load noise. Most of the heavy work runs in native code (rolldown) and in workerd, not in the Vitest process's baseline tier.
What it isn't:
isolate: falseruntime reuse. That reuse happens inside workerd, which has its own V8, while the crash is in Node's V8.apps/connect/test/oauth.test.tsalone prints 26 for 26 passing tests, and connect doesn't useisolate: false. They're log noise inside workerd, not a symptom of this bug.How it was tested
The stress runs above, plus
vp checkon the changed files. An independent review then checked that the Wrangler path resolves under pnpm locally and in CI, and that loadingvite.config.tselsewhere (lint, fmt, build, dev) is harmless: it only turns off one of V8's JIT tiers. I also checked that a runtime--no-sparkplugbeats--always-sparkplugset before it, and the other way round, so the amplified runs with the fix really ran without Sparkplug.Checklist
vp checkandvp testpass