Repository navigation
[v24.x] deps: V8: backport 0b94a9fd23ba - #65753
Closed
ruangustavo wants to merge 1 commit into
Closed
ruangustavo wants to merge 1 commit into
ruangustavo wants to merge 1 commit into
Conversation
Collaborator
|
Review requested:
|
richardlau
approved these changes
Sep 3, 2026
Collaborator
Collaborator
Collaborator
Collaborator
Author
|
Hey @richardlau, let me know how can I help you fixing the pipeline :) |
This comment was marked as outdated.
This comment was marked as outdated.
Collaborator
Collaborator
aduh95
force-pushed
the
v24.x-staging
branch
3 times, most recently
from
September 9, 2026 21:47
7fc125b to
e711d3f
Compare
Original commit message:
[leaptiering] Fix BaselineOutOfLinePrologue builtin
... which tried to preserve kJavaScriptCallDispatchHandleRegister even
on configurations where it's not used which resulted in a random value
on the stack discoverable by GC.
This issue triggered only on non-sandbox configuration with enabled
leaptiering.
Drive-by: fix MacroAssembler::GenerateTailCallToReturnedCode() on riscv
port which wasn't preserving dispatch handle as all the other ports do.
Bug: 42204201
Fixed: 413769394
Change-Id: If146b0b7a6cf972ed5a881142f40980774f19cba
Reviewed-on: https://chromium-review.googlesource.com/c/v8/v8/+/6587010
Commit-Queue: Igor Sheludko <ishell@chromium.org>
Reviewed-by: Olivier Flückiger <olivf@chromium.org>
Cr-Commit-Position: refs/heads/main@{#100512}
Node.js 24 builds V8 with leaptiering enabled and the sandbox disabled,
so V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE is not defined and the JS
calling convention does not carry the dispatch handle register (x4 on
arm64). BaselineOutOfLinePrologue and GenerateTailCallToReturnedCode
still pushed that register as a tagged slot of an INTERNAL frame, so
whatever value the caller left there is dereferenced by
ClearStaleLeftTrimmedPointerVisitor during mark-compact root scanning
and crashes the process with SIGSEGV (seen as jest workers dying).
Refs: v8/v8@0b94a9f
Fixes: nodejs#62393
ruangustavo
force-pushed
the
backport-v8-0b94a9fd23ba
branch
from
September 10, 2026 03:47
a315622 to
f47b5f8
Compare
Author
|
Hey @richardlau, I rebased onto the current |
This comment was marked as spam.
This comment was marked as spam.
bbertucc
added a commit
to EqualifyEverything/equalify-iris
that referenced
this pull request
Sep 14, 2026
…segfault (#477) #405's crash has an upstream name — nodejs/node#62393, V8 CL 0b94a9fd23ba: the BaselineOutOfLinePrologue builtin left a random value on the stack where the GC found it. `ClearStaleLeftTrimmedPointerVisitor` reads it as a heap pointer and faults at 0xe, the address in all five local reports. Sparkplug generates that prologue, so `--no-sparkplug` removes the path. The two decisions #405 was holding both answer no. Node cannot be bumped into the fix: the backport (nodejs/node#65753) is open, not landed, so no released 24.x has it, and one upstream report has the crash live on v26.7.0. CI needs no dead-child retry: every occurrence here and upstream is macOS arm64, and every workflow runs ubuntu-latest. Cost is nothing measurable — 55.8 s mean either way over two runs each locally, ~3% on Linux x64 over the two heaviest jsdom files, against a 22-minute review step. The flag is on the test script only, on purpose: `npm start` and `npm run dev` keep the Sparkplug path, because one dev server dying is loud where a dead test child reads as a clean run with a short pass count. A test pins the flag, and after round 1 it pins its *position*: `node --test "glob" --no-sparkplug` exits 0, warns nothing, and the child's `execArgv` does not carry the flag, so a presence-only check could go green through the exact regression it exists to catch. Drop the flag and that test together when nodejs/node#65753 ships in a 24.x release. Two rounds, both approved; round 1's three notes fixed, round 2 clean. Refs #405 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
richardlau
approved these changes
Sep 16, 2026
Contributor
Failed to start CI�[36m⠋�[39m Getting reviews from nodejs/node/pull/65753 �[36m⠋�[39m Getting commits from nodejs/node/pull/65753 �[36m⠙�[39m Validating Jenkins credentials �[36m⠙�[39m Validating Jenkins credentials ✔ Jenkins credentials valid �[36m⠹�[39m Getting comments from nodejs/node/pull/65753 �[36m⠸�[39m Querying data for job/node-test-pull-request/77190/ �[36m⠸�[39m Querying data for job/node-test-pull-request/77190/ �[36m⠸�[39m Querying API for job/node-test-pull-request/77190/ SyntaxError: Unexpected token '<', ..." https://github.com/nodejs/node/actions/runs/35160272536 |
Collaborator
Collaborator
Collaborator
Collaborator
aduh95
force-pushed
the
v24.x-staging
branch
from
September 19, 2026 14:06
00d2960 to
4396c9b
Compare
Author
|
Hey @richardlau, this is the second time the pipeline has failed at the same stage. lmk how I can fix it, I can’t see the error details. |
Collaborator
Merged
5 tasks done
richardlau
approved these changes
Sep 29, 2026
richardlau
pushed a commit
that referenced
this pull request
Sep 29, 2026
Original commit message:
[leaptiering] Fix BaselineOutOfLinePrologue builtin
... which tried to preserve kJavaScriptCallDispatchHandleRegister even
on configurations where it's not used which resulted in a random value
on the stack discoverable by GC.
This issue triggered only on non-sandbox configuration with enabled
leaptiering.
Drive-by: fix MacroAssembler::GenerateTailCallToReturnedCode() on riscv
port which wasn't preserving dispatch handle as all the other ports do.
Bug: 42204201
Fixed: 413769394
Change-Id: If146b0b7a6cf972ed5a881142f40980774f19cba
Reviewed-on: https://chromium-review.googlesource.com/c/v8/v8/+/6587010
Commit-Queue: Igor Sheludko <ishell@chromium.org>
Reviewed-by: Olivier Flückiger <olivf@chromium.org>
Cr-Commit-Position: refs/heads/main@{#100512}
Node.js 24 builds V8 with leaptiering enabled and the sandbox disabled,
so V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE is not defined and the JS
calling convention does not carry the dispatch handle register (x4 on
arm64). BaselineOutOfLinePrologue and GenerateTailCallToReturnedCode
still pushed that register as a tagged slot of an INTERNAL frame, so
whatever value the caller left there is dereferenced by
ClearStaleLeftTrimmedPointerVisitor during mark-compact root scanning
and crashes the process with SIGSEGV (seen as jest workers dying).
Refs: v8/v8@0b94a9f
Fixes: #62393
PR-URL: #65753
Reviewed-By: Richard Lau <richard.lau@ibm.com>
Member
|
Landed in fb0d0ca. |
nickruigrok
added a commit
to grasplabs/work
that referenced
this pull request
Sep 29, 2026
…s them Node 24's V8 pushes a stale register in Sparkplug's BaselineOutOfLinePrologue, which the GC reads as a pointer and segfaults on (nodejs/node#62393; the backport is open in nodejs/node#65753). It killed the Vitest process mid-run with exit 139, and can kill the Wrangler child that bundles connect for core's tests, which Wrangler's launcher reports as success.
nickruigrok
added a commit
to grasplabs/work
that referenced
this pull request
Sep 29, 2026
…s them (#219) ## What and why `vp test` sometimes exited with code 139 (SIGSEGV) partway through a full local run, with no test failing, and passed on a rerun. The cause isn't our code, workerd or vitest-pool-workers. It's a V8 bug in Node 24: **nodejs/node#62393**. Sparkplug's `BaselineOutOfLinePrologue` pushes a register that Node's V8 build never fills (leaptiering on, sandbox off), so its stack slot holds junk such as `0x7`. If a major GC runs at that moment (reached through a stack-guard interrupt), `ClearStaleLeftTrimmedPointerVisitor` reads the junk as a heap pointer and faults at `0x6`/`0xe`. The fix is V8 `0b94a9fd23ba`, which is in Node 25+. The backport to 24.x, nodejs/node#65753, is **open, not landed**, and reporters on the issue still hit the crash on 24.21. So no Node release inside `devEngines` (`^24.11.0`) fixes it yet, and there's nothing to pin or upgrade. The mitigation is to switch the baseline tier off (`--no-sparkplug`) in the two Node processes that hit it: - **The Vitest process** (`vite.config.ts`). `vp test` spawns `node vitest.mjs` itself and gives us no place to pass Node flags, and `NODE_OPTIONS` rejects `--no-sparkplug`. So the root config calls `v8.setFlagsFromString("--no-sparkplug")`, which Vitest loads first. I checked that the runtime flag stops baseline compilation just like the CLI flag (`%ActiveTierIsSparkplug` stays false). - **The Wrangler CLI that core's global setup runs to bundle connect** (`apps/core/test/build-connect.ts`). It hit the same crash in an amplified run. Wrangler's launcher turns its child dying from a signal into exit code 0 (`code === null ? 0`), so the only symptom was a misleading `ENOENT … grasp-os-connect-*/index.js` from the global setup. The setup now runs `node --no-sparkplug wrangler.js`, and the launcher passes `execArgv` on to the CLI. Both comments say to drop this once a Node 24 release in `devEngines` carries the backport. **Limits.** A runtime flag stops only *new* baseline compilation. The few hundred functions Node compiles to baseline while starting up, before Vitest loads the config (module loader internals such as `createRequire` and `initializeImportMeta`), keep their Sparkplug code. So this makes the crash rare, not impossible: 0 in 15 mitigated runs against about 1 in 10 without it. The flag also covers only `vp test` from the repo root: a run with an app's own config (`vp -C apps/core test`) and Vitest's `forks` pool children don't load the root config. None of those crashed in these runs. If an exit 139 still turns up before the backport ships, the next step is running Vitest ourselves as `node --no-sparkplug`, not dropping this. Wrangler's launcher reports any signal death of its CLI as success, not just this segfault (an OOM kill does the same). So `bundleConnect` now says Wrangler wrote no bundle and points at a possible crash, instead of failing on a bare missing-file error. I checked this by killing the CLI with SIGSEGV from a preload. ## Evidence **Crash signature.** macOS kept crash reports for every segfault: `~/Library/Logs/DiagnosticReports/node-*.ips` (8 since 2026-09-27, 4 of them today). Every one has the same faulting frames: ```text EXC_BAD_ACCESS (SIGSEGV) KERN_INVALID_ADDRESS at 0x0000000000000006 / 0x000000000000000e v8::internal::ClearStaleLeftTrimmedPointerVisitor::VisitRootPointers v8::internal::InternalFrame::Iterate v8::internal::Isolate::Iterate v8::internal::Heap::IterateRoots v8::internal::MarkCompactCollector::MarkRoots / MarkLiveObjects / CollectGarbage v8::internal::Heap::CollectGarbage v8::internal::StackGuard::HandleInterrupts v8::internal::Runtime_StackGuardWithGap Builtins_CEntry_Return1_ArgvOnStack_NoBuiltinExit Builtins_BaselineOutOfLinePrologue ``` The crashed process is the `node …/vitest/vitest.mjs run` child of `vp test`: its parent is `node`, it runs 31 `rolldown-worker` threads, and it dies 12–40 s after launch. That window is when the Vitest process is busiest (peak RSS about 4.4 GB, transforms and module graph building). Node 24.15.0, macOS 26.6.2, arm64. **Stress runs.** Full `vp test`, with 8 `yes > /dev/null` CPU burners, while other sessions ran their own suites on the same machine (load average above 20): | Setup | Runs | Exit 139 | Other failures | | --- | --- | --- | --- | | `main`, unmitigated | 10 | **1** (after 19 s, no test failed) | 1 console test over its 5 s timeout; 1 run invalid (I ran two suites in one checkout, which race on `dist/test-assets`) | | unmitigated, amplified (`--always-sparkplug --no-maglev`, the upstream repro's flags, set by a `NODE_OPTIONS=--import` preload) | 5 | 0 in Vitest; **1 in the Wrangler bundling child**, reported only as the `ENOENT` above | none | | this branch | 10 | **0** | 1 `agent.test.ts` `vi.waitFor` over its 1 s default | | this branch, amplified | 5 | **0** | none | A fifth crash report today (12:42) came from another worktree's unmitigated suite that was running at the same time. **Cost.** None measurable. Mean of the passing runs: 271 s unmitigated vs 278 s mitigated under load, and 409 s vs 401 s amplified, both within the load noise. Most of the heavy work runs in native code (rolldown) and in workerd, not in the Vitest process's baseline tier. **What it isn't:** - *Memory growth.* The Vitest process's RSS peaks at about 4.4 GB early, then falls to 0.5–2 GB, with no upward trend across a run. The workerd runtimes stay flat (about 19 × 200 MB). - *`isolate: false` runtime reuse.* That reuse happens inside workerd, which has its own V8, while the crash is in Node's V8. - *The "code had hung … canceled this request" messages* (about 550–800 per run, in every run, crashing or not). workerd logs one whenever a test's RPC call into connect rejects as the test expects. `apps/connect/test/oauth.test.ts` alone prints 26 for 26 passing tests, and connect doesn't use `isolate: false`. They're log noise inside workerd, not a symptom of this bug. ## How it was tested The stress runs above, plus `vp check` on the changed files. An independent review then checked that the Wrangler path resolves under pnpm locally and in CI, and that loading `vite.config.ts` elsewhere (lint, fmt, build, dev) is harmless: it only turns off one of V8's JIT tiers. I also checked that a runtime `--no-sparkplug` beats `--always-sparkplug` set before it, and the other way round, so the amplified runs with the fix really ran without Sparkplug. ## Checklist - [x] `vp check` and `vp test` pass - [x] Schema changes are expand-only (no drops or renames in the same release) - [x] New external calls go through connect; new model calls through the model gateway - [x] No secrets, client names or client configuration - [x] Risky or user-visible changes ship behind a feature flag (test tooling only)
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Backport of v8/v8@0b94a9f ("[leaptiering] Fix BaselineOutOfLinePrologue builtin") to the V8 13.6 in
v24.x. Applies cleanly todeps/v8;v8_embedder_stringbumped to-node.54.Fixes: #62393
What crashes
A register that nothing guarantees, pushed on the stack as if it were a tagged pointer, then read by the GC:
sequenceDiagram participant M as JS caller (Sparkplug code) participant CL as CompileLazy (TurboFan/CSA-generated) participant P as BaselineOutOfLinePrologue (hand-written asm) participant GC as mark-compact M->>CL: first call of f() Note over CL: x4 is NOT part of the JS linkage here<br/>(V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE undefined)<br/>register allocator uses x4 as scratch → x4 = 0x7 CL->>P: tail call into freshly compiled baseline code Note over P: stack guard slow path:<br/>Push(x4, new_target) // "dispatch handles always look like Smis" P->>GC: Runtime_StackGuardWithGap → CollectGarbage Note over GC: InternalFrame::Iterate → ClearStaleLeftTrimmedPointerVisitor<br/>IsHeapObject(0x7) is true (low bit set)<br/>reads map_word at 0x7 - 1 GC--xGC: SIGSEGV, KERN_INVALID_ADDRESS at 0x6Node.js builds V8 with leaptiering on and the sandbox off. In that configuration
V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLEis not defined (src/common/globals.h), sokJavaScriptCallDispatchHandleRegister(x4on arm64,r15on x64) is not carried by JS calls. Two hand-written builtins still gated the push onV8_ENABLE_LEAPTIERING_BOOL. The upstream commit message says it directly: "This issue triggered only on non-sandbox configuration with enabled leaptiering."Why only v24:
ClearStaleLeftTrimmedPointerVisitorstarted visiting stack roots in V8 12.6 (v8/v8@22c404b8bbbb, 2024-05). Node 22 ships V8 12.4, so the same stale slot was never dereferenced there. The fix landed in V8 main after the 13.6 branch cut and is present in Node 25/26.Evidence
Instrumented
v24.20.0(macOS arm64) logging every non-Smi, non-pointer value inINTERNALframe slots right before the GC visitor ran. Three independent crashes, identical frame:Matching crash report for the same worker:
EXC_BAD_ACCESS / KERN_INVALID_ADDRESS at 0x0000000000000006inClearStaleLeftTrimmedPointerVisitor::VisitRootPointers←InternalFrame::Iterate.Stack walk at that point (callee = first call of a small module-level helper, i.e. the
CompileLazypath):A temporary check at the builtin entry (
x4 == closure.dispatch_handle, abort otherwise) fires even for:node --always-sparkplug -e 'function f(a,b){return a+b}; f(1,2)'so this is not memory corruption: in this configuration the register simply never holds the handle.
The fix (same as upstream, all ports)
With the linkage macro undefined the slot now receives
padreg(xzr, i.e. Smi 0) instead of whatever was inx4. Same change on x64 (#ifdef V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLEaround the push/pop), loong64, mips64 and riscv.Verification
Same jest suite (322 suites / 5184 tests), 4 workers,
--no-maglev --always-sparkplug --max-old-space-size=512, macOS arm64:The last column matters more than the crash count: the instrumentation logged any non-pointer tagged value in an
INTERNALframe slot whether or not a GC happened to hit it.I have not run V8 CI /
make test-v8; please trigger it.Refs: v8/v8@0b94a9f