Skip to content

[v24.x] deps: V8: backport 0b94a9fd23ba - #65753

Closed
ruangustavo wants to merge 1 commit into
nodejs:v24.x-stagingfrom
ruangustavo:backport-v8-0b94a9fd23ba
Closed

ruangustavo wants to merge 1 commit into
nodejs:v24.x-stagingfrom
ruangustavo:backport-v8-0b94a9fd23ba

Conversation

@ruangustavo

Copy link
Copy Markdown

Backport of v8/v8@0b94a9f ("[leaptiering] Fix BaselineOutOfLinePrologue builtin") to the V8 13.6 in v24.x. Applies cleanly to deps/v8; v8_embedder_string bumped to -node.54.

Fixes: #62393

What crashes

A register that nothing guarantees, pushed on the stack as if it were a tagged pointer, then read by the GC:

sequenceDiagram
    participant M as JS caller (Sparkplug code)
    participant CL as CompileLazy (TurboFan/CSA-generated)
    participant P as BaselineOutOfLinePrologue (hand-written asm)
    participant GC as mark-compact
    M->>CL: first call of f()
    Note over CL: x4 is NOT part of the JS linkage here<br/>(V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE undefined)<br/>register allocator uses x4 as scratch → x4 = 0x7
    CL->>P: tail call into freshly compiled baseline code
    Note over P: stack guard slow path:<br/>Push(x4, new_target)  // "dispatch handles always look like Smis"
    P->>GC: Runtime_StackGuardWithGap → CollectGarbage
    Note over GC: InternalFrame::Iterate → ClearStaleLeftTrimmedPointerVisitor<br/>IsHeapObject(0x7) is true (low bit set)<br/>reads map_word at 0x7 - 1
    GC--xGC: SIGSEGV, KERN_INVALID_ADDRESS at 0x6
Loading

Node.js builds V8 with leaptiering on and the sandbox off. In that configuration V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE is not defined (src/common/globals.h), so kJavaScriptCallDispatchHandleRegister (x4 on arm64, r15 on x64) is not carried by JS calls. Two hand-written builtins still gated the push on V8_ENABLE_LEAPTIERING_BOOL. The upstream commit message says it directly: "This issue triggered only on non-sandbox configuration with enabled leaptiering."

Why only v24: ClearStaleLeftTrimmedPointerVisitor started visiting stack roots in V8 12.6 (v8/v8@22c404b8bbbb, 2024-05). Node 22 ships V8 12.4, so the same stale slot was never dereferenced there. The fix landed in V8 main after the 13.6 branch cut and is present in Node 25/26.

Evidence

Instrumented v24.20.0 (macOS arm64) logging every non-Smi, non-pointer value in INTERNAL frame slots right before the GC visitor ran. Three independent crashes, identical frame:

builtin = BaselineOutOfLinePrologue   code_offset = 116   INTERNAL frame, 6 slots from sp:

[0] 0x7000000000    SmiTag(frame_size)              ok
[1] 0x0             padreg (PushArgument)           ok
[2] 0x10b697280011  new_target                      ok, valid HeapObject
[3] 0x7             kJavaScriptCallDispatchHandleRegister   <-- stale; callee's real dispatch_handle was 0x143c900
[4] 0x0             padreg (EnterFrame)             ok
[5] 0x2e            INTERNAL frame marker (23 << 1) ok

Matching crash report for the same worker: EXC_BAD_ACCESS / KERN_INVALID_ADDRESS at 0x0000000000000006 in ClearStaleLeftTrimmedPointerVisitor::VisitRootPointers ← InternalFrame::Iterate.

Stack walk at that point (callee = first call of a small module-level helper, i.e. the CompileLazy path):

 0  EXIT      CEntry_Return1_ArgvOnStack_NoBuiltinExit
 1  INTERNAL  BaselineOutOfLinePrologue                          <-- bad slot here
 2  BASELINE  forEach            (axios.cjs)   dispatch_handle 0x143c900
 3  BASELINE  <module wrapper>   (axios.cjs)   dispatch_handle 0x143c400
 4  TURBOFAN  _execModule        (jest-runtime)

A temporary check at the builtin entry (x4 == closure.dispatch_handle, abort otherwise) fires even for:

node --always-sparkplug -e 'function f(a,b){return a+b}; f(1,2)'

so this is not memory corruption: in this configuration the register simply never holds the handle.

The fix (same as upstream, all ports)

 deps/v8/src/builtins/arm64/builtins-arm64.cc          Generate_BaselineOutOfLinePrologue
 deps/v8/src/codegen/arm64/macro-assembler-arm64.cc    GenerateTailCallToReturnedCode
-    Register maybe_dispatch_handle = V8_ENABLE_LEAPTIERING_BOOL
+    Register maybe_dispatch_handle = V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE_BOOL
                                          ? kJavaScriptCallDispatchHandleRegister
                                          : padreg;
     static_assert(kJSDispatchHandleShift > 0);
+    __ AssertSmi(maybe_dispatch_handle);
     __ Push(maybe_dispatch_handle, new_target);

With the linkage macro undefined the slot now receives padreg (xzr, i.e. Smi 0) instead of whatever was in x4. Same change on x64 (#ifdef V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE around the push/pop), loong64, mips64 and riscv.

Verification

Same jest suite (322 suites / 5184 tests), 4 workers, --no-maglev --always-sparkplug --max-old-space-size=512, macOS arm64:

build runs runs with a SIGSEGV worker invalid slots logged by instrumentation
v24.20.0 official, 12 workers, under load 5 5 n/a
v24.20.0 + instrumentation, no fix 4 1 3
v24.20.0 + instrumentation + this patch 6 0 0

The last column matters more than the crash count: the instrumentation logged any non-pointer tagged value in an INTERNAL frame slot whether or not a GC happened to hit it.

I have not run V8 CI / make test-v8; please trigger it.

Refs: v8/v8@0b94a9f

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

Review requested:

  • @nodejs/gyp
  • @nodejs/security-wg
  • @nodejs/v8-update

@nodejs-github-bot nodejs-github-bot added build Issues and PRs related to Node.js builds or CI infrastructure. needs-ci PRs that need a full CI run. v24.x Issues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch. v8 engine Issues and PRs related to the V8 dependency. labels Sep 3, 2026
@Renegade334 Renegade334 changed the title deps: V8: backport 0b94a9fd23ba [v24.x] deps: V8: backport 0b94a9fd23ba Sep 3, 2026
@richardlau richardlau added the request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. label Sep 3, 2026
@github-actions github-actions Bot removed the request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. label Sep 3, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@ruangustavo

Copy link
Copy Markdown
Author

Hey @richardlau, let me know how can I help you fixing the pipeline :)

@nodejs-github-bot

This comment was marked as outdated.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@aduh95
aduh95 force-pushed the v24.x-staging branch 3 times, most recently from 7fc125b to e711d3f Compare September 9, 2026 21:47
Original commit message:

    [leaptiering] Fix BaselineOutOfLinePrologue builtin

    ... which tried to preserve kJavaScriptCallDispatchHandleRegister even
    on configurations where it's not used which resulted in a random value
    on the stack discoverable by GC.
    This issue triggered only on non-sandbox configuration with enabled
    leaptiering.

    Drive-by: fix MacroAssembler::GenerateTailCallToReturnedCode() on riscv
    port which wasn't preserving dispatch handle as all the other ports do.

    Bug: 42204201
    Fixed: 413769394

    Change-Id: If146b0b7a6cf972ed5a881142f40980774f19cba
    Reviewed-on: https://chromium-review.googlesource.com/c/v8/v8/+/6587010
    Commit-Queue: Igor Sheludko <ishell@chromium.org>
    Reviewed-by: Olivier Flückiger <olivf@chromium.org>
    Cr-Commit-Position: refs/heads/main@{#100512}

Node.js 24 builds V8 with leaptiering enabled and the sandbox disabled,
so V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE is not defined and the JS
calling convention does not carry the dispatch handle register (x4 on
arm64). BaselineOutOfLinePrologue and GenerateTailCallToReturnedCode
still pushed that register as a tagged slot of an INTERNAL frame, so
whatever value the caller left there is dereferenced by
ClearStaleLeftTrimmedPointerVisitor during mark-compact root scanning
and crashes the process with SIGSEGV (seen as jest workers dying).

Refs: v8/v8@0b94a9f
Fixes: nodejs#62393
@ruangustavo
ruangustavo force-pushed the backport-v8-0b94a9fd23ba branch from a315622 to f47b5f8 Compare September 10, 2026 03:47
@ruangustavo

Copy link
Copy Markdown
Author

Hey @richardlau, I rebased onto the current v24.x-staging: the branch was rewritten after v24.21.0 and the PR started showing 97 commits that aren't part of this change. It's the same single commit as before. Could you add request-ci again? The last run (77190) was green.

@Romiiansyah

This comment was marked as spam.

bbertucc added a commit to EqualifyEverything/equalify-iris that referenced this pull request Sep 14, 2026
…segfault (#477)

#405's crash has an upstream name — nodejs/node#62393, V8 CL 0b94a9fd23ba: the
BaselineOutOfLinePrologue builtin left a random value on the stack where the GC
found it. `ClearStaleLeftTrimmedPointerVisitor` reads it as a heap pointer and
faults at 0xe, the address in all five local reports. Sparkplug generates that
prologue, so `--no-sparkplug` removes the path.

The two decisions #405 was holding both answer no. Node cannot be bumped into the
fix: the backport (nodejs/node#65753) is open, not landed, so no released 24.x has
it, and one upstream report has the crash live on v26.7.0. CI needs no dead-child
retry: every occurrence here and upstream is macOS arm64, and every workflow runs
ubuntu-latest.

Cost is nothing measurable — 55.8 s mean either way over two runs each locally,
~3% on Linux x64 over the two heaviest jsdom files, against a 22-minute review
step. The flag is on the test script only, on purpose: `npm start` and `npm run
dev` keep the Sparkplug path, because one dev server dying is loud where a dead
test child reads as a clean run with a short pass count.

A test pins the flag, and after round 1 it pins its *position*: `node --test
"glob" --no-sparkplug` exits 0, warns nothing, and the child's `execArgv` does not
carry the flag, so a presence-only check could go green through the exact
regression it exists to catch. Drop the flag and that test together when
nodejs/node#65753 ships in a 24.x release.

Two rounds, both approved; round 1's three notes fixed, round 2 clean.

Refs #405

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@richardlau richardlau added the request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. label Sep 16, 2026
@github-actions github-actions Bot added request-ci-failed Starting CI with the request-ci label failed and requires manual intervention. and removed request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. labels Sep 16, 2026
@github-actions

Copy link
Copy Markdown
Contributor
Failed to start CI
�[36m⠋�[39m Getting reviews from nodejs/node/pull/65753
�[36m⠋�[39m Getting commits from nodejs/node/pull/65753
�[36m⠙�[39m Validating Jenkins credentials
�[36m⠙�[39m Validating Jenkins credentials
✔  Jenkins credentials valid
�[36m⠹�[39m Getting comments from nodejs/node/pull/65753
�[36m⠸�[39m Querying data for job/node-test-pull-request/77190/
�[36m⠸�[39m Querying data for job/node-test-pull-request/77190/
�[36m⠸�[39m Querying API for job/node-test-pull-request/77190/
SyntaxError: Unexpected token '<', ..."    
  https://github.com/nodejs/node/actions/runs/35160272536

@richardlau richardlau added request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. and removed request-ci-failed Starting CI with the request-ci label failed and requires manual intervention. labels Sep 17, 2026
@github-actions github-actions Bot removed the request-ci Add this label to start a Jenkins CI on a PR. Only starts once the PR has an approving review. label Sep 17, 2026
@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

@ruangustavo

Copy link
Copy Markdown
Author

Hey @richardlau, this is the second time the pipeline has failed at the same stage. lmk how I can fix it, I can’t see the error details.

@nodejs-github-bot

Copy link
Copy Markdown
Collaborator

richardlau pushed a commit that referenced this pull request Sep 29, 2026
Original commit message:

    [leaptiering] Fix BaselineOutOfLinePrologue builtin

    ... which tried to preserve kJavaScriptCallDispatchHandleRegister even
    on configurations where it's not used which resulted in a random value
    on the stack discoverable by GC.
    This issue triggered only on non-sandbox configuration with enabled
    leaptiering.

    Drive-by: fix MacroAssembler::GenerateTailCallToReturnedCode() on riscv
    port which wasn't preserving dispatch handle as all the other ports do.

    Bug: 42204201
    Fixed: 413769394

    Change-Id: If146b0b7a6cf972ed5a881142f40980774f19cba
    Reviewed-on: https://chromium-review.googlesource.com/c/v8/v8/+/6587010
    Commit-Queue: Igor Sheludko <ishell@chromium.org>
    Reviewed-by: Olivier Flückiger <olivf@chromium.org>
    Cr-Commit-Position: refs/heads/main@{#100512}

Node.js 24 builds V8 with leaptiering enabled and the sandbox disabled,
so V8_JS_LINKAGE_INCLUDES_DISPATCH_HANDLE is not defined and the JS
calling convention does not carry the dispatch handle register (x4 on
arm64). BaselineOutOfLinePrologue and GenerateTailCallToReturnedCode
still pushed that register as a tagged slot of an INTERNAL frame, so
whatever value the caller left there is dereferenced by
ClearStaleLeftTrimmedPointerVisitor during mark-compact root scanning
and crashes the process with SIGSEGV (seen as jest workers dying).

Refs: v8/v8@0b94a9f
Fixes: #62393
PR-URL: #65753
Reviewed-By: Richard Lau <richard.lau@ibm.com>
@richardlau

Copy link
Copy Markdown
Member

Landed in fb0d0ca.

@richardlau richardlau closed this Sep 29, 2026
nickruigrok added a commit to grasplabs/work that referenced this pull request Sep 29, 2026
…s them

Node 24's V8 pushes a stale register in Sparkplug's BaselineOutOfLinePrologue,
which the GC reads as a pointer and segfaults on (nodejs/node#62393; the
backport is open in nodejs/node#65753). It killed the Vitest process mid-run
with exit 139, and can kill the Wrangler child that bundles connect for core's
tests, which Wrangler's launcher reports as success.
nickruigrok added a commit to grasplabs/work that referenced this pull request Sep 29, 2026
…s them (#219)

## What and why

`vp test` sometimes exited with code 139 (SIGSEGV) partway through a
full local run, with no test failing, and passed on a rerun. The cause
isn't our code, workerd or vitest-pool-workers. It's a V8 bug in Node
24: **nodejs/node#62393**.

Sparkplug's `BaselineOutOfLinePrologue` pushes a register that Node's V8
build never fills (leaptiering on, sandbox off), so its stack slot holds
junk such as `0x7`. If a major GC runs at that moment (reached through a
stack-guard interrupt), `ClearStaleLeftTrimmedPointerVisitor` reads the
junk as a heap pointer and faults at `0x6`/`0xe`. The fix is V8
`0b94a9fd23ba`, which is in Node 25+. The backport to 24.x,
nodejs/node#65753, is **open, not landed**, and reporters on the issue
still hit the crash on 24.21. So no Node release inside `devEngines`
(`^24.11.0`) fixes it yet, and there's nothing to pin or upgrade.

The mitigation is to switch the baseline tier off (`--no-sparkplug`) in
the two Node processes that hit it:

- **The Vitest process** (`vite.config.ts`). `vp test` spawns `node
vitest.mjs` itself and gives us no place to pass Node flags, and
`NODE_OPTIONS` rejects `--no-sparkplug`. So the root config calls
`v8.setFlagsFromString("--no-sparkplug")`, which Vitest loads first. I
checked that the runtime flag stops baseline compilation just like the
CLI flag (`%ActiveTierIsSparkplug` stays false).
- **The Wrangler CLI that core's global setup runs to bundle connect**
(`apps/core/test/build-connect.ts`). It hit the same crash in an
amplified run. Wrangler's launcher turns its child dying from a signal
into exit code 0 (`code === null ? 0`), so the only symptom was a
misleading `ENOENT … grasp-os-connect-*/index.js` from the global setup.
The setup now runs `node --no-sparkplug wrangler.js`, and the launcher
passes `execArgv` on to the CLI.

Both comments say to drop this once a Node 24 release in `devEngines`
carries the backport.

**Limits.** A runtime flag stops only *new* baseline compilation. The
few hundred functions Node compiles to baseline while starting up,
before Vitest loads the config (module loader internals such as
`createRequire` and `initializeImportMeta`), keep their Sparkplug code.
So this makes the crash rare, not impossible: 0 in 15 mitigated runs
against about 1 in 10 without it. The flag also covers only `vp test`
from the repo root: a run with an app's own config (`vp -C apps/core
test`) and Vitest's `forks` pool children don't load the root config.
None of those crashed in these runs. If an exit 139 still turns up
before the backport ships, the next step is running Vitest ourselves as
`node --no-sparkplug`, not dropping this.

Wrangler's launcher reports any signal death of its CLI as success, not
just this segfault (an OOM kill does the same). So `bundleConnect` now
says Wrangler wrote no bundle and points at a possible crash, instead of
failing on a bare missing-file error. I checked this by killing the CLI
with SIGSEGV from a preload.

## Evidence

**Crash signature.** macOS kept crash reports for every segfault:
`~/Library/Logs/DiagnosticReports/node-*.ips` (8 since 2026-09-27, 4 of
them today). Every one has the same faulting frames:

```text
EXC_BAD_ACCESS (SIGSEGV)  KERN_INVALID_ADDRESS at 0x0000000000000006 / 0x000000000000000e
v8::internal::ClearStaleLeftTrimmedPointerVisitor::VisitRootPointers
v8::internal::InternalFrame::Iterate
v8::internal::Isolate::Iterate
v8::internal::Heap::IterateRoots
v8::internal::MarkCompactCollector::MarkRoots / MarkLiveObjects / CollectGarbage
v8::internal::Heap::CollectGarbage
v8::internal::StackGuard::HandleInterrupts
v8::internal::Runtime_StackGuardWithGap
Builtins_CEntry_Return1_ArgvOnStack_NoBuiltinExit
Builtins_BaselineOutOfLinePrologue
```

The crashed process is the `node …/vitest/vitest.mjs run` child of `vp
test`: its parent is `node`, it runs 31 `rolldown-worker` threads, and
it dies 12–40 s after launch. That window is when the Vitest process is
busiest (peak RSS about 4.4 GB, transforms and module graph building).
Node 24.15.0, macOS 26.6.2, arm64.

**Stress runs.** Full `vp test`, with 8 `yes > /dev/null` CPU burners,
while other sessions ran their own suites on the same machine (load
average above 20):

| Setup | Runs | Exit 139 | Other failures |
| --- | --- | --- | --- |
| `main`, unmitigated | 10 | **1** (after 19 s, no test failed) | 1
console test over its 5 s timeout; 1 run invalid (I ran two suites in
one checkout, which race on `dist/test-assets`) |
| unmitigated, amplified (`--always-sparkplug --no-maglev`, the upstream
repro's flags, set by a `NODE_OPTIONS=--import` preload) | 5 | 0 in
Vitest; **1 in the Wrangler bundling child**, reported only as the
`ENOENT` above | none |
| this branch | 10 | **0** | 1 `agent.test.ts` `vi.waitFor` over its 1 s
default |
| this branch, amplified | 5 | **0** | none |

A fifth crash report today (12:42) came from another worktree's
unmitigated suite that was running at the same time.

**Cost.** None measurable. Mean of the passing runs: 271 s unmitigated
vs 278 s mitigated under load, and 409 s vs 401 s amplified, both within
the load noise. Most of the heavy work runs in native code (rolldown)
and in workerd, not in the Vitest process's baseline tier.

**What it isn't:**

- *Memory growth.* The Vitest process's RSS peaks at about 4.4 GB early,
then falls to 0.5–2 GB, with no upward trend across a run. The workerd
runtimes stay flat (about 19 × 200 MB).
- *`isolate: false` runtime reuse.* That reuse happens inside workerd,
which has its own V8, while the crash is in Node's V8.
- *The "code had hung … canceled this request" messages* (about 550–800
per run, in every run, crashing or not). workerd logs one whenever a
test's RPC call into connect rejects as the test expects.
`apps/connect/test/oauth.test.ts` alone prints 26 for 26 passing tests,
and connect doesn't use `isolate: false`. They're log noise inside
workerd, not a symptom of this bug.

## How it was tested

The stress runs above, plus `vp check` on the changed files. An
independent review then checked that the Wrangler path resolves under
pnpm locally and in CI, and that loading `vite.config.ts` elsewhere
(lint, fmt, build, dev) is harmless: it only turns off one of V8's JIT
tiers. I also checked that a runtime `--no-sparkplug` beats
`--always-sparkplug` set before it, and the other way round, so the
amplified runs with the fix really ran without Sparkplug.

## Checklist

- [x] `vp check` and `vp test` pass
- [x] Schema changes are expand-only (no drops or renames in the same
release)
- [x] New external calls go through connect; new model calls through the
model gateway
- [x] No secrets, client names or client configuration
- [x] Risky or user-visible changes ship behind a feature flag (test
tooling only)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Issues and PRs related to Node.js builds or CI infrastructure. needs-ci PRs that need a full CI run. v8 engine Issues and PRs related to the V8 dependency. v24.x Issues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants