Repository navigation
test_runner: child process exits 0xC0000005 on Windows after its tests pass #65756
Description
Activity
@dirkwa
Thanks for reporting. This looks like a real child process crash rather than a test assertion failure.Could you check a few things to help narrow it down?
- Does this still reproduce with Node.js v24.20.0 and the latest v26.x?
- Does it reproduce with
--test-isolation=none? - If the repository / failing workflow run is public, could you share a link to it?
The
--test-isolation=noneresult would be especially useful, since the default test runner mode runs test files in child Node.js processes. If the crash disappears without process isolation, that would help narrow this toward the child process lifecycle / teardown path.Thanks for looking at this. Answers so far:
Public run: yes — here is a retained occurrence from a different day, hitting a different test file (consistent with the affected file varying between occurrences): https://github.com/dirkwa/signalk-container/actions/runs/32586760771/job/97064199368 —
dist\test\classifyVolumeSources.test.jsreported failed at:1:1while all of its own tests pass, on node v24.19.0. The repository is https://github.com/dirkwa/signalk-container; the JUnit artifact for that run has expired, but the job log is public.Versions: every occurrence I have observed was on v24.19.0 — CI had not yet picked up 24.20.0 (released Aug 26), so I have no signal either way on newer versions. Given the ~1-in-8 rate a single green run means nothing, so I have started a stress workflow repeating the suite 30× per leg on
windows-latestacross 24.19.0 / 24.20.0 / 26.8.1: https://github.com/dirkwa/signalk-container/actions/runs/34055954804--test-isolation=none: the suite has some cross-file state that fails assertions under that mode, so a clean pass/fail A/B is not possible as-is. But since the crash signature is a process exiting 0xC0000005 rather than an assertion failure, the same stress matrix includes isolation=none legs and watches exit codes specifically — the runs go through aspawnSyncdriver so the raw 32-bit status is captured (bash truncates NTSTATUS), and child crashes are detected via theexitCode:field thattestCodeFailureentries carry in the JUnit output.I will report the matrix results here when it finishes.
Update: the flake just reproduced again — in the regular CI run triggered by the very push that added the stress harness, so there is now a fresh public occurrence whose JUnit artifact is still retained (7 days from 2026-09-06):
- Run: https://github.com/dirkwa/signalk-container/actions/runs/34055955130 (Windows / Node 24 job)
- Node v24.19.0, different file again:
dist\test\resolveHostPath.test.jsreported failed at:1:1, all of its own tests✔ - The
test-results-windows-latest-node24artifact's junit.xml carries the file-level failure:
[Error: test failed] { code: 'ERR_TEST_FAILURE', failureType: 'testCodeFailure', cause: 'test failed', exitCode: 3221225477, signal: null }That makes four observed occurrences, each hitting a different test file, all on v24.19.0 so far.
Stress matrix interim: all three
--test-isolation=nonelegs (24.19.0 / 24.20.0 / 26.8.1, 30 runs each = 90 runs) completed with zero abnormal exits — consistent with the crash living in the child-process path. The default-isolation legs are still running; I will post the full matrix when they finish.Stress matrix results are in: https://github.com/dirkwa/signalk-container/actions/runs/34055954804 — 30 suite repetitions per leg on
windows-latest(~1330 tests / ~250 files per repetition), harness on branchstress-node65756. Exit statuses were captured raw via aspawnSyncdriver; a "hit" is a child-crash signature: anexitCode:field inside atestCodeFailureentry in the JUnit output, and/or the main process exiting outside {0,1}.Node isolation hits / 30 runs 24.19.0 default (child processes) 1 24.20.0 default (child processes) 2 26.8.1 default (child processes) 0 24.19.0 --test-isolation=none0 24.20.0 --test-isolation=none0 26.8.1 --test-isolation=none0 All three hits have the identical signature: main runner exits 1, one file-level testcase carries
exitCode: 3221225477(0xC0000005), the crashed file's own tests all pass (failures="0"in its testsuite), and the affected file differs per occurrence (libpodNetworkBackendInfo.test.js×2,getLiveResources.test.js×1 — different files again from the four earlier occurrences).So, answering the three questions:
- v24.20.0: yes, still reproduces (2/30). v26.8.1: 0/30 — but at the observed ~5% per-suite rate, 30 clean runs are only weak evidence of absence (~0.95³⁰ ≈ 21% chance of seeing nothing even at the same rate). I can run a longer 26.x-only batch if useful.
--test-isolation=none: 0 hits in 90 runs across all three versions, and the main process never exited abnormally. Consistent with the crash living in the child-process lifecycle/teardown path, as suspected.- Public run links are above and in the previous comments; the hit runs' spec + JUnit logs are uploaded as artifacts on the stress run with 30-day retention (
HIT-run-*.{spec.log,junit.xml}insidestress-24.19.0-isolation-defaultandstress-24.20.0-isolation-default).
Happy to run further variations (more 26.x iterations,
--test-concurrencysweeps, or a crash-dump collection approach if you can suggest one that works on hosted runners).@dirkwa
Thanks, this is really helpful. The--test-isolation=noneresults make the child-process path look suspicious.If you’re up for one more check, could you try running each test file as a separate top-level process from the stress driver, e.g.
node --test --test-isolation=none <file>for each file? That might help tell whether simply starting/stopping lots of Node processes is enough to trigger it, or whether the test runner’s process-isolation path is involved.Also, another ~30 runs on 26.8.1 would be useful if it’s not too costly.
Round 2 done: https://github.com/dirkwa/signalk-container/actions/runs/34080855285
Per your suggestion, the driver now has a
perfilemode: it spawnsnode --test --test-isolation=none <file>itself for each of the ~250 test files per iteration — the same process churn as the default runner mode, but without the runner's process-isolation path. Every per-file exit status is captured raw viaspawnSync; any status outside {0,1} or a signal counts as a hit.Node mode result 24.19.0 perfile (driver-spawned, ~7,500 processes) 0 hits, every process exited 0 24.20.0 perfile (driver-spawned, ~7,500 processes) 0 hits, every process exited 0 26.8.1 default (runner child processes), 30 more runs 0 hits (now 0/60 cumulative) So ~15,000 short-lived Node processes on the two versions that crash under the runner's default isolation (3 hits / 60 runs there) produced not a single abnormal exit when spawned directly. On the observed per-child rate (~3/15,000) the expected count for the perfile legs was ~3, so P(0) ≈ e⁻³ ≈ 5% if plain process start/stop were the trigger — pointing toward the runner's process-isolation plumbing (IPC channel / child lifecycle / teardown) rather than generic process churn.
Two caveats: the perfile children differ from runner children in more than the isolation path (no IPC pipe to a parent runner, plain stdio piping,
--test-isolation=noneinside the child), and 26.8.1's cumulative 0/60 is likewise suggestive (~0.95⁶⁰ ≈ 5% chance if the rate were unchanged) but not proof of absence.Harness is on the same branch; happy to run further variations — e.g. per-file with default isolation (
node --test <file>, one runner + one child per file) to bisect between the IPC/teardown path and child count, if that would help.Thanks, this narrows it down a lot. The per-file result makes generic process churn look unlikely, so the runner-managed isolation path seems more suspicious.
The per-file default-isolation test (
node --test <file>) sounds like a good next step to me.Round 3 (per-file with default isolation) is done: https://github.com/dirkwa/signalk-container/actions/runs/34392228050 — driver spawns
node --test <file>per test file, 30 iterations per leg.Node mode result 24.19.0 perfile-default (one runner, one child, per file) 1 hit / 30 runs 24.20.0 perfile-default (one runner, one child, per file) 0 / 30 The hit is
dist\test\selectImagesToReap.test.json iteration 23: its own 18 tests all pass (tests="18" failures="0") while the file-level testcase carries[Error: test failed] { code: 'ERR_TEST_FAILURE', failureType: 'testCodeFailure', cause: 'test failed', exitCode: 3221225477, signal: null }— identical to every earlier occurrence.
Methodology note worth flagging. In this mode the crashed child is reported by its runner, which then exits 1 — indistinguishable from an ordinary assertion failure by exit status alone. Detecting the crash requires reading the
testCodeFailureentry'sexitCode/signalfrom the runner's own output; my first cut of this leg keyed on exit status and would have reported a false clean 0/30. I validated the corrected detector against a fixture that aborts after its test passes (caught,signal: 'SIGABRT') and against a plain assertion failure (correctly not flagged).What the three rounds together say:
node --test --test-isolation=none <file>per file does not crash;node --test <file>per file does. One runner managing a single isolation child is enough to reproduce it, so the accumulation of many children under one runner is not required. That points at the per-child isolation lifecycle rather than multi-child bookkeeping.Correction to my earlier comments. I wrote "~250 test files"; the actual unit-suite count is 90 (96 files in
dist/including stale output, plus 16 integration files thatnpm testdoes not run). So the round-2 perfile legs were ~2,700 processes each, not ~15,000, and the "expected ~3 hits, P(0) ≈ 5%" estimate there was built on that inflated denominator. Restated honestly: 0 crashes in ~5,400 perfile processes versus 1 in ~2,700 perfile-default — suggestive in the same direction, but a weaker separation than I implied.Harness is on the same branch if further variations would help.
Reacted by Yuya InoueFollowing the earlier report of 0 crashes on v26.8.1 #65756 (comment), I extended the stress run to a maximum of 300 iterations and observed the same
0xC0000005child-process crash.I captured a native dump in that run. The analysis showed a write to a freed Realm through
CppgcMixin::~CppgcMixin()duringContextifyScriptdestruction. This matches the lifetime issue described and addressed in #65778.To check whether that fix also addresses the crashes on v24, I applied the changes (c909c63) to v24.20.0 and ran the Windows stress workload:
- With the fix v24.20.0: all 300 iterations passed, with no native crashes or test failures.
- Unmodified v24.20.0: a child process crashed with
3221225477(0xC0000005) at iteration 21, and the run stopped at the first crash.
Both builds used the same build configuration and pinned fixture, with default isolation,
--test-concurrency=1, and spec/JUnit reporters.These results suggest that #65778 addresses this crash on v24 as well.
- addedwindowsIssues and PRs related to the Windows platform.Issues and PRs related to the Windows platform.v24.xIssues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch.Issues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch.v26.xIssues that can be reproduced on v26.x or PRs targeting the v26.x-staging branch.Issues that can be reproduced on v26.x or PRs targeting the v26.x-staging branch.
on Sep 13, 2026 I've asked in #65778 (comment) whether this fix can be backported to v26.x and v24.x.
Awesome, thanks a lot for your effort
Reacted by Yuya InoueRomiiansyah commented
on Sep 19, 2026 on Sep 19, 2026 via email · Hidden as spamshow commentMore actions
Version
v24.19.0
Platform
windows-2025-vs2026(GitHub-hostedwindows-latestrunner), x64Subsystem
test_runner
What steps will reproduce the bug?
Run a multi-file suite under
node --teston Windows:Intermittently one test file is reported as failed even though every test in it
passed. I have not found a minimal reproducer — it is timing-dependent and does
not reproduce on demand.
How often does it reproduce? Is there a required condition?
Intermittent, perhaps 1 run in 8 across a ~1330-test suite. Only on the Windows
leg; the same commit passes on Linux, arm64 (QEMU) and macOS in the same
matrix, and passes on Windows when the job is re-run with no change.
--test-concurrency=1does not prevent it.What is the expected behavior? Why is that the expected behavior?
A test file whose tests all pass should be reported as passing. The process
exit status should reflect the test results.
What do you see instead?
The file is reported failed at
:1:1— the whole-file marker rather than anyassertion — while its own tests all show
✔:The JUnit reporter contradicts this in the same run — zero failures overall:
and the suite for that file reports every test passing:
while the file-level testcase in the same document carries:
3221225477is0xC0000005—STATUS_ACCESS_VIOLATION. The test file's childprocess appears to crash after completing its tests, rather than exiting
non-zero from user code.
Additional information
change under test each time. The run above was on a PR touching a different
file entirely.
which suggests the crash also truncates reporter accounting.
exitCodevalue from one retained CI artifact (7-day retentionexpired the earlier ones). The behavioural pattern — all tests pass, file
reported failed, clean re-run — is from several occurrences.
I am happy to gather more if it would help narrow this: a specific reporter
combination,
--test-concurrencysetting, or an approach to collecting a crashdump on a GitHub-hosted Windows runner.