Skip to content

test_runner: child process exits 0xC0000005 on Windows after its tests pass #65756

Description

@dirkwa

Version

v24.19.0

Platform

windows-2025-vs2026 (GitHub-hosted windows-latest runner), x64

Subsystem

test_runner

What steps will reproduce the bug?

Run a multi-file suite under node --test on Windows:

node --test --test-concurrency=1 \
     --test-reporter=spec --test-reporter-destination=stdout \
     --test-reporter=junit --test-reporter-destination=test-results/junit.xml \
     "dist/test/*.test.js"

Intermittently one test file is reported as failed even though every test in it
passed. I have not found a minimal reproducer — it is timing-dependent and does
not reproduce on demand.

How often does it reproduce? Is there a required condition?

Intermittent, perhaps 1 run in 8 across a ~1330-test suite. Only on the Windows
leg; the same commit passes on Linux, arm64 (QEMU) and macOS in the same
matrix, and passes on Windows when the job is re-run with no change.

--test-concurrency=1 does not prevent it.

What is the expected behavior? Why is that the expected behavior?

A test file whose tests all pass should be reported as passing. The process
exit status should reflect the test results.

What do you see instead?

The file is reported failed at :1:1 — the whole-file marker rather than any
assertion — while its own tests all show ✔:

✔ resolveSignalkNetworks (7.2616ms)
✖ dist\test\resolveSignalkNetworks.test.js (249.1325ms)
...
ℹ tests 1327
ℹ pass 1323
ℹ fail 1
  'test failed'

The JUnit reporter contradicts this in the same run — zero failures overall:

tests=1340  failures=0  errors=0

and the suite for that file reports every test passing:

<testsuite name="resolveSignalkNetworks" tests="5" failures="0" errors="0">

while the file-level testcase in the same document carries:

<testcase name="dist\test\resolveSignalkNetworks.test.js" time="0.249133">
  <failure type="testCodeFailure" message="test failed">
[Error: test failed] {
  code: 'ERR_TEST_FAILURE',
  failureType: 'testCodeFailure',
  cause: 'test failed',
  exitCode: 3221225477,
  signal: null
}
  </failure>
</testcase>

3221225477 is 0xC0000005 — STATUS_ACCESS_VIOLATION. The test file's child
process appears to crash after completing its tests, rather than exiting
non-zero from user code.

Additional information

  • The affected file differs between occurrences and has been unrelated to the
    change under test each time. The run above was on a PR touching a different
    file entirely.
  • The console and JUnit totals disagree in an affected run (1327 vs 1340),
    which suggests the crash also truncates reporter accounting.
  • I have the exitCode value from one retained CI artifact (7-day retention
    expired the earlier ones). The behavioural pattern — all tests pass, file
    reported failed, clean re-run — is from several occurrences.

I am happy to gather more if it would help narrow this: a specific reporter
combination, --test-concurrency setting, or an approach to collecting a crash
dump on a GitHub-hosted Windows runner.

Activity

  1. inoway46 commented on Sep 6, 2026

    @inoway46
    Contributor

    @dirkwa
    Thanks for reporting. This looks like a real child process crash rather than a test assertion failure.

    Could you check a few things to help narrow it down?

    • Does this still reproduce with Node.js v24.20.0 and the latest v26.x?
    • Does it reproduce with --test-isolation=none?
    • If the repository / failing workflow run is public, could you share a link to it?

    The --test-isolation=none result would be especially useful, since the default test runner mode runs test files in child Node.js processes. If the crash disappears without process isolation, that would help narrow this toward the child process lifecycle / teardown path.

  2. dirkwa commented on Sep 6, 2026

    @dirkwa
    Author

    Thanks for looking at this. Answers so far:

    Public run: yes — here is a retained occurrence from a different day, hitting a different test file (consistent with the affected file varying between occurrences): https://github.com/dirkwa/signalk-container/actions/runs/32586760771/job/97064199368 — dist\test\classifyVolumeSources.test.js reported failed at :1:1 while all of its own tests pass, on node v24.19.0. The repository is https://github.com/dirkwa/signalk-container; the JUnit artifact for that run has expired, but the job log is public.

    Versions: every occurrence I have observed was on v24.19.0 — CI had not yet picked up 24.20.0 (released Aug 26), so I have no signal either way on newer versions. Given the ~1-in-8 rate a single green run means nothing, so I have started a stress workflow repeating the suite 30× per leg on windows-latest across 24.19.0 / 24.20.0 / 26.8.1: https://github.com/dirkwa/signalk-container/actions/runs/34055954804

    --test-isolation=none: the suite has some cross-file state that fails assertions under that mode, so a clean pass/fail A/B is not possible as-is. But since the crash signature is a process exiting 0xC0000005 rather than an assertion failure, the same stress matrix includes isolation=none legs and watches exit codes specifically — the runs go through a spawnSync driver so the raw 32-bit status is captured (bash truncates NTSTATUS), and child crashes are detected via the exitCode: field that testCodeFailure entries carry in the JUnit output.

    I will report the matrix results here when it finishes.

  3. dirkwa commented on Sep 6, 2026

    @dirkwa
    Author

    Update: the flake just reproduced again — in the regular CI run triggered by the very push that added the stress harness, so there is now a fresh public occurrence whose JUnit artifact is still retained (7 days from 2026-09-06):

    [Error: test failed] { code: 'ERR_TEST_FAILURE', failureType: 'testCodeFailure', cause: 'test failed', exitCode: 3221225477, signal: null }
    

    That makes four observed occurrences, each hitting a different test file, all on v24.19.0 so far.

    Stress matrix interim: all three --test-isolation=none legs (24.19.0 / 24.20.0 / 26.8.1, 30 runs each = 90 runs) completed with zero abnormal exits — consistent with the crash living in the child-process path. The default-isolation legs are still running; I will post the full matrix when they finish.

  4. dirkwa commented on Sep 6, 2026

    @dirkwa
    Author

    Stress matrix results are in: https://github.com/dirkwa/signalk-container/actions/runs/34055954804 — 30 suite repetitions per leg on windows-latest (~1330 tests / ~250 files per repetition), harness on branch stress-node65756. Exit statuses were captured raw via a spawnSync driver; a "hit" is a child-crash signature: an exitCode: field inside a testCodeFailure entry in the JUnit output, and/or the main process exiting outside {0,1}.

    Node isolation hits / 30 runs
    24.19.0 default (child processes) 1
    24.20.0 default (child processes) 2
    26.8.1 default (child processes) 0
    24.19.0 --test-isolation=none 0
    24.20.0 --test-isolation=none 0
    26.8.1 --test-isolation=none 0

    All three hits have the identical signature: main runner exits 1, one file-level testcase carries exitCode: 3221225477 (0xC0000005), the crashed file's own tests all pass (failures="0" in its testsuite), and the affected file differs per occurrence (libpodNetworkBackendInfo.test.js ×2, getLiveResources.test.js ×1 — different files again from the four earlier occurrences).

    So, answering the three questions:

    1. v24.20.0: yes, still reproduces (2/30). v26.8.1: 0/30 — but at the observed ~5% per-suite rate, 30 clean runs are only weak evidence of absence (~0.95³⁰ ≈ 21% chance of seeing nothing even at the same rate). I can run a longer 26.x-only batch if useful.
    2. --test-isolation=none: 0 hits in 90 runs across all three versions, and the main process never exited abnormally. Consistent with the crash living in the child-process lifecycle/teardown path, as suspected.
    3. Public run links are above and in the previous comments; the hit runs' spec + JUnit logs are uploaded as artifacts on the stress run with 30-day retention (HIT-run-*.{spec.log,junit.xml} inside stress-24.19.0-isolation-default and stress-24.20.0-isolation-default).

    Happy to run further variations (more 26.x iterations, --test-concurrency sweeps, or a crash-dump collection approach if you can suggest one that works on hosted runners).

  5. inoway46 commented on Sep 7, 2026

    @inoway46
    Contributor

    @dirkwa
    Thanks, this is really helpful. The --test-isolation=none results make the child-process path look suspicious.

    If you’re up for one more check, could you try running each test file as a separate top-level process from the stress driver, e.g. node --test --test-isolation=none <file> for each file? That might help tell whether simply starting/stopping lots of Node processes is enough to trigger it, or whether the test runner’s process-isolation path is involved.

    Also, another ~30 runs on 26.8.1 would be useful if it’s not too costly.

  6. dirkwa commented on Sep 7, 2026

    @dirkwa
    Author

    Round 2 done: https://github.com/dirkwa/signalk-container/actions/runs/34080855285

    Per your suggestion, the driver now has a perfile mode: it spawns node --test --test-isolation=none <file> itself for each of the ~250 test files per iteration — the same process churn as the default runner mode, but without the runner's process-isolation path. Every per-file exit status is captured raw via spawnSync; any status outside {0,1} or a signal counts as a hit.

    Node mode result
    24.19.0 perfile (driver-spawned, ~7,500 processes) 0 hits, every process exited 0
    24.20.0 perfile (driver-spawned, ~7,500 processes) 0 hits, every process exited 0
    26.8.1 default (runner child processes), 30 more runs 0 hits (now 0/60 cumulative)

    So ~15,000 short-lived Node processes on the two versions that crash under the runner's default isolation (3 hits / 60 runs there) produced not a single abnormal exit when spawned directly. On the observed per-child rate (~3/15,000) the expected count for the perfile legs was ~3, so P(0) ≈ e⁻³ ≈ 5% if plain process start/stop were the trigger — pointing toward the runner's process-isolation plumbing (IPC channel / child lifecycle / teardown) rather than generic process churn.

    Two caveats: the perfile children differ from runner children in more than the isolation path (no IPC pipe to a parent runner, plain stdio piping, --test-isolation=none inside the child), and 26.8.1's cumulative 0/60 is likewise suggestive (~0.95⁶⁰ ≈ 5% chance if the rate were unchanged) but not proof of absence.

    Harness is on the same branch; happy to run further variations — e.g. per-file with default isolation (node --test <file>, one runner + one child per file) to bisect between the IPC/teardown path and child count, if that would help.

  7. inoway46 commented on Sep 7, 2026

    @inoway46
    Contributor

    Thanks, this narrows it down a lot. The per-file result makes generic process churn look unlikely, so the runner-managed isolation path seems more suspicious.

    The per-file default-isolation test (node --test <file>) sounds like a good next step to me.

  8. dirkwa commented on Sep 9, 2026

    @dirkwa
    Author

    Round 3 (per-file with default isolation) is done: https://github.com/dirkwa/signalk-container/actions/runs/34392228050 — driver spawns node --test <file> per test file, 30 iterations per leg.

    Node mode result
    24.19.0 perfile-default (one runner, one child, per file) 1 hit / 30 runs
    24.20.0 perfile-default (one runner, one child, per file) 0 / 30

    The hit is dist\test\selectImagesToReap.test.js on iteration 23: its own 18 tests all pass (tests="18" failures="0") while the file-level testcase carries

    [Error: test failed] { code: 'ERR_TEST_FAILURE', failureType: 'testCodeFailure', cause: 'test failed', exitCode: 3221225477, signal: null }
    

    — identical to every earlier occurrence.

    Methodology note worth flagging. In this mode the crashed child is reported by its runner, which then exits 1 — indistinguishable from an ordinary assertion failure by exit status alone. Detecting the crash requires reading the testCodeFailure entry's exitCode/signal from the runner's own output; my first cut of this leg keyed on exit status and would have reported a false clean 0/30. I validated the corrected detector against a fixture that aborts after its test passes (caught, signal: 'SIGABRT') and against a plain assertion failure (correctly not flagged).

    What the three rounds together say: node --test --test-isolation=none <file> per file does not crash; node --test <file> per file does. One runner managing a single isolation child is enough to reproduce it, so the accumulation of many children under one runner is not required. That points at the per-child isolation lifecycle rather than multi-child bookkeeping.

    Correction to my earlier comments. I wrote "~250 test files"; the actual unit-suite count is 90 (96 files in dist/ including stale output, plus 16 integration files that npm test does not run). So the round-2 perfile legs were ~2,700 processes each, not ~15,000, and the "expected ~3 hits, P(0) ≈ 5%" estimate there was built on that inflated denominator. Restated honestly: 0 crashes in ~5,400 perfile processes versus 1 in ~2,700 perfile-default — suggestive in the same direction, but a weaker separation than I implied.

    Harness is on the same branch if further variations would help.

  9. inoway46 commented on Sep 13, 2026

    @inoway46
    Contributor

    Following the earlier report of 0 crashes on v26.8.1 #65756 (comment), I extended the stress run to a maximum of 300 iterations and observed the same 0xC0000005 child-process crash.

    I captured a native dump in that run. The analysis showed a write to a freed Realm through CppgcMixin::~CppgcMixin() during ContextifyScript destruction. This matches the lifetime issue described and addressed in #65778.

    To check whether that fix also addresses the crashes on v24, I applied the changes (c909c63) to v24.20.0 and ran the Windows stress workload:

    • With the fix v24.20.0: all 300 iterations passed, with no native crashes or test failures.
    • Unmodified v24.20.0: a child process crashed with 3221225477 (0xC0000005) at iteration 21, and the run stopped at the first crash.

    Both builds used the same build configuration and pinned fixture, with default isolation, --test-concurrency=1, and spec/JUnit reporters.

    These results suggest that #65778 addresses this crash on v24 as well.

  10. added
    windowsIssues and PRs related to the Windows platform.
    v24.xIssues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch.
    v26.xIssues that can be reproduced on v26.x or PRs targeting the v26.x-staging branch.
    on Sep 13, 2026
  11. inoway46 commented on Sep 13, 2026

    @inoway46
    Contributor

    I've asked in #65778 (comment) whether this fix can be backported to v26.x and v24.x.

  12. dirkwa commented on Sep 13, 2026

    @dirkwa
    Author

    Awesome, thanks a lot for your effort

  13. Romiiansyah commented on Sep 19, 2026

    @Romiiansyah
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    v24.xIssues that can be reproduced on v24.x or PRs targeting the v24.x-staging branch.v26.xIssues that can be reproduced on v26.x or PRs targeting the v26.x-staging branch.windowsIssues and PRs related to the Windows platform.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions