Skip to content

fs: heap buffer overflow in readFileSync(fd, "utf8") when a pipe returns short reads #66341

Description

@bartech-lab

Version

v26.10.0 (also present in v26.8.0 through v26.10.0 and v24.21.0; not present in v26.7.0 or v22.23.3)

Platform

Linux archpc 7.2.4-arch1-2 #1 SMP PREEMPT_DYNAMIC Tue, 08 Sep 2026 10:22:31 +0000 x86_64 GNU/Linux
Arch Linux, nodejs 26.10.0-1, glibc 2.44

Subsystem

fs

What steps will reproduce the bug?

#!/usr/bin/env bash
# Producer: 24 short writes (~96 KiB, each read back as a short read), then one large write.
node -e '
const chunk = "x".repeat(4000);
let i = 0;
const t = setInterval(() => {
  process.stdout.write(chunk);
  if (++i === 24) { clearInterval(t); process.stdout.write("y".repeat(400000)); }
}, 5);
' | node -e 'console.log(require("fs").readFileSync(0, "utf8").length)'
echo "exit=${PIPESTATUS[1]}"

How often does it reproduce? Is there a required condition?

10 of 10 runs with the script above. The input must be a pipe (or FIFO, or /dev/stdin backed by one) that delivers more than 64 KiB in reads shorter than 8 KiB before the first full 8 KiB read. Regular files are not affected. The encoding does not matter: the input above is pure ASCII.

Real-world trigger: some-test-runner | node -e 'require("fs").readFileSync(0, "utf8")...', where the producer writes in small chunks. It crashed at 7 of 60 runs with a ~420 KB input written in 4093-byte chunks.

What is the expected behavior? Why is that the expected behavior?

Prints 496000 and exits 0, as v26.7.0 does and as readFileSync(0).toString("utf8") does on v26.10.0.

What do you see instead?

malloc(): invalid size (unsorted)
exit=134

SIGABRT from glibc. Symbolized stack (from a real crash, via debuginfod.archlinux.org):

#6  malloc_printerr (str="malloc(): invalid size (unsorted)") at malloc.c:5093
#7  _int_malloc (av=<main_arena>, bytes=427486) at malloc.c:3683
#10 __GI___libc_realloc (oldmem=0x0, bytes=427486) at malloc.c:3198
#11 node::UncheckedRealloc<char16_t> (pointer=0x0, n=<optimized out>) at ../../src/util-inl.h:267
#13 node::MaybeStackBuffer<unsigned short, 256ul>::AllocateSufficientStorage (storage=213743) at ../../src/util-inl.h:544
#14 node::StringBytes::Encode (...) at ../../src/string_bytes.cc:659
#15 node::fs::ReadFileUtf8 (args=...) at ../../src/node_file.cc:3559

The abort happens later than the corruption. The corruption itself is a heap buffer overflow in ReadFileUtf8.

Additional information

Cause, in ReadFileUtf8 (src/node_file.cc#L3511-L3523 at v26.10.0):

  • While big == nullptr, every short read (r < 8192) is appended to result and the loop continues. On a pipe, result can grow past 64 KiB this way.
  • On the first full 8 KiB read, the code allocates big_cap = kMinChunk (64 KiB) and runs memcpy(big, result.data(), result.size()) with no check that result.size() <= big_cap. That write overflows the heap buffer.
  • After that, big_len > big_cap, so big_len == big_cap never matches, and big_cap - big_len in the next uv_buf_init underflows.

A possible fix: size the first allocation from what is already buffered, for example big_cap = std::max(kMinChunk, result.size() + sizeof(buffer)), and make the growth check big_len >= big_cap.

Workaround: fs.readFileSync(0).toString("utf8") avoids the fast path.

Activity

  1. darc-games-gt commented on Oct 1, 2026

    @darc-games-gt

    Independent confirmation from a third party, on a production CI pipeline. Posting because the fix PR (#66343) has no reviewer assigned yet and this seems like the kind of report that helps a maintainer decide it is real.

    Environment

    • Node v24.21.0 (listed as affected above)
    • GitHub Actions, ubuntu-24.04 runner image
    • glibc's allocator aborts the process; the message is emitted by libc, not by Node (verified with strings against libc.so.6: the string is present there, and absent from the node binary)

    Shape of the call

    printf '%s' "$BIG_JSON" | node -e 'JSON.parse(fs.readFileSync(0, "utf8"))'
    

    where $BIG_JSON is the output of turbo run … --dry=json — a few megabytes, well past the 64 KiB threshold described in the analysis. I cannot give an exact byte count here because I could not re-measure it on a current checkout, so I am deliberately not quoting one.

    Symptom, verbatim from the CI log

    malloc(): invalid size (unsorted)
    <script>: line 172: printf: write error: Broken pipe
    

    The Broken pipe line is the writer noticing that the reader died, so it is a consequence, not a second fault.

    Behaviour

    • It broke a required status check and kept it broken until changed. Not a rare flake in our case.
    • Not reproducible locally, which cost the colleague who diagnosed it a lot of time. That is consistent with the analysis: it depends on the pipe delivering short reads, and a local pipe tends to behave differently from one under CI load.

    What fixed it on our side

    Writing the JSON to a temp file and reading it with readFileSync(path) instead of readFileSync(0). Green since, with no recurrence in the three subsequent runs of the same job.

    Worth noting for anyone arriving here from a search: the file-based read does not take a different code path in ReadFileUtf8 — I checked the source at tag v24.21.0, and uv_fs_fstat is called inside the same loop and only informs capacity. It avoids the crash because a regular file returns full 8 KiB reads, so the continue branch that accumulates without a bound is never taken, and result still fits in kMinChunk when the heap buffer is allocated. So the workaround removes the trigger, not the defect.

    Thanks to both of you for the analysis and the patch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions