Skip to content

The grain is extended and rebased without reading it whole - #189

Merged
torstei merged 2 commits into
mainfrom
bugfix/grain-bounded-walk
Oct 2, 2026
Merged

torstei merged 2 commits into
mainfrom
bugfix/grain-bounded-walk

Conversation

@torstei

@torstei torstei commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Problem

extend_grain and rebase_head read the entire .grain on every call to learn how many records it holds and where the last one ends; append_page read and rewrote it per received page; and extend_grain also parsed all of .rings. The .grain states neither number, so every write rediscovered them.

The writer's maintenance tick calls extend_grain whenever the flushed-chunk count changes, i.e. on every chunk flush. A store with a ~130 MB grain therefore allocated ~130 MB, transiently, every few seconds (peak RSS ~195 MB), and a host running many such stores sees the sum.

Change

.grain.commit beside the grain records count and end after each completed write. It is bound to the grain's inode, its 16 header bytes and the 16 bytes at its first record, so a grain replaced or head-dropped by anyone else (an older binary during a downgrade, say) no longer matches and is walked once instead of trusted. Anything past the committed end is a torn write and is cut before the next append.

The commit is renamed into place without an fsync: a lost or older one describes a prefix of the file, which costs re-deriving the difference and never correctness.

  • extend_grain: seek to the committed end and append. No commit means one buffered forward walk, then a commit.
  • rebase_head: walk only the k dropped records; carry the commit through both the COLLAPSE_RANGE and the rewrite path.
  • append_page: append in place.
  • read_index_tail: the rings are a fixed stride after a declared header, so the chunk count is the file length and the tail is one read. Header validation is shared with parse_index_versioned via index_layout.

No format change: the grain's bytes are identical, so no new magic and older binaries read it as before.

Measured

A store with an 89 MB grain, one chunk appended:

peak RSS
before 96 MB
after, cold (no commit yet) 9 MB
after, warm 9 MB

Tests

Extending in steps equals a one-shot build; a torn tail is cut with and without a commit; a commit for another file is not believed; a lagging commit is repaired; the commit survives rebase_head on both the collapse (ext4) and rewrite (tmpfs) paths; adopted pages build a byte-identical grain; the rings tail equals the full parse from that point, including v1 numbering, a partial trailing record and refused headers. Mutating the code to ignore the commit fails the_commit_governs_where_extending_resumes.

Not in this PR

  • grain::load still reads the whole file and copies every filter, so a --has query holds about twice the grain.
  • A query cannot jump to the filters for a time window; it scans them all.
  • The VM tests have not been run locally.

🤖 Generated with Claude Code

torstei and others added 2 commits October 2, 2026 22:13
extend_grain and rebase_head read the entire .grain on every call just to
learn how many records it held and where the last one ended; append_page
read and rewrote it per received page. At a few hundred MB that was a
transient allocation of the same size on every chunk flush.

A .grain.commit beside the grain now records both numbers after each
completed write, bound to the grain's inode and the bytes at its two ends,
so a grain replaced or rebased by anyone else is walked once instead of
trusted. Whatever lies past the committed end is a torn write and is cut
before the next append. The commit is renamed into place without an fsync:
a lost or older one describes a prefix of the file and costs re-deriving
the difference.

extend_grain: seek to the committed end and append.
rebase_head: walk only the k dropped records, carry the commit through
both the COLLAPSE_RANGE and the rewrite path.
append_page: append in place.

.rings is still read whole by extend_grain, and grain::load still copies
every filter; neither is changed here.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
extend_grain parsed all of .rings on every call to learn the chunk count
and take the records past what the grain covers. The records are a fixed
stride after a declared header, so the count is the file's length and the
tail is a single read.

read_index_tail returns both. Its header validation is the one
parse_index_versioned already did, factored into index_layout so the two
cannot disagree about what a rings file is.

On a store with an 89 MB grain, appending one chunk peaked at 96 MB
before this series and 9 MB after.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@torstei
torstei merged commit a169879 into main Oct 2, 2026
7 checks passed
@torstei
torstei deleted the bugfix/grain-bounded-walk branch October 2, 2026 20:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant