Skip to content

Reclaim log-stream bytes nothing cites any more - #1993

Merged
ppXD merged 4 commits into
mainfrom
feat/a-sealed-result-pins-what-it-cites-and-a-reaper-collects-the-rest
Sep 19, 2026
Merged

ppXD merged 4 commits into
mainfrom
feat/a-sealed-result-pins-what-it-cites-and-a-reaper-collects-the-rest

Conversation

@ppXD

@ppXD ppXD commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • A sealed paired-qualification result now records what it cites, in its own transaction: paired_qualification_result_pin (migration 0235) holds one row per agent run, log stream, cleanup receipt and offloaded event payload behind the observations it was computed from. 0235 backfills the same four closures for results sealed earlier, so the table is complete rather than complete-going-forward.
  • pinned_artifact_id joins ArtifactReferenceOracle.ReferenceSites, so the existing artifact reaper can never collect an artifact a sealed result pins.
  • New retention plane for durable records that are not artifacts: a committed rule table (DurableRetentionPolicy), one pure decision (DurableRetentionDecision.Decide — age floor, then an independent quarantine, and every answer but a definite "nothing cites this" keeps), and a generic claim/decide/act loop (DurableRetentionReaper) over per-plane cursors. Hourly at :45 via ReapExpiredDurableRecordsCommand (INonTransactionalCommand). One class is registered, because one cursor exists.
  • One sweep reaches every eligible record, not just the per-tenant head: the fair claim returns one row per team, so the loop asks again until the batch is full (ids already settled this tick are excluded, so a drain still in progress cannot spin).
  • The only cursor is Retention/Cursors/LogStreamRetentionCursor: for a terminal stream past 30 days that no pin cites and whose objects no other stream or workflow_artifact row also names, it purges the routed CAS bytes first and then stamps agent_run_log_stream.purged_at. The head row deliberately outlives its bytes — the read path answers Purged with its own code and the Room says purged, rather than either of them reporting a policy reclamation as data loss.
  • Every decision is settled, including the keeps. A record nothing can ever reclaim would otherwise be re-claimed by the oldest-first query on every tick, and a few hundred of them deployment-wide would fill the batch for ever while the sweep reported a healthy claim count.
  • Migration 0236 adds the one guard arm that makes the two new columns writable at all: agent_run_log_stream_guard() refuses every update to a terminal row, so a retention statement is admitted explicitly — those two columns only, revision advancing, and a purge that has no previously recorded retain_until or that tries to be undone is refused. A statement that touches anything else reads the refusal it always did, word for word.

Test plan

  • Unit: full suite 10828 passed / 1 skipped / 0 failed. New: the rule table pinned to its literal windows; the citation-site list pinned plus a drift detector over the EF model that reds on any future *artifact_object_id column; Pinned_row_is_never_collected, First_unreferenced_observation_only_quarantines, Young_row_keeps_whatever_its_verdict, A_citation_clears_the_quarantine_it_contradicts. Room fold: a purged stream is not folded as integrity-verified.
  • Integration, real Postgres + real local-rwx destination: 579 passed / 0 failed across Retention|AgentRunLog|Qualification|Room|Reconcil|Sweep|ArtifactLocationVerifier in one invocation (23 of them new).
  • Mutations run, not just claimed — each named test goes red and green again: pin probe dropped; sibling-segment probe dropped; keeps not settled; purge tombstone unchecked on the read path; an empty reclaimable set read as drained; a partial drain deferred like a refusal; the guard reading NEW.retain_until instead of OLD; a purge by object id rather than placement; the claim not looped; a listed citation site with no probe beside it.
  • Frontend: pnpm build and pnpm lint clean (0 errors), vitest 193 passed. "Purged" is in the RoomAgentLogStatus union AND in the log reader's AgentRunLogReadAvailability union + its runtime Set (without both, the backend's 410 decoded as InvalidResponse), purgedAt rides on the stream summary, and the reader renders its own words for it.
  • dotnet build backend/CodeSpace.sln — 0 errors, after rebase onto main.

Deploy notes

  • Rolling-safe. agent_run_log_stream.retain_until / purged_at are nullable ADD COLUMNs with no default and no index — metadata-only, no table rewrite — and NULL is exactly what an older binary writes and reads. paired_qualification_result_pin is a new table nothing older touches.
  • 0235's backfill runs once inside the migration transaction; it is bounded by the number of sealed results already present.
  • 0236 is a CREATE OR REPLACE of agent_run_log_stream_guard() at its own number (the function has been redefined at 0129, 0132, 0133, 0201, 0230, 0232 — one number each, which MigrationDiscoveryTests.No_two_migrations_at_one_number_redefine_the_same_function enforces). It only widens: every update the previous body admitted is still admitted, with its message unchanged.
  • Nothing is reclaimed for at least 30 days plus a further day of quarantine after this ships, so the first sweep that can delete anything is well after a rollback window.
  • The drain reclaims PLACEMENTS, not objects: two byte-identical writes deduplicate to one content-addressed object with a placement each (the object key is chosen per write, so both can sit at the same destination), and each is named in its own purge request. A placement the destination will not give up — a Corrupt one, which is claimable but never deletable — is a keep, logged with the refusal's own code every sweep.

Follow-ups

  • The other planes that still have no rule and no reaper: cleanup receipts (0229, Orphaned never until compensated), capture gaps (0231, kept with their stream), qualification evidence (0218–0225), reconciled budget reservations (0104). Each gets its rule in the same change as its cursor.
  • artifact_transfer_intent needs a schema change first, not just a cursor: a terminal intent still holding temporary_object_key is permanently unreachable, because ArtifactCasRuntimeCoordinator.Resume.cs:~187-190 gates the clearing write on a live lease that a terminal row never has. Relaxing artifact_cas_transfer_guard() (defined at 0127, 0128, 0131, 0226) is its own migration and its own negative test.
  • Row-level GC of segments and verification manifests is deliberately not here: agent_run_log_segment_guard() refuses every non-INSERT and agent_run_log_verification refuses DELETE, both pinned by a test in this PR. The bytes — which is what actually accumulates — are reclaimed without them.
  • The claim query has no supporting index (a partial index on agent_run_log_stream (completed_at, id) WHERE purged_at IS NULL AND state <> 'Open' wants CREATE INDEX CONCURRENTLY, which cannot appear in a transactional DbUp script).

The artifact plane was the only one with a retention rule and a reaper. An
agent run's log streams, their segments and the routed CAS objects behind
them had neither: once a run ended, its archive sat at its destination for
good, and nothing in the schema could even answer whether something still
needed it.

Two things were missing, and the first is why the second is safe. A sealed
paired-qualification result is a statement ABOUT a set of runs, expressed
nowhere as a foreign key — so "is this stream still cited" was unanswerable,
and an unanswerable question has to mean keep. The pin table answers it, and
is written in the seal's own transaction so a committed seal always has its
pins. Migration 0235 backfills the same four closures for results sealed
before it existed, because a pin table that only fills forward would make
every older result read as citing nothing.

The reaper then reclaims only what both waits and a fail-closed citation
probe have cleared, bytes first and the tombstone last. The head row
deliberately outlives its bytes: a reader that finds nothing cannot tell
"reclaimed by policy" from "capture lost it", so the Room reads Purged
instead of meeting a stream that will not load.
`.gitignore:19` ignores every directory named `Logs/`, so the cursor — the
only file in this change that actually reclaims anything — was silently
skipped by `git add` and never appeared in `git status`. The build and the
tests were green locally because the file was on disk; CI would have failed
to compile, because the integration test references the class.

`Cursors/` is also the name rule 18.3 asks for: the sub-folder is named for
the variant axis, which here is which plane a cursor sweeps, not which
plane the first one happens to be.
A review of the first cut found four ways the plane could look like it was
working while doing nothing, or say something it had not earned.

A keep wrote nothing at all, so a stream nothing can ever reclaim — one a
sealed result pins, or one whose bytes another stream also names — was
re-claimed by the oldest-first query on every tick. A few hundred of those
deployment-wide would have filled the batch for ever while the sweep
reported a healthy claim count. Every decision is now settled, including the
keeps, and the claim query leaves a record alone for its recheck interval.

The read path answered a policy reclamation in the vocabulary of data loss:
the segment walk found no available location and reported ArtifactMissing,
byte-identical to an object that vanished. Worse, the metadata kept
projecting the surviving manifest receipt as an integrity proof over bytes
that were gone. A purged stream now answers with its own code before the
walk, and withholds the proof.

The citer list was neither complete nor pinned. It is enumerated, covered by
a drift detector over the EF model, and it deliberately does NOT include a
transfer intent: ck_artifact_transfer_intent_outcome makes that column null
on every non-terminal row, so the only rows naming an object are finished
transfers — and a transfer is what wrote each segment, so probing it would
have kept every log stream in the deployment for ever.

Three smaller reversals of meaning: a stream whose bytes were lost by some
other path is no longer tombstoned as purged; a drain that is still making
progress no longer backs off like a refusal; and the guard now requires the
quarantine to have been recorded by an EARLIER statement, so one sweep
cannot both propose and execute a collection.

The rule table lists one class, because one cursor exists. The five classes
it advertised without one are gone until their cursors arrive.
@ppXD
ppXD force-pushed the feat/a-sealed-result-pins-what-it-cites-and-a-reaper-collects-the-rest branch from 5e6c189 to a69c4d5 Compare September 19, 2026 05:19
… team

Three things the first rework got wrong, each of which quietly kept bytes
nobody wanted.

Deduplication is per WRITE, not per stream: two byte-identical segments are
one content-addressed object with a placement each, and because the object
key is chosen per write both can sit at the same destination. The drain
refused any object with more than one placement, so a capture that repeated
itself was unreclaimable for good — and the suite could not see it, because
every test wrote distinct bytes per segment. The drain now walks placements
and names each one in its purge request, which is what the coordinator asks
for when an object has more than one. The refusal it replaced was also
misreported: "holds bytes at N destinations" was one destination, N keys.

The per-tenant fair claim returns one record per team, so claiming once
capped the sweep at one stream per team per tick — about a dozen a day
against the two a single run writes. The loop asks again until the batch is
full, skipping ids already settled this tick so a drain that is still making
progress cannot spin.

And the log reader's own frontend rejected the new Purged verdict as a
malformed response, which is worse than the wrong-but-well-formed answer it
replaced: a 410 an operator could read became "invalid log response".
@ppXD
ppXD merged commit abd3828 into main Sep 19, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant