Skip to content

Reclaim the settled cleanup receipts nothing cites any more - #2001

Merged
ppXD merged 3 commits into
mainfrom
feat/reclaim-the-durable-records-nothing-cites-any-more
Sep 20, 2026
Merged

ppXD merged 3 commits into
mainfrom
feat/reclaim-the-durable-records-nothing-cites-any-more

Conversation

@ppXD

@ppXD ppXD commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • One new cursor on the durable-retention plane, registered together with its class (a class exists only with a cursor): Cursors/CleanupReceiptRetentionCursor, 30 d. A receipt is a ledger row whose whole content is a statement about something that is over, so reclaiming it means DELETING it — there is nothing to purge first, and nothing a tombstone could say that a missing settled receipt does not.
  • Reclaimable outcomes are exactly the two RunCleanupReceipt.IsSettled already names: Completed and Compensated, held against that definition by a test rather than against a second opinion about it. Orphaned is addressed to a sweep on the host that owes the teardown and can still become Compensated. Unknown is neither settled nor beyond reach: the ledger's upsert (IRunCleanupLedger.cs:51-55) fences only the two settled outcomes, so a live orphan a host cannot even attempt to tear down moves Orphaned → Unknown and leaves 0229's orphan queue for good — and the Room COUNTS those rows (RoomProjector.SummarizeRecovery), over every receipt of a run and with no age bound, to say how many resources are in an unknown cleanup state.
  • Citer: paired_qualification_result_pin.pinned_id — the pin table from Reclaim log-stream bytes nothing cites any more #1993. Pinned by a list, cross-checked against the EF model (a future column naming a receipt reds), against ClassifyAsync's own probe body (a listed site with no probe reds), and against the claim's raw SQL (the IN list and the pin kind are one rule written twice, and two copies drift unobserved).
  • Migration 0237: one nullable retain_until on agent_run_cleanup_receipt, plus a partial index matching the claim's DISTINCT ON (team_id) … ORDER BY team_id, recorded_at, id shape. The index is built in-script rather than CONCURRENTLY because DbUp runs each script in one transaction; the SHARE-lock window is stated in the migration, and the table holds a handful of rows per abandoned run.
  • Two repairs to the shared loop. The batch budget is per cursor, not spent in class order — one plane's backlog could otherwise starve every plane behind it for ever while the sweep reported a healthy count. And SettleAsync now receives the same window the claim used, so a deleting statement repeats the time predicates that admitted the row.

Not built, with the evidence

  • Capture gaps — workflow_run_capture_gap_guard() (0146:253-255) refuses DELETE because "a removable gap makes a complete manifest reachable by deleting the evidence"; a gap can only honestly go with the manifest whose verdict it qualifies.
  • Paired-qualification results — paired_qualification_result_immutable (0219:42-44) refuses UPDATE and DELETE, and a result's citers are not enumerable in columns at all: QualificationReceipt carries no column naming a result and records its cohort and metrics as JSON, which is the shape this plane's charter excludes.
  • Budget reservations — the Room's run-scoped readers are UNWINDOWED: RoomProjector.BudgetAsync and TerminalEvidenceAsync select every reservation of a run with no time predicate and derive CommittedUsd, CapUsd, UnbudgetedUsd and UnresolvedClaims from them, which RoomNarrative renders as "$X committed". Reclaiming the oldest rows of a long run would leave that figure quietly wrong rather than absent. It needs a durable per-run summary to outlive the rows first — a money-plane change, not a retention slice. (The 90 d floor would have been three times the only cap window, TeamCostCap.RollingThirtyDaysSpan; that arithmetic was never the problem.)
  • artifact_transfer_intent still needs artifact_cas_transfer_guard() relaxed.

The policy table states all four refusals where a reader will look.

Test plan

  • Unit: full suite 10897 passed / 1 skipped / 0 failed.
  • Integration (real Postgres): 501 passed / 1 skipped / 0 failed across Retention|Budget|Room|Recovery|AgentRunLog in one invocation; 13 of them new. Own-row assertions throughout — no sweep tallies — and the class deletes what it staged in teardown.
  • Mutations run, each named test red then green again: Unknown back in the reclaimable set (both by claim and through the Room's own fold); the pin probe dropped; Orphaned admitted; the delete's repeated time predicates dropped; only Quarantine and Collect settled; the SQL and the allow-list drifted apart; one batch budget shared across cursors.

Deploy notes

  • Rolling-safe. retain_until is a nullable ADD COLUMN with no default — metadata-only, no rewrite — and NULL is what an older binary writes and reads. agent_run_cleanup_receipt carries no trigger, so no guard had to learn about it.
  • Nothing is reclaimed for at least 30 days plus a further day of quarantine after this ships, so the first sweep that can delete anything is well past a rollback window.
  • The cursor rides the existing hourly ReapExpiredDurableRecordsCommand; there is no new job and no new schedule.

Two more planes that had a rule in a design document and no reaper.

A cleanup receipt is a ledger row, so reclaiming it means deleting it —
there is nothing to purge first and nothing a tombstone could say that a
missing settled receipt does not already say. Only the outcomes nothing can
change again are candidates. Orphaned is never one: it names the host that
still owes the teardown, it can still become Compensated, and removing it is
the only way to make a leaked resource invisible. Unknown IS one, because the
only sweep that revisits a receipt selects Orphaned alone — waiting for an
Unknown row to settle waits for ever.

A budget claim is reclaimed only from the four states nothing can change
again, named as an allow-list so a new state is out of scope until someone
puts it there rather than in scope the moment it is not called live. Ninety
days is three times the only window a team cap is ever measured over, which
a test pins against that window rather than leaving to arithmetic in a
reader's head. Two citers keep a claim alive — the physical model-call
receipt that spent under it, and a child claim accounted beneath it — and
both are probed rather than left to the ON DELETE RESTRICT behind them,
because a refusal that arrives as a constraint violation is a sweep-wide
error while one that arrives as a verdict is a single row kept and counted.

Both cursors exclude an uncollectable row in the CLAIM query instead of
settling it, which is why one deadline column each is enough: nothing that
can never be collected occupies a batch slot.

Capture gaps and paired-qualification results are NOT reaped, and the policy
table says so where a reader will look. Each is refused deletion by its own
trigger for a reason that trigger states, and a result's citers are not
enumerable in columns at all — a qualification receipt records its cohort
and its metrics as JSON, which is exactly the shape this plane's charter
excludes.
@ppXD
ppXD force-pushed the feat/reclaim-the-durable-records-nothing-cites-any-more branch from 7d5af2b to 8345243 Compare September 19, 2026 23:34
The structure survived review; each cursor's central safety argument did
not, because both were claims about readers that a reader in the Room
falsifies at the call site.

An Unknown receipt is neither settled nor beyond reach. The ledger's upsert
fences only Completed and Compensated, so a live orphan on a host that
cannot even attempt a teardown moves Orphaned to Unknown and leaves the
orphan queue for good — and the Room COUNTS those rows, over every receipt
of a run and with no age bound, to say how many resources are in an unknown
cleanup state. Draining them per-team-head would have walked that count down
over days and then removed the card, putting a wrong number where the truth
used to be. The reclaimable set is now exactly the two RunCleanupReceipt
already calls settled, and a test reads the Room's own fold before and after
a sweep.

The budget cursor is gone. RoomProjector.BudgetAsync sums EVERY reservation
of a run with no time predicate to state what that run committed, so
reclaiming the oldest rows of a long run leaves that figure quietly wrong
rather than absent. Reclaiming spend records needs a durable per-run summary
to outlive them, which is a money-plane change and not a retention slice; it
is recorded as a refusal beside the capture gap and the qualification
result, with the reader that refutes it named.

Three smaller repairs. The deleting statement now repeats the two time
predicates its comment always claimed it did, so a decision taken about the
row a receipt used to be cannot delete the row it has become. The batch
budget is per cursor rather than shared in class order, where one plane's
backlog could starve every plane behind it for ever while the sweep reported
a healthy count. And a keep the cursor cannot record any other way pushes
its own deadline forward, so an unanswerable citation question stops being
re-asked on every tick.
@ppXD
ppXD force-pushed the feat/reclaim-the-durable-records-nothing-cites-any-more branch from 8345243 to 56e30b5 Compare September 19, 2026 23:36
@ppXD ppXD changed the title Reclaim the settled receipts and terminal budget claims nothing cites Reclaim the settled cleanup receipts nothing cites any more Sep 19, 2026
The rework introduced a way for a receipt to skip the second wait. The
deferral it added applied its recheck deadline to every keep, including a
CITATION — whose decision deliberately carries no deadline, because a
quarantine that elapsed while something still pointed at the row never
waited for anything. A pin landing between a claim and its classification
therefore left a future deadline behind; once the pin went away with its
result, the first uncited observation past that deadline would collect the
row. The log-stream cursor already honoured the null, so the two disagreed
on a shared contract.

It also deleted rather than shrank the guard that a class exists only with
its cursor, and a summary recording that two test files across the backend
and the web client pin the same wire strings. Both are back.

Three riders. The drift test now reads migration 0237 as well, because the
partial index predicate is a third copy of the settled set and an index that
silently stops matching the claim is a sweep that starts seq-scanning. The
method extract stops at the next member of any visibility, not only the next
private one, so a scoped check cannot quietly become a whole-file one — and
it now asserts its own boundary rather than leaving that to a reader.
@ppXD
ppXD merged commit 5f7bda0 into main Sep 20, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant