Repository navigation
feat(runtime): compact with prefix reuse, rolling, and window-derived budgets - #61
Conversation
Checkpoint compaction previously replaced the agent system prompt with a dedicated compaction prompt, which discarded the session's cacheable prefix and made every compaction a cold request. It also sized the compaction output budget from the checkpoint text budget alone, so a reasoning model could spend the whole ceiling thinking and never write the checkpoint. Compaction requests now reuse the session's stable prefix item by item and append the directive and payload as trailing user messages, each in its own boundary tag (<merry_compaction_instructions>, <merry_compaction_payload>). Structured output, tool-free shape, and the install transaction are unchanged, and the request stays outside the agent loop. Budget and failure handling: - reserve reasoning tokens from the request input instead of the text budget, so the ceiling grows with the window the runtime measured; - fit every request before sending it: grant what the window allows on the first attempt, then shrink the covered window until the request and its full reserve fit; - degrade a truncated candidate by retrying once with a larger reserve and a covered window that can host it, instead of repeating the request; - keep archive-only reduction available when no replacement fits, and report the budget failure instead of silently skipping; - never inherit the primary model's reasoning effort: compaction has its own configurable level. Provider and diagnostics: - normalize Responses incomplete_details.reason into a typed finish detail so truncation reports max_output_tokens or content_filter; - keep the compaction payload estimator a documented conservative bound and cover it with a test against the authoritative measurement. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all.
Review follow-up that removes duplicated code and state introduced while fitting compaction into the model window: - parse reasoning-effort config values in one place, shared by provider entries, provider defaults, and runtime compaction, so accepted values and the diagnostic shape cannot drift; - rename AutomaticCompactionConfig to CompactionConfig and document that enabled/policy drive the hard-watermark path while reasoning_effort applies to every compaction request; - split the compaction *request* coverage cap out of CompactionWindowBudget into CompactionCoverageBudget, which answers a different question than the budget for the request compaction installs, and use Option instead of a u64::MAX sentinel; - drop the planner's redundant saw_completed_turn state and replace the empty-coverage boolean with a named EmptyCoverage meaning; - disambiguate the two bytes-per-token constants so the estimation ratio and the accepted-output byte ceiling cannot be confused. Verified with cargo fmt --all --check, cargo clippy, and the merry-runtime unit and integration suites.
Review follow-up on cohesion. The compaction runtime grew into one 930-line module and the provider step carried ~185 lines of compaction orchestration, which pushed provider_step.rs to the 1000-line design signal. - `runtime/auto_compaction/` now splits by responsibility: the module root owns the shared request types and session preparation, and `prefix`, `fit`, `plan`, `generate`, `install`, `manual`, and `phase` each own one part of the job; - the automatic hard-watermark path moves into `phase`, so the provider step builds one description of the step, calls the phase, and only recompiles its request afterwards. provider_step.rs drops from 969 to 775 lines; - the window-fit arithmetic moves next to the reserve it serves in `compaction.rs`, so all compaction budget math lives in one file, and its tests move with it into the existing budget test module; - the fitter no longer repeats the window invariant: `validate_compaction_model_window` is the single gate for whether a request may be sent, and the fitter only chooses which covered window to try. Behaviour is unchanged: the full workspace suite passes (58 suites, 2171 tests) alongside cargo fmt --all --check and clippy with -D warnings.
A real 272k-token session looped between two failures instead of reducing context: the window could not host the full request, the refit gave up 1.25x the excess input and collapsed coverage to zero, the planner then degraded to archive-only (which never replaces the checkpoint), the body stayed at ~236k, and compaction re-ran every couple of minutes. The 1m-window run shows the same shape at a larger scale. Two defects, both fixed: - the refit released a multiple of the excess input, which overshoots the allowance and can take all the coverage with it. It now releases the excess plus a small safety margin, and makes real progress when the measured input already fits but the request failed on its output side; - the reasoning reserve had no floor, so a small request received a ceiling the model could not finish in. Observed attempts truncated at 34,022 and 44,337 token ceilings for 49,051 and 90,308 token inputs, while 59,624 and 66,956 finished larger requests, so the reserve now floors at the smaller of a share of the window and a multiple of the checkpoint text budget, and a truncated retry raises the floor as well as the share. For the session that exposed it, the first attempt now covers 134,641 payload tokens instead of 25,224 (5.3x) and asks for 81,600 output tokens instead of 34,022 (2.4x), which fits the window; the degraded retry asks for 141,440. The output ceiling is also bounded by the model's declared output limit so a large window cannot ask for output the provider would reject. Regression tests pin both defects: the coverage-collapse case, the small-input ceiling, and the solver's agreement with the ceiling it promises. All three fail against the previous arithmetic. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2174 tests).
The directive told the model what to preserve and said not to limit the entry count, but it never said the checkpoint is a compression. A real checkpoint came out at 21,158 of its 21,760 allowed tokens across 161 entries, including session metadata, tool-call counts, and per-failure retellings: an execution record, not a summary, and exactly the ledger material the design keeps out of a checkpoint. The directive now states the goal and the discipline: - it is a compression task, and the checkpoint must end up far shorter than the turns it replaces, keeping only what changes what a later turn would do; - merge related facts into one entry instead of one entry per turn, file, or tool call, and drop commands, call counts, session metadata, file listings, and step-by-step execution; - keep entries short, aim well below the output limit, and treat that limit as a safety ceiling rather than a target to fill; - preserve a literal value exactly only when later work depends on it. The design constraints stay: no fixed entry count, sections may be empty, exact literals where they matter, one ref minimum per entry, ref values only from available_ref_ids, the eight sections plus handoffs, and the payload treated as data. One provider-boundary fixture declared no compaction window, so it inherited the deliberately tiny primary window and no longer had room for the request once the directive grew. It now declares the window the design requires of a compaction model, which is what let it stay a compaction test rather than a budget test. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2174 tests).
The directive now leads with the mission, then keeps, drops, handoff, and citation contracts as numbered sections. The compression goal is explicit: compress all covered history into a dense checkpoint, keep only what future turns strictly need, merge facts that share a decision or boundary, aim for one sentence per entry, and treat a section as empty when nothing survives it. The drop list is correspondingly explicit, because a real checkpoint came out at 21,158 of its 21,760 allowed tokens across 161 entries that included session metadata, tool-call counts, and per-failure retellings. Every claim in the new text was checked against the runtime before adopting it: the candidate schema admits only keep and replace handoffs, entry refs carry length(min = 1), rationale/new_ids/reason are required but nullable, refs are validated against the payload's model-supplied ids, and empty section arrays are accepted. Tests now assert those contracts instead of long verbatim sentences, so the directive can be reworded without losing the guards that matter: refs only from available_ref_ids, never an empty refs array, no drop handoffs, the payload as data, the compression goal, the noise list, and the no-tools no-reply rules. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2174 tests).
Compaction could not run on a session whose history holds archived tool results. The request shows a short artifact notice for those results, but the compaction payload sent the archived body instead, so the payload grew far beyond the request body it was meant to summarize: session aaf69cd7, 272k window 880 archived results : 1,977,060 bytes = 494,265 tokens, sent in full request body : 360,847 tokens compaction payload : 916,710 tokens With that payload no retention choice could host the compaction request: a payload small enough for the compaction window required retaining more raw history than the hard watermark allows, so every candidate failed and the step ended with "no compaction window fits the compaction request budget". The compactor read cold storage the conversation had already replaced, which also contradicted the runtime's own projection: projected_token_estimate already counts those results as notices. The payload now carries the same notice the request shows, and the planner's per-item estimate follows it, so the estimate and the built payload agree. The notice keeps the status, the artifact id, and the ref, so a checkpoint entry can still cite the ref and read the body on demand, and the full transcript keeps the exact content. For the session above the payload drops from 916,710 to about 448,845 tokens, so the fitter covers about 42% of the history, retains the rest, and lands at ~219k against a 237k hard watermark instead of failing. This reverses behavior main deliberately asserted: that the notice is provider-only and the compactor reads the exact content. That rule cannot hold when archived bodies dwarf the window, because it makes compaction impossible exactly when the session needs it most. The test that pinned it now asserts the new contract, with a larger archived body so a regression is unambiguous. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2174 tests).
Shrinking the context window reported "no compaction window fits the compaction request budget" and ended the step instead of compacting. History collected under a wide window no longer fits the narrowed one, and one reduction can only cover what the compaction request can host, so the planner rejected every candidate: the retained history alone exceeded the body budget, and the planner treated that as "nothing fits" rather than as "one pass is not enough". Compaction now rolls. Each pass covers as much history as the compaction window can host, and the provider step repeats passes until the recompiled request lands back under the hard watermark, bounded at four passes with a diagnostic that reports the count and that the retained history does not fit. - `RetainedFit` names how strictly a pass must land inside the body budget: `Required` keeps manual compaction's promise, and `Deferred` lets a rolling pass install the largest covered window and leave the budget to the next pass. - `build_rolling_compaction_preparation` expresses that for the automatic path; manual compaction and the existing builder keep the required semantics. - the provider step runs the reduction loop and emits one compaction lifecycle pair per pass, so a rolling reduction is visible as more than one compaction. Covered by a test that reproduces the reported failure: history is seeded under a 1M window, the window shrinks to 64k, and the step must complete after more than one reduction. Against the previous behavior the same test fails with the exact reported diagnostic, "no compaction window fits the compaction request budget". Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2176 tests).
The bound of four assumed a shrink that needs two or three passes. That arithmetic does not hold: how much a later pass covers changes with the checkpoint and the covered range, so the passes a given shrink needs are only known as they run. At a 272k window a pass covers roughly 140k to 160k tokens of history, so a session that ran near a 1M window needs several passes and a wider window reduced further needs more; four would have failed those sessions with the same "no compaction window fits" diagnostic this bound exists to avoid. Twelve is deliberately generous and keeps its only real job: stopping a history that cannot be reduced at all from spending model calls forever. The rolling test now scripts one candidate per allowed pass, so exhausting the script fails loudly instead of falling back to a non-candidate response, and it asserts the pass count stays inside the bound. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2176 tests).
A session keeps its history when its context window shrinks, so the body can sit far above the new window. Rolling compaction covers one window-worth per pass and re-summarizes the previous checkpoint every pass, which is what a wide window reduced to a small one needs several of. The new strategy covers the whole history in one pass instead, shortening the older covered tool results so the payload stops growing with the number of tool calls. The runtime picks between them from the request it is about to build: body / window > 1.5 -> one pass over the whole history otherwise -> rolling, so the shared prefix stays cached The ratio is measured against the window in the current request budget, which after a shrink is the new window. No record of a previous window is needed: ordinary turns compact at the hard watermark, so a body this far above the window means the window moved or earlier compaction did not land. - `CompactionStrategy` and `CitationCompactionPolicy::strategy_for` own the choice; `one_shot_window_percent` (150) and `one_shot_retained_tool_exchanges` (5) are configurable, and zero always rolls. - `CompactionShape` names how one pass covers and shapes its payload, replacing the loose retention flags at the session builders: `SinglePass` for manual compaction, `Rolling`, and `OneShot`. - A one-shot payload counts retention per tool exchange, never per item, because a call and its result are one pair and the runtime rejects a window that carries only one of them. The shortened results keep their status, artifact id, and ref. - When no one-shot payload fits even with every covered result shortened, the pass falls back to rolling rather than failing, and shape changes do not consume the coverage-tightening budget. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all. New coverage: the strategy threshold on both sides, the tunables, the payload pairing and shortening, and an end-to-end case where a window shrinks far below a tool-heavy history and the step completes after exactly one compaction call; that case fails with one-shot disabled.
One-shot compaction could not reduce a real session in one pass, so it fell back to rolling and left the body a few thousand tokens under the watermark. The payload was dominated by tool call arguments, which the pass still sent in full: session aaf69cd7, covered window, 272k window covered call arguments 225,800 tokens sent in full covered result bodies 47,800 tokens shortened to notices covered text 25,200 tokens previous checkpoint 53,300 tokens carried every pass input budget with the reserve floor 190,400 tokens The non-shortenable part alone was 304,300 tokens, so no covered window could fit, and the observed attempt sent 427,277 input tokens before giving up. The runtime log for that attempt is what a rebuild after this change should stop producing: runtime.compaction.one_shot_refit followed by a rolling refit that covers only a fraction of the history. A one-shot pass now replaces the arguments of every covered call outside the retained newest exchanges with a marker, alongside the result notice it already sent. The tool name, call id, status, artifact id, and ref stay, so the checkpoint still records what ran and can cite it, and the exact arguments remain in the transcript for the user and for tooling. Rolling is unchanged: it is the cheap path that preserves the shared prefix, so it keeps results exact. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all. The payload test now also asserts that only the newest covered exchange keeps its call arguments.
A one-shot pass now leaves the covered tool exchanges out of the payload instead of shortening them. Measured on session aaf69cd7 the covered window held 1,324 exchanges: their arguments alone were 225,800 tokens against a 190,400 token input budget, and even a minimal per-exchange marker cost about 95,000 tokens, so nothing per exchange could stay. The oldest covered exchanges now contribute their conversation text only. The decision lives where the payload is assembled, not in the per-item adapter: `CompactionHistoryItem::to_compaction_turn_item` stays a pure item conversion with no retention flag, and `citation_compaction_input_from_history` skips the exchanges outside the retained newest ones. Rolling is unaffected and still sends exchanges whole, because it is the cheap path that reuses the provider cache. Omission applies to tool exchanges only. Covered user and assistant text always travels, and a tool call and its result are dropped together, because a window that carried only one of them is rejected as stale and because a ref that leaves the payload can no longer be cited. The end-to-end test caught that distinction while this was being written: applying the filter to every item dropped the conversation text as well, and the checkpoint was then rejected with "checkpoint entry c1 references unknown ref h0". Its fixture now uses zero-padded ids so an assertion cannot match one id as a prefix of another. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2181 tests).
The previous summary ceiling treated a 128k 3% soft target as the hard limit, so session aaf69cd7 failed with "8569 tokens, above output limit 3840". Summary acceptance now has an independent hard ceiling from the destination window: 3% soft (512-8192) and 10% hard (1024-16384, at most 1/8 of tiny windows). Exceeding the soft target is allowed; exceeding the hard limit repairs. When the session request fits, compaction appends a tail directive and keeps tools, order, and hashes unchanged so the provider can reuse its prefix cache. When it does not, the request is rebuilt from user and assistant history, keeping the newest whole tool exchanges that fit, then rolling text only as a last resort. The retained tail is the largest complete suffix of 5, 4, 3, 2, or 1 turns that fits the destination body budget. Manual compaction and previews now use that same budget instead of an unbounded window, so a five-turn default no longer keeps an oversized tail. Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, cargo test --all --offline (58 suites, 2198 passed, 9 ignored), and the Python SDK ruff/ty/pytest/uv build checks.
…l budgets A 128k session rejected an 8569-token checkpoint because the 3% prompt target was also the hard limit. Soft guidance is now 5% of the destination window (512-8192). Hard acceptance stays at 10% below 32k so the compaction request still fits, and is max(2.5x guidance, 15%) from 32k up, capped at 20480 so a 2M window does not grow a 300k summary. Installation targets fixed input plus that hard ceiling plus a 10% raw tail, not half the watermark. Rolling continues until that body, an indivisible tail below the watermark, or twelve passes. Manual compaction uses the same loop.
olicesx
left a comment
There was a problem hiding this comment.
Default-config auto-compact is going to bounce.
This PR’s own live numbers say compaction did not finish below ~59k output (budget.rs: 59,624 / 66,956 ceilings succeeded; 34k / 44k truncated). For the 272k fallback window that is also the TUI default, the new floor therefore asks for 20480 summary + 59840 reasoning = 80320 max_output_tokens.
examples/config.toml still points both primary and compaction at gpt-4.1-mini (documented 32k max output). OpenAI capabilities leave max_output_tokens = None, so the new clamp in compile_fitted_compaction_request is unwrap_or(u64::MAX) and never fires. HTTP 400 is InvalidRequest and is not retryable; auto-compaction only runs at the hard watermark, so that failure aborts the step with the session still over the watermark.
If the API instead silently caps at 32k, reasoning room is ~12k — the case this PR already measured as unable to finish — and the truncation retry asks for even more because the declared cap is None.
The prefix-reuse path the PR leads with is real only when the compaction window is substantially larger than the session request (the cache tests use 256k vs 64k). Same-model / same-window, which is the example, falls through to Payload/OneShot and then hits this ceiling.
The fitter has to know the model’s real output cap here, not only when capabilities.max_output_tokens is set. Until that is wired, default-config auto-compact is going to bounce. The 59k floor is only honest for a compaction model that can actually emit 59k.
Secondary, smaller: compact_context_once returns on ArchiveOnly without installing; the automatic path archives. A tool-heavy tail can survive a manual compact.
| limits.window_tokens, | ||
| estimated_input_tokens, | ||
| ) | ||
| .min(limits.max_output_tokens.unwrap_or(u64::MAX)); |
There was a problem hiding this comment.
This clamp is a no-op on the default OpenAI path: capabilities leave max_output_tokens as None, so this is u64::MAX.
For the 272k fallback window the reserved ceiling is 20480 + 59840 = 80320. gpt-4.1-mini (the example compaction model) is 32k. HTTP 400 is InvalidRequest and not retryable, and this only runs at the hard watermark.
| /// request. Real attempts truncated at 34,022 and 44,337 token ceilings for | ||
| /// 49,051 and 90,308 token inputs, while 59,624 and 66,956 token ceilings | ||
| /// finished for 151,458 and 180,787 token inputs. No ceiling below roughly 59,000 | ||
| /// tokens finished, whatever the request size. |
There was a problem hiding this comment.
These live ceilings (59k–67k succeeded, 34k–44k truncated) are already above gpt-4.1-mini’s 32k cap. The 22% / 272k floor that follows from this is only honest for a model that can emit that much.
Summary
Replace the old compaction path with a destination-window budget that can reuse the session prefix cache, fall back to a one-shot rebuild, and roll until the installed body is actually small enough.
max(2.5× guidance, 15%)from 32k up, capped at 20480 so a 2M window does not grow a 300k summary.Test plan
cargo fmt --all --checkcargo clippy --all-targets --all-features -- -D warningscargo test -p merry-runtime --offline compactioncargo test -p merry-cli --offline --bin merry -- --skip sandbox::