Skip to content

feat(runtime): compact with prefix reuse, rolling, and window-derived budgets - #61

Merged
locez merged 14 commits into
mainfrom
feat/compaction-request-budget
Sep 20, 2026
Merged

locez merged 14 commits into
mainfrom
feat/compaction-request-budget

Conversation

@locez

@locez locez commented Sep 18, 2026

Copy link
Copy Markdown
Owner

Summary

Replace the old compaction path with a destination-window budget that can reuse the session prefix cache, fall back to a one-shot rebuild, and roll until the installed body is actually small enough.

  • When the session request still fits, append a tail directive and keep tools, order, and hashes unchanged.
  • When it does not, rebuild from user/assistant history, keep the newest whole tool exchanges that fit, then roll text only as a last resort.
  • Soft prompt guidance is 5% of the destination window (512–8192). Hard acceptance is 10% below 32k so the compaction request still fits, and max(2.5× guidance, 15%) from 32k up, capped at 20480 so a 2M window does not grow a 300k summary.
  • Installation targets fixed input + hard ceiling + 10% raw tail, not half the watermark. Rolling continues until that body, an indivisible tail below the watermark, or twelve passes.
  • A 128k session that rendered 8569 tokens no longer fails against a 3840-token “hard” limit that was actually the soft target.

Test plan

  • cargo fmt --all --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo test -p merry-runtime --offline compaction
  • cargo test -p merry-cli --offline --bin merry -- --skip sandbox::
  • CI on this PR

Checkpoint compaction previously replaced the agent system prompt with a
dedicated compaction prompt, which discarded the session's cacheable prefix
and made every compaction a cold request. It also sized the compaction
output budget from the checkpoint text budget alone, so a reasoning model
could spend the whole ceiling thinking and never write the checkpoint.

Compaction requests now reuse the session's stable prefix item by item and
append the directive and payload as trailing user messages, each in its own
boundary tag (<merry_compaction_instructions>, <merry_compaction_payload>).
Structured output, tool-free shape, and the install transaction are
unchanged, and the request stays outside the agent loop.

Budget and failure handling:
- reserve reasoning tokens from the request input instead of the text
  budget, so the ceiling grows with the window the runtime measured;
- fit every request before sending it: grant what the window allows on the
  first attempt, then shrink the covered window until the request and its
  full reserve fit;
- degrade a truncated candidate by retrying once with a larger reserve and
  a covered window that can host it, instead of repeating the request;
- keep archive-only reduction available when no replacement fits, and
  report the budget failure instead of silently skipping;
- never inherit the primary model's reasoning effort: compaction has its
  own configurable level.

Provider and diagnostics:
- normalize Responses incomplete_details.reason into a typed finish detail
  so truncation reports max_output_tokens or content_filter;
- keep the compaction payload estimator a documented conservative bound and
  cover it with a test against the authoritative measurement.

Verified with cargo fmt --all --check, cargo clippy --all-targets
--all-features -- -D warnings, and cargo test --all.
Review follow-up that removes duplicated code and state introduced while
fitting compaction into the model window:

- parse reasoning-effort config values in one place, shared by provider
  entries, provider defaults, and runtime compaction, so accepted values and
  the diagnostic shape cannot drift;
- rename AutomaticCompactionConfig to CompactionConfig and document that
  enabled/policy drive the hard-watermark path while reasoning_effort applies
  to every compaction request;
- split the compaction *request* coverage cap out of CompactionWindowBudget
  into CompactionCoverageBudget, which answers a different question than the
  budget for the request compaction installs, and use Option instead of a
  u64::MAX sentinel;
- drop the planner's redundant saw_completed_turn state and replace the
  empty-coverage boolean with a named EmptyCoverage meaning;
- disambiguate the two bytes-per-token constants so the estimation ratio and
  the accepted-output byte ceiling cannot be confused.

Verified with cargo fmt --all --check, cargo clippy, and the merry-runtime
unit and integration suites.
Review follow-up on cohesion. The compaction runtime grew into one 930-line
module and the provider step carried ~185 lines of compaction orchestration,
which pushed provider_step.rs to the 1000-line design signal.

- `runtime/auto_compaction/` now splits by responsibility: the module root owns
  the shared request types and session preparation, and `prefix`, `fit`, `plan`,
  `generate`, `install`, `manual`, and `phase` each own one part of the job;
- the automatic hard-watermark path moves into `phase`, so the provider step
  builds one description of the step, calls the phase, and only recompiles its
  request afterwards. provider_step.rs drops from 969 to 775 lines;
- the window-fit arithmetic moves next to the reserve it serves in
  `compaction.rs`, so all compaction budget math lives in one file, and its
  tests move with it into the existing budget test module;
- the fitter no longer repeats the window invariant: `validate_compaction_model_window`
  is the single gate for whether a request may be sent, and the fitter only
  chooses which covered window to try.

Behaviour is unchanged: the full workspace suite passes (58 suites, 2171 tests)
alongside cargo fmt --all --check and clippy with -D warnings.
A real 272k-token session looped between two failures instead of reducing
context: the window could not host the full request, the refit gave up 1.25x
the excess input and collapsed coverage to zero, the planner then degraded to
archive-only (which never replaces the checkpoint), the body stayed at ~236k,
and compaction re-ran every couple of minutes. The 1m-window run shows the
same shape at a larger scale.

Two defects, both fixed:

- the refit released a multiple of the excess input, which overshoots the
  allowance and can take all the coverage with it. It now releases the excess
  plus a small safety margin, and makes real progress when the measured input
  already fits but the request failed on its output side;
- the reasoning reserve had no floor, so a small request received a ceiling the
  model could not finish in. Observed attempts truncated at 34,022 and 44,337
  token ceilings for 49,051 and 90,308 token inputs, while 59,624 and 66,956
  finished larger requests, so the reserve now floors at the smaller of a share
  of the window and a multiple of the checkpoint text budget, and a truncated
  retry raises the floor as well as the share.

For the session that exposed it, the first attempt now covers 134,641 payload
tokens instead of 25,224 (5.3x) and asks for 81,600 output tokens instead of
34,022 (2.4x), which fits the window; the degraded retry asks for 141,440.
The output ceiling is also bounded by the model's declared output limit so a
large window cannot ask for output the provider would reject.

Regression tests pin both defects: the coverage-collapse case, the small-input
ceiling, and the solver's agreement with the ceiling it promises. All three
fail against the previous arithmetic.

Verified with cargo fmt --all --check, cargo clippy --all-targets
--all-features -- -D warnings, and cargo test --all (58 suites, 2174 tests).
The directive told the model what to preserve and said not to limit the entry
count, but it never said the checkpoint is a compression. A real checkpoint came
out at 21,158 of its 21,760 allowed tokens across 161 entries, including session
metadata, tool-call counts, and per-failure retellings: an execution record, not
a summary, and exactly the ledger material the design keeps out of a checkpoint.

The directive now states the goal and the discipline:

- it is a compression task, and the checkpoint must end up far shorter than the
  turns it replaces, keeping only what changes what a later turn would do;
- merge related facts into one entry instead of one entry per turn, file, or
  tool call, and drop commands, call counts, session metadata, file listings,
  and step-by-step execution;
- keep entries short, aim well below the output limit, and treat that limit as a
  safety ceiling rather than a target to fill;
- preserve a literal value exactly only when later work depends on it.

The design constraints stay: no fixed entry count, sections may be empty, exact
literals where they matter, one ref minimum per entry, ref values only from
available_ref_ids, the eight sections plus handoffs, and the payload treated as
data.

One provider-boundary fixture declared no compaction window, so it inherited the
deliberately tiny primary window and no longer had room for the request once the
directive grew. It now declares the window the design requires of a compaction
model, which is what let it stay a compaction test rather than a budget test.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2174 tests).
The directive now leads with the mission, then keeps, drops, handoff, and
citation contracts as numbered sections. The compression goal is explicit:
compress all covered history into a dense checkpoint, keep only what future
turns strictly need, merge facts that share a decision or boundary, aim for one
sentence per entry, and treat a section as empty when nothing survives it.

The drop list is correspondingly explicit, because a real checkpoint came out at
21,158 of its 21,760 allowed tokens across 161 entries that included session
metadata, tool-call counts, and per-failure retellings.

Every claim in the new text was checked against the runtime before adopting it:
the candidate schema admits only keep and replace handoffs, entry refs carry
length(min = 1), rationale/new_ids/reason are required but nullable, refs are
validated against the payload's model-supplied ids, and empty section arrays are
accepted.

Tests now assert those contracts instead of long verbatim sentences, so the
directive can be reworded without losing the guards that matter: refs only from
available_ref_ids, never an empty refs array, no drop handoffs, the payload as
data, the compression goal, the noise list, and the no-tools no-reply rules.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2174 tests).
Compaction could not run on a session whose history holds archived tool
results. The request shows a short artifact notice for those results, but the
compaction payload sent the archived body instead, so the payload grew far
beyond the request body it was meant to summarize:

  session aaf69cd7, 272k window
  880 archived results : 1,977,060 bytes = 494,265 tokens, sent in full
  request body         : 360,847 tokens
  compaction payload   : 916,710 tokens

With that payload no retention choice could host the compaction request: a
payload small enough for the compaction window required retaining more raw
history than the hard watermark allows, so every candidate failed and the step
ended with "no compaction window fits the compaction request budget". The
compactor read cold storage the conversation had already replaced, which also
contradicted the runtime's own projection: projected_token_estimate already
counts those results as notices.

The payload now carries the same notice the request shows, and the planner's
per-item estimate follows it, so the estimate and the built payload agree. The
notice keeps the status, the artifact id, and the ref, so a checkpoint entry can
still cite the ref and read the body on demand, and the full transcript keeps
the exact content.

For the session above the payload drops from 916,710 to about 448,845 tokens, so
the fitter covers about 42% of the history, retains the rest, and lands at
~219k against a 237k hard watermark instead of failing.

This reverses behavior main deliberately asserted: that the notice is
provider-only and the compactor reads the exact content. That rule cannot hold
when archived bodies dwarf the window, because it makes compaction impossible
exactly when the session needs it most. The test that pinned it now asserts the
new contract, with a larger archived body so a regression is unambiguous.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2174 tests).
Shrinking the context window reported "no compaction window fits the compaction
request budget" and ended the step instead of compacting. History collected under
a wide window no longer fits the narrowed one, and one reduction can only cover
what the compaction request can host, so the planner rejected every candidate:
the retained history alone exceeded the body budget, and the planner treated that
as "nothing fits" rather than as "one pass is not enough".

Compaction now rolls. Each pass covers as much history as the compaction window
can host, and the provider step repeats passes until the recompiled request lands
back under the hard watermark, bounded at four passes with a diagnostic that
reports the count and that the retained history does not fit.

- `RetainedFit` names how strictly a pass must land inside the body budget:
  `Required` keeps manual compaction's promise, and `Deferred` lets a rolling
  pass install the largest covered window and leave the budget to the next pass.
- `build_rolling_compaction_preparation` expresses that for the automatic path;
  manual compaction and the existing builder keep the required semantics.
- the provider step runs the reduction loop and emits one compaction lifecycle
  pair per pass, so a rolling reduction is visible as more than one compaction.

Covered by a test that reproduces the reported failure: history is seeded under a
1M window, the window shrinks to 64k, and the step must complete after more than
one reduction. Against the previous behavior the same test fails with the exact
reported diagnostic, "no compaction window fits the compaction request budget".

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2176 tests).
The bound of four assumed a shrink that needs two or three passes. That
arithmetic does not hold: how much a later pass covers changes with the
checkpoint and the covered range, so the passes a given shrink needs are only
known as they run. At a 272k window a pass covers roughly 140k to 160k tokens of
history, so a session that ran near a 1M window needs several passes and a wider
window reduced further needs more; four would have failed those sessions with the
same "no compaction window fits" diagnostic this bound exists to avoid.

Twelve is deliberately generous and keeps its only real job: stopping a history
that cannot be reduced at all from spending model calls forever.

The rolling test now scripts one candidate per allowed pass, so exhausting the
script fails loudly instead of falling back to a non-candidate response, and it
asserts the pass count stays inside the bound.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2176 tests).
A session keeps its history when its context window shrinks, so the body can sit
far above the new window. Rolling compaction covers one window-worth per pass and
re-summarizes the previous checkpoint every pass, which is what a wide window
reduced to a small one needs several of. The new strategy covers the whole
history in one pass instead, shortening the older covered tool results so the
payload stops growing with the number of tool calls.

The runtime picks between them from the request it is about to build:

  body / window > 1.5  -> one pass over the whole history
  otherwise            -> rolling, so the shared prefix stays cached

The ratio is measured against the window in the current request budget, which
after a shrink is the new window. No record of a previous window is needed:
ordinary turns compact at the hard watermark, so a body this far above the window
means the window moved or earlier compaction did not land.

- `CompactionStrategy` and `CitationCompactionPolicy::strategy_for` own the choice;
  `one_shot_window_percent` (150) and `one_shot_retained_tool_exchanges` (5) are
  configurable, and zero always rolls.
- `CompactionShape` names how one pass covers and shapes its payload, replacing the
  loose retention flags at the session builders: `SinglePass` for manual
  compaction, `Rolling`, and `OneShot`.
- A one-shot payload counts retention per tool exchange, never per item, because a
  call and its result are one pair and the runtime rejects a window that carries
  only one of them. The shortened results keep their status, artifact id, and ref.
- When no one-shot payload fits even with every covered result shortened, the pass
  falls back to rolling rather than failing, and shape changes do not consume the
  coverage-tightening budget.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all. New coverage: the strategy threshold on both
sides, the tunables, the payload pairing and shortening, and an end-to-end case
where a window shrinks far below a tool-heavy history and the step completes after
exactly one compaction call; that case fails with one-shot disabled.
One-shot compaction could not reduce a real session in one pass, so it fell back
to rolling and left the body a few thousand tokens under the watermark. The
payload was dominated by tool call arguments, which the pass still sent in full:

  session aaf69cd7, covered window, 272k window
  covered call arguments  225,800 tokens   sent in full
  covered result bodies    47,800 tokens   shortened to notices
  covered text             25,200 tokens
  previous checkpoint      53,300 tokens   carried every pass
  input budget with the reserve floor    190,400 tokens

The non-shortenable part alone was 304,300 tokens, so no covered window could
fit, and the observed attempt sent 427,277 input tokens before giving up. The
runtime log for that attempt is what a rebuild after this change should stop
producing: runtime.compaction.one_shot_refit followed by a rolling refit that
covers only a fraction of the history.

A one-shot pass now replaces the arguments of every covered call outside the
retained newest exchanges with a marker, alongside the result notice it already
sent. The tool name, call id, status, artifact id, and ref stay, so the
checkpoint still records what ran and can cite it, and the exact arguments remain
in the transcript for the user and for tooling. Rolling is unchanged: it is the
cheap path that preserves the shared prefix, so it keeps results exact.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all. The payload test now also asserts that only
the newest covered exchange keeps its call arguments.
A one-shot pass now leaves the covered tool exchanges out of the payload instead
of shortening them. Measured on session aaf69cd7 the covered window held 1,324
exchanges: their arguments alone were 225,800 tokens against a 190,400 token
input budget, and even a minimal per-exchange marker cost about 95,000 tokens, so
nothing per exchange could stay. The oldest covered exchanges now contribute their
conversation text only.

The decision lives where the payload is assembled, not in the per-item adapter:
`CompactionHistoryItem::to_compaction_turn_item` stays a pure item conversion with
no retention flag, and `citation_compaction_input_from_history` skips the exchanges
outside the retained newest ones. Rolling is unaffected and still sends exchanges
whole, because it is the cheap path that reuses the provider cache.

Omission applies to tool exchanges only. Covered user and assistant text always
travels, and a tool call and its result are dropped together, because a window that
carried only one of them is rejected as stale and because a ref that leaves the
payload can no longer be cited.

The end-to-end test caught that distinction while this was being written: applying
the filter to every item dropped the conversation text as well, and the checkpoint
was then rejected with "checkpoint entry c1 references unknown ref h0". Its fixture
now uses zero-padded ids so an assertion cannot match one id as a prefix of another.

Verified with cargo fmt --all --check, cargo clippy --all-targets --all-features
-- -D warnings, and cargo test --all (58 suites, 2181 tests).
The previous summary ceiling treated a 128k 3% soft target as the hard
limit, so session aaf69cd7 failed with "8569 tokens, above output limit
3840". Summary acceptance now has an independent hard ceiling from the
destination window: 3% soft (512-8192) and 10% hard (1024-16384, at most
1/8 of tiny windows). Exceeding the soft target is allowed; exceeding the
hard limit repairs.

When the session request fits, compaction appends a tail directive and
keeps tools, order, and hashes unchanged so the provider can reuse its
prefix cache. When it does not, the request is rebuilt from user and
assistant history, keeping the newest whole tool exchanges that fit, then
rolling text only as a last resort.

The retained tail is the largest complete suffix of 5, 4, 3, 2, or 1
turns that fits the destination body budget. Manual compaction and
previews now use that same budget instead of an unbounded window, so a
five-turn default no longer keeps an oversized tail.

Verified with cargo fmt --all --check, cargo clippy --all-targets
--all-features -- -D warnings, cargo test --all --offline (58 suites,
2198 passed, 9 ignored), and the Python SDK ruff/ty/pytest/uv build
checks.
…l budgets

A 128k session rejected an 8569-token checkpoint because the 3% prompt
target was also the hard limit. Soft guidance is now 5% of the destination
window (512-8192). Hard acceptance stays at 10% below 32k so the compaction
request still fits, and is max(2.5x guidance, 15%) from 32k up, capped at
20480 so a 2M window does not grow a 300k summary.

Installation targets fixed input plus that hard ceiling plus a 10% raw tail,
not half the watermark. Rolling continues until that body, an indivisible
tail below the watermark, or twelve passes. Manual compaction uses the same
loop.

@olicesx olicesx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Default-config auto-compact is going to bounce.

This PR’s own live numbers say compaction did not finish below ~59k output (budget.rs: 59,624 / 66,956 ceilings succeeded; 34k / 44k truncated). For the 272k fallback window that is also the TUI default, the new floor therefore asks for 20480 summary + 59840 reasoning = 80320 max_output_tokens.

examples/config.toml still points both primary and compaction at gpt-4.1-mini (documented 32k max output). OpenAI capabilities leave max_output_tokens = None, so the new clamp in compile_fitted_compaction_request is unwrap_or(u64::MAX) and never fires. HTTP 400 is InvalidRequest and is not retryable; auto-compaction only runs at the hard watermark, so that failure aborts the step with the session still over the watermark.

If the API instead silently caps at 32k, reasoning room is ~12k — the case this PR already measured as unable to finish — and the truncation retry asks for even more because the declared cap is None.

The prefix-reuse path the PR leads with is real only when the compaction window is substantially larger than the session request (the cache tests use 256k vs 64k). Same-model / same-window, which is the example, falls through to Payload/OneShot and then hits this ceiling.

The fitter has to know the model’s real output cap here, not only when capabilities.max_output_tokens is set. Until that is wired, default-config auto-compact is going to bounce. The 59k floor is only honest for a compaction model that can actually emit 59k.

Secondary, smaller: compact_context_once returns on ArchiveOnly without installing; the automatic path archives. A tool-heavy tail can survive a manual compact.

limits.window_tokens,
estimated_input_tokens,
)
.min(limits.max_output_tokens.unwrap_or(u64::MAX));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This clamp is a no-op on the default OpenAI path: capabilities leave max_output_tokens as None, so this is u64::MAX.

For the 272k fallback window the reserved ceiling is 20480 + 59840 = 80320. gpt-4.1-mini (the example compaction model) is 32k. HTTP 400 is InvalidRequest and not retryable, and this only runs at the hard watermark.

/// request. Real attempts truncated at 34,022 and 44,337 token ceilings for
/// 49,051 and 90,308 token inputs, while 59,624 and 66,956 token ceilings
/// finished for 151,458 and 180,787 token inputs. No ceiling below roughly 59,000
/// tokens finished, whatever the request size.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These live ceilings (59k–67k succeeded, 34k–44k truncated) are already above gpt-4.1-mini’s 32k cap. The 22% / 272k floor that follows from this is only honest for a model that can emit that much.

@locez
locez merged commit 0e3199f into main Sep 20, 2026
8 checks passed
@locez
locez deleted the feat/compaction-request-budget branch September 20, 2026 01:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants