diff --git a/reference/configuration/options.md b/reference/configuration/options.md
index f128baeb..9a91d793 100644
--- a/reference/configuration/options.md
+++ b/reference/configuration/options.md
@@ -384,6 +384,7 @@ agent:
enabled: true
model: default
maxTurns: 50
+ maxToolResultBytes: 65536
autoApprove: false
allowDestructive: false
user: hdb_agent
@@ -393,6 +394,7 @@ agent:
- `provider` — Recorded on the session but **not yet used to route the model call** — only `model` reaches the provider. Set the provider through the [`models`](../models/overview.md#configuration) configuration instead
- `model` — Model id override, passed through to the model call; _Default_: the [`models`](../models/overview.md#configuration) generative default
- `maxTurns` — Maximum tool-call iterations in a single run; _Default_: `50`
+- `maxToolResultBytes` — Largest tool result, in bytes of JSON, the agent adds to its conversation. A larger result is cut to this size, ending with a note that gives the original size and tells the model how to ask for less. The file-reading tools return pages of at most half this size, measured as JSON, so a page arrives whole. Lower it for a model with a small context window. It must be an integer from `1024` to `1048576`; at startup an invalid value is logged and the default is used. See [When a conversation outgrows the model's context window](../operations-api/operations.md#when-a-conversation-outgrows-the-models-context-window); _Default_: `65536`
- `maxCostUsd` — Intended per-run cost ceiling. **Not enforced** — it is a stored setting only, and nothing checks spend against it; _Default_: `5.00`
- `autoApprove` — Run without per-action approval gates; _Default_: `false`
- `allowDestructive` — Include the tools marked destructive in the agent's toolset: `write_file`, the inspector's code-evaluation tools, and any operations tool carrying MCP's [`destructiveHint`](../mcp/tool-metadata.md) (`drop_table`, `delete`, `restart`, `set_configuration`, ...). When `false` they are removed entirely rather than gated. That hint comes from a curated set in core which does not cover every damaging operation, so this is not a complete safety boundary on its own — see [Agent operations](../operations-api/operations.md#agent); _Default_: `false`
@@ -401,7 +403,7 @@ agent:
- `httpFetch` — Whether the agent has its `http_fetch` tool, and which hosts it may reach: `true`, `false`, or `{ allow: [...] }`. Read at startup only. See [Restricting `http_fetch`](#restricting-http_fetch); _Default_: `true`
- `systemPromptAppend` — Operator text appended to the agent's system prompt
-`enabled`, `provider`, `model`, `maxTurns`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend` can also be changed at runtime with [`set_agent_config`](../operations-api/operations.md#set_agent_config), which applies in memory only. `enabled` is the exception worth knowing: it cannot switch the agent on, because with the agent disabled at startup no agent operation is registered at all.
+`enabled`, `provider`, `model`, `maxTurns`, `maxToolResultBytes`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend` can also be changed at runtime with [`set_agent_config`](../operations-api/operations.md#set_agent_config), which applies in memory only. `enabled` is the exception worth knowing: it cannot switch the agent on, because with the agent disabled at startup no agent operation is registered at all.
### Restricting `http_fetch`
diff --git a/reference/operations-api/operations.md b/reference/operations-api/operations.md
index eb05c42d..1f49d260 100644
--- a/reference/operations-api/operations.md
+++ b/reference/operations-api/operations.md
@@ -1712,6 +1712,17 @@ A session's `status` is one of:
`completed` also covers hitting the `agent.maxTurns` ceiling — in that case `lastError` reads `Reached maxTurns= without a final answer.`, so check it before treating a completed session as finished.
+### When a conversation outgrows the model's context window
+
+
+
+Every model request replays the session's whole conversation, so a single large tool result can fill the model's context window, and then every later request is rejected. Two rules keep a session usable:
+
+- **Tool results are capped where they are stored.** A result larger than [`agent.maxToolResultBytes`](../configuration/options.md#agent) (default 64 KiB) is cut to that size before it is added to `messages`, ending with a note that gives its original size and tells the model how to ask for less. Only the cut form is kept, so `get_agent_session` shows what the model saw. The file-reading tools return pages sized to fit under the cap, measured as JSON: `read_file` returns up to `lineCount` whole lines from `startLine`, and while the file continues it gives `nextLine` and `nextOffset`, which the agent passes back as `startLine` and `offset` to read the next page without rescanning the file. So it can work through a log of any size one page at a time. A line longer than a page comes back in parts, continued the same way, with every part but the last flagged `lineTruncated`.
+- **A rejected request is retried once.** When the provider rejects a request because it does not fit the model's context window (detected for the OpenAI, Anthropic and Bedrock backends), the agent cuts every result over 2 KiB to 2 KiB, in the most recent group of tool results that has one, keeping its beginning plus a note saying it was cut, records the cut in `messages`, and sends the request again. If nothing is left to cut, or the retry is rejected as well, the run ends `error` with a `lastError` that starts `The conversation no longer fits the model's context window`. Shorten the prompt, or start a new session. With fallback models configured, a context-window rejection from a fallback that follows an unrelated failure of the first model is reported as that first failure, and is not retried.
+
+Earlier versions added every tool result in full, so one large result left the session unusable. Prompting such a session again usually recovers it, because its oversized result is cut on the first rejection; if it still ends `error`, start a new session.
+
### `agent_prompt`
Sends a prompt to the agent. Omit `session_id` to start a new session; supply one to continue an existing conversation. Returns immediately with the session id and `"status": "running"`.
@@ -1809,7 +1820,9 @@ One gap is worth knowing: changing `allowDestructive` with [`set_agent_config`](
### `set_agent_config`
-Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only.
+
+
+Updates agent settings and returns the resulting configuration. Accepts any of `enabled`, `provider`, `model`, `maxTurns`, `maxToolResultBytes`, `maxCostUsd`, `autoApprove`, `allowDestructive`, and `systemPromptAppend`; keys not supplied are left unchanged. Each field is described under [`agent`](../configuration/options.md#agent). A request that includes `httpFetch` is rejected with a 400 and nothing in it is applied: the [`http_fetch` policy](../configuration/options.md#restricting-http_fetch) is read at startup only. A `maxToolResultBytes` that is not an integer from `1024` to `1048576` is rejected the same way.
```json
{ "operation": "set_agent_config", "autoApprove": false, "maxTurns": 20 }
@@ -1818,7 +1831,7 @@ Updates agent settings and returns the resulting configuration. Accepts any of `
Three limits are worth knowing:
- **The change is in-memory and not persisted.** It applies for the life of the process and is lost on restart; edit `harper-config.yaml` for a durable change.
-- **A run already in flight keeps the settings it started with** — its toolset, `autoApprove`, `model`, and `systemPromptAppend` are all captured at start. Changes take effect on the next run. To stop a run immediately, use `cancel_agent_run`.
+- **A run already in flight keeps the settings it started with** — its toolset, `autoApprove`, `model`, `maxToolResultBytes`, and `systemPromptAppend` are all captured at start. Changes take effect on the next run. To stop a run immediately, use `cancel_agent_run`.
- **`enabled` is not a kill switch.** It cannot turn the agent on — if it was off at startup, this operation does not exist. Setting it to `false` only makes subsequent `agent_prompt` calls return 409; a run already in flight continues, and `approve_agent_action` still resumes a paused one. Use `cancel_agent_run` to stop a run.
### MCP access
diff --git a/release-notes/v5-lincoln/5.4.md b/release-notes/v5-lincoln/5.4.md
index 32120367..87d0048f 100644
--- a/release-notes/v5-lincoln/5.4.md
+++ b/release-notes/v5-lincoln/5.4.md
@@ -18,6 +18,12 @@ A `subscribe` request can now carry the `databaseGeneration` its position came f
A durable MQTT session now resumes through the same checked replay. When a reconnect's saved position names history that audit retention has removed, Harper deletes the session and reports `sessionPresent: false` (or, if the problem appears after `CONNACK`, sends an MQTT v5 `DISCONNECT` with reason code `0x83` and closes the connection) instead of delivering a partial catch-up. A position saved before a restore or copy, or on another cluster node, resumes best-effort, as before. A quiet topic's position is kept current, so retention on other tables does not reset it, and positions no longer skip unacknowledged messages that share a transaction. Resubscribing to a topic the session already holds continues from its saved position, and QoS 0 subscriptions are kept with the session and resume live. Each connection's Last Will is kept separately, so a connection that is taken over publishes its own will, never the newer connection's. See [Durable Sessions](/reference/v5/mqtt/overview#durable-sessions).
+## Built-in Agent
+
+### Capped and Paged Agent Tool Results
+
+The built-in agent added every tool result to its conversation in full, so one large result, such as a log read whole, could fill the model's context window and leave the session unusable. A tool result is now cut to `agent.maxToolResultBytes` (default `65536`) before it is stored, and `read_file` pages through a file by line, with no size limit, returning a cursor (`nextLine`, `nextOffset`) to continue from. When a provider still rejects a request as too long for the model, the agent cuts the most recent tool results to 2 KiB and retries once, and prompting a session left unusable by an earlier version usually recovers it the same way. The agent also stops sending its system prompt twice per request. See [When a conversation outgrows the model's context window](/reference/v5/operations-api/operations#when-a-conversation-outgrows-the-models-context-window).
+
## Component Deploys
### Certifying a Release in a Canary Worker