Bound per-conversation and per-query cost in the agent (review, area 3) - #314
Merged
Merged
Conversation
- History sent to the model is the recent part: up to 40 messages and 60K characters, plus a handoff's seeded first turn. Nothing was trimmed; threads overflowed the context window (8K-char messages at turn 47, an 8K-emoji message by turn 6) and then failed every turn. - BM25 scores a deduplicated query of at most 64 tokens. rank_bm25 scans every document per query token, repeats included: ~16s of CPU for a 2,000-character question, slowing every session. Documents are scored as before. - OpenAI embeddings have a 30s timeout; with none, a stalled endpoint pinned shared worker threads until every to_thread queued. - The graph compiles once, under a lock: concurrent first requests each opened a Postgres pool. thread_holds_analysis and forget_thread initialize it instead of answering 'no' or doing nothing -- after a restart the web-search guard on analysis threads failed open. Answer sweep 16/16 and the analysis handoff browser test pass on this branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
These fixes come from the max-level review of the agent and retrieval code. Each was reproduced by the reviewer.
chat_historywas never trimmed. Threads overflowed the context window and then failed every turn: 8K-character messages at turn 47, an 8K-emoji message (~24k tokens) by turn 6.agent.history.recent: up to 40 messages and 60K characters, plus a handoff's seeded first turn. Applied wherever history reaches a model.BoundedBM25Retriever: the query is deduplicated and capped at 64 tokens; document scoring is unchanged_ensure_graphunder a lockthread_holds_analysisreturned "no" until the graph was initialized, so after a restart the web-search guard failed open;forget_threadwas a no-op_ensure_graphOn this branch, the answer sweep passes 16/16 and the analysis-handoff browser test passes end to end (the seeded turn survives the trimming).
Not in this PR: Postgres guest threads left behind by a restart need a sweep, which is a deployment decision (production's compose only). Answer quality items follow separately.
🤖 Generated with Claude Code