Repository navigation
Answer length: the prompt instructs exhaustiveness, and the document budget should be per-caller #172
Description
Activity
Measured: a token budget would be tighter than a document count, but is not currently binding
From triaging #139, which proposes token-aware truncation with limits in
config.yml.With the per-collection document cap now on
main, measured over six real questions on Release95:40 docs 5,796 – 9,531 tokens (1.6x spread)So a fixed document count is a loose proxy — the same 40 documents vary by 64% in tokens. A token budget bounds the thing that actually costs money and context.
But it is not urgent: 9,531 tokens is 7% of gpt-4o-mini's 128k window, down from ~32k before the cap. Adding a second limit now would tighten a bound that is not binding.
tiktokenis already in the lockfile (transitive vialangchain-openai), so this would cost no new dependency when it is wanted.The more valuable idea from #139
Putting the limits in
config.ymlrather than a module constant:retriever: context_truncation: max_docs: 15 max_tokens: 12000
That is the same direction as the per-caller budget in this issue — search-results integration wants fewer documents than chat — and it makes tuning a config change rather than a code change.
It is blocked on the same plumbing as #151 (wiring
AgentGraphto the YAML config):HybridRetrieveris built insidecreate_profile_graphs, which has noConfig. Worth doing all three together rather than piecemeal.Recommendation
Keep the document cap as-is. When the config plumbing lands, move
MAX_DOCUMENTS_PER_COLLECTIONintoconfig.ymlalongside amax_tokensbackstop, per-caller.🤖 Generated with Claude Code
Measured: how long the answers actually are
First end-to-end run of the full RAG chain (question in, answer out) after the document cap landed. Six questions from
tests/golden/questions.txt, Release95, gpt-4o-mini, counted withtiktokenrather than estimated.tokens words citations question 413 ~240 4 What does CDK5 phosphorylate in Alzheimer's... 822 ~470 7 How does TP53 regulate PTEN transcription? 576 ~330 5 What is the role of CDK12 in DNA repair... 780 ~450 10 Which proteins are in the RNA polymerase II... 457 ~260 2 What happens during Golgi fragmentation... 429 ~250 1 How is oxidative stress handled by peroxiredoxins? mean 580 tokens / 332 words, range 413-822 mean 4.8 citations; 5 of 6 answers carry a Sources sectionWhat this says
Length tracks how much distinct material comes back. The shortest answer had 1 citation, the longest had 10. That is the prompt working as written — it instructs the model to be comprehensive over whatever it is given.
So retrieval is no longer the binding constraint on answer length. Cutting documents further would lose grounding without necessarily shortening the prose, because the instruction is to be thorough about whatever arrives.
332 words is more moderate than it felt. Worth saying plainly: the answers are long-ish for a chat surface, but this is a couple of dense paragraphs plus a citation list, not runaway output.
The cheap experiment
A one-line prompt change — something to the effect of "prioritise the most relevant findings rather than covering every document provided" — then re-run this measurement. No retrieval risk, ~10 minutes. The wording is a voice decision, so it wants a human.
Caveat
Answers are not reproducible run to run even at temperature 0; two runs of the same six questions gave 411-826 and 413-822. Treat these as approximate, and use the same six questions when comparing before and after.
🤖 Generated with Claude Code
Two related findings from the retrieval work in PR #167. Recording rather than acting: both are product decisions.
1. Answer length has two levers; only one has been pulled
MAX_DOCUMENTS_PER_COLLECTIONnow caps the context at 40 documents (~6.3k tokens), down from ~222 (~32k). That helps, butsrc/retrievers/reactome/prompt.pyindependently instructs the model to be exhaustive over whatever it receives:plus a mandatory de-duplicated Sources list on every answer.
Nothing in the prompt says how long an answer should be, or that the most relevant few matter more than covering everything supplied. So the model will still try to cover all 40 documents.
The prompt is the more direct lever for length, and it does not cost retrieval recall — documents can stay in context for grounding and citation while the narrative covers fewer of them. Wording is a voice decision, so leaving it to whoever owns the prompt.
Observed in practice: "the answers provided by the chat are too long ... having the 11th pathway or reaction on a list is probably not useful at all. You would want to talk about the top ten or top five."
2. The document budget belongs to the caller, not the retriever
MAX_DOCUMENTS_PER_COLLECTIONis currently a module constant, which assumes every consumer wants the same amount of context. That will not hold:Same retriever, different budget per caller. That is a per-request parameter, not a constant, and it is the same seam as the headless-API extraction: the caller states how much context it wants instead of the retriever deciding for everyone.
Worth building that way during the retriever rewrite rather than accreting a second constant afterwards.
Note
The current value of 10 per collection is a starting point matching what a single retriever returns, not a tuned number. It should be calibrated against answer quality — likely lower. See #171 and the ragas evaluation work.
🤖 Generated with Claude Code