Skip to content

Answer length: the prompt instructs exhaustiveness, and the document budget should be per-caller #172

Description

@adamjohnwright

Two related findings from the retrieval work in PR #167. Recording rather than acting: both are product decisions.

1. Answer length has two levers; only one has been pulled

MAX_DOCUMENTS_PER_COLLECTION now caps the context at 40 documents (~6.3k tokens), down from ~222 (~32k). That helps, but src/retrievers/reactome/prompt.py independently instructs the model to be exhaustive over whatever it receives:

  • "answer the user's questions comprehensively, mechanistically, and with precision"
  • "Provide an information-rich narrative that explains not only what is happening but also how and why"
  • "Comprehensiveness: Capture all mechanistically relevant details available in Reactome"

plus a mandatory de-duplicated Sources list on every answer.

Nothing in the prompt says how long an answer should be, or that the most relevant few matter more than covering everything supplied. So the model will still try to cover all 40 documents.

lever state
documents reaching the model 222 → 40 (done)
prompt telling it to be exhaustive unchanged

The prompt is the more direct lever for length, and it does not cost retrieval recall — documents can stay in context for grounding and citation while the narrative covers fewer of them. Wording is a voice decision, so leaving it to whoever owns the prompt.

Observed in practice: "the answers provided by the chat are too long ... having the 11th pathway or reaction on a list is probably not useful at all. You would want to talk about the top ten or top five."

2. The document budget belongs to the caller, not the retriever

MAX_DOCUMENTS_PER_COLLECTION is currently a module constant, which assumes every consumer wants the same amount of context. That will not hold:

  • Search-results integration — the search results are the comprehensive part, so the chat answer alongside them should draw on fewer documents.
  • Chat — can afford more, being the primary surface.
  • Analysis summarisation — different again; the input is a result set rather than a question.

Same retriever, different budget per caller. That is a per-request parameter, not a constant, and it is the same seam as the headless-API extraction: the caller states how much context it wants instead of the retriever deciding for everyone.

Worth building that way during the retriever rewrite rather than accreting a second constant afterwards.

Note

The current value of 10 per collection is a starting point matching what a single retriever returns, not a tuned number. It should be calibrated against answer quality — likely lower. See #171 and the ragas evaluation work.

🤖 Generated with Claude Code

Activity

  1. adamjohnwright commented on Sep 4, 2026

    @adamjohnwright
    ContributorAuthor

    Measured: a token budget would be tighter than a document count, but is not currently binding

    From triaging #139, which proposes token-aware truncation with limits in config.yml.

    With the per-collection document cap now on main, measured over six real questions on Release95:

    40 docs   5,796 – 9,531 tokens   (1.6x spread)
    

    So a fixed document count is a loose proxy — the same 40 documents vary by 64% in tokens. A token budget bounds the thing that actually costs money and context.

    But it is not urgent: 9,531 tokens is 7% of gpt-4o-mini's 128k window, down from ~32k before the cap. Adding a second limit now would tighten a bound that is not binding.

    tiktoken is already in the lockfile (transitive via langchain-openai), so this would cost no new dependency when it is wanted.

    The more valuable idea from #139

    Putting the limits in config.yml rather than a module constant:

    retriever:
      context_truncation:
        max_docs: 15
        max_tokens: 12000

    That is the same direction as the per-caller budget in this issue — search-results integration wants fewer documents than chat — and it makes tuning a config change rather than a code change.

    It is blocked on the same plumbing as #151 (wiring AgentGraph to the YAML config): HybridRetriever is built inside create_profile_graphs, which has no Config. Worth doing all three together rather than piecemeal.

    Recommendation

    Keep the document cap as-is. When the config plumbing lands, move MAX_DOCUMENTS_PER_COLLECTION into config.yml alongside a max_tokens backstop, per-caller.

    🤖 Generated with Claude Code

  2. adamjohnwright commented on Sep 4, 2026

    @adamjohnwright
    ContributorAuthor

    Measured: how long the answers actually are

    First end-to-end run of the full RAG chain (question in, answer out) after the document cap landed. Six questions from tests/golden/questions.txt, Release95, gpt-4o-mini, counted with tiktoken rather than estimated.

      tokens   words   citations   question
         413     ~240       4      What does CDK5 phosphorylate in Alzheimer's...
         822     ~470       7      How does TP53 regulate PTEN transcription?
         576     ~330       5      What is the role of CDK12 in DNA repair...
         780     ~450      10      Which proteins are in the RNA polymerase II...
         457     ~260       2      What happens during Golgi fragmentation...
         429     ~250       1      How is oxidative stress handled by peroxiredoxins?
    
      mean 580 tokens / 332 words, range 413-822
      mean 4.8 citations; 5 of 6 answers carry a Sources section
    

    What this says

    Length tracks how much distinct material comes back. The shortest answer had 1 citation, the longest had 10. That is the prompt working as written — it instructs the model to be comprehensive over whatever it is given.

    So retrieval is no longer the binding constraint on answer length. Cutting documents further would lose grounding without necessarily shortening the prose, because the instruction is to be thorough about whatever arrives.

    332 words is more moderate than it felt. Worth saying plainly: the answers are long-ish for a chat surface, but this is a couple of dense paragraphs plus a citation list, not runaway output.

    The cheap experiment

    A one-line prompt change — something to the effect of "prioritise the most relevant findings rather than covering every document provided" — then re-run this measurement. No retrieval risk, ~10 minutes. The wording is a voice decision, so it wants a human.

    Caveat

    Answers are not reproducible run to run even at temperature 0; two runs of the same six questions gave 411-826 and 413-822. Treat these as approximate, and use the same six questions when comparing before and after.

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions