Skip to content

Fix what the model is told (review, area 1a) - #303

Merged
adamjohnwright merged 1 commit into
mainfrom
review-1a-model-inputs
Oct 3, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
review-1a-model-inputs

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

This is the third group of fixes from the max-level code review of src/api and src/handoff.

Finding Fix Test (each sabotaged and seen to fail)
A handoff seeded the summary's data without the rules the summary was written under, so follow-ups were answered from input known to produce "12 of 1280". summary_instruction() and DATA_RULES live in summarise.py. The endpoint and the seed both use them. Seed carries the rules
"Every one of them is significant" was added whenever the count was inexact, including when only 2 of 12 were significant. Said only when every shown pathway is significant. Otherwise UNKNOWN_COUNT_INSTRUCTION. 12/12, 2/12 and 0/12 cases
A temporary failure at claim time told the model the result was deleted. The seed is given the outcome and says "temporary" for a failure. Failed vs gone
Summaries ignored config.yml's llm: block and the temperature rules, so fixed-temperature models failed every summary. Content blocks were dropped, and the stream still reported summarised. _summary_llm() resolves the model the way the chat does. _text() reads content blocks. An empty stream is a failure. o3's temperature; blocks read; empty stream not summarised
Two concurrent requests generated different summaries, and the later one overwrote the earlier. The aggregate fallback regenerated instead of reusing. Single-flight per key. The store never overwrites. The fallback reuses a stored aggregate. Two readers, one model call, one text; store keeps the first

Verified

  • ./checks.sh passes.
  • The browser test passes against a local instance: an analysis summary from the real endpoint, a handoff, the same summary shown, a follow-up answered from the data, and an unknown id reported.

Also removes new_thread_id, which nothing called.

🤖 Generated with Claude Code

- A handoff's seeded data now carries the rules and instructions the
  summary was written under. The chat model reads the same data for every
  follow-up, and without them answered from the input that once produced
  "12 significant out of 1280". One function, summary_instruction, now
  builds them for the endpoint and the seed.
- "Every one of them is significant" is said only when it is true. When
  the count is inexact for another reason, a neutral instruction says the
  overall count cannot be determined.
- A temporary fetch failure at claim time is no longer reported to the
  model as the result having been deleted.
- Summaries choose their model the way the chat does: config.yml's llm
  block and resolve_temperature. Fixed-temperature models rejected 0.0,
  failing every summary while the chat worked. Content-block chunks are
  read, and a stream that yields no text is a failure, not 'summarised'.
- Concurrent requests for one summary wait for the first and reuse it; a
  stored summary is never overwritten; an aggregate fallback reuses a
  stored aggregate summary. A reader could be served, on reload or in
  Continue in chat, a summary they were never shown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 82d3998 into main Oct 3, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the review-1a-model-inputs branch October 3, 2026 12:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant