Skip to content

Make the tests that passed for the wrong reasons honest (review, area 1a) - #304

Merged
adamjohnwright merged 1 commit into
mainfrom
review-1a-honest-tests
Oct 3, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
review-1a-honest-tests

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

This is the last group from the max-level review of src/api and src/handoff: tests that passed for reasons other than what they are named for. Each new or changed test was run against the mutation the reviewer used to show the old one was weak, and it fails under that mutation:

Test Was weak because Mutation it now catches
test_refuses_an_unknown_kind It posted kind: "search", so the 422 came from the missing answer_id kind: Literal["analysis", "bogus"], which skips the presence check
test_a_search_request_cannot_smuggle_an_analysis_field Same cause: no answer_id It now sends a valid answer_id and asserts no AnalysisHandoff is minted
test_cannot_widen_the_disclosure_tier get("anything") is None is always true It now asserts nothing was minted (HandoffStore.__len__)
test_the_rate_limit_applies (new) The fixture allowed 10,000, so deleting the limiter passed Limiter removed
test_what_is_kept_is_the_stripped_text…, …held_back…, …each_id_leads_to_its_own_text (new) The stub had no anchors, no sources list and nothing held back The held flush dropped from what is kept
test_what_the_model_is_sent_names_no_file_and_no_token (new) Nothing tested what the seed sends to the model data = dict(fetched.result)

./checks.sh passes.

🤖 Generated with Claude Code

From the review of src/api and src/handoff (area 1a). Each new or changed
test was checked against the mutation the reviewer used to show the old
one was weak, and fails under it:

- Unknown handoff kind: posted kind 'search', whose 422 came from the
  missing answer_id. Now kind 'bogus' with a stored summary and no human
  claim; routing unknown kinds to the analysis model now fails it.
- Smuggling analysis fields into a search request: now with a valid
  answer_id, asserting no analysis handoff is minted.
- Widening the tier asserted get('anything') is None, which is always
  true. Now: nothing was minted.
- The handoff limiter had no test (the fixture allows 10,000). Now three
  mints against a limit of two.
- What is kept for Continue in chat: the stub had no anchors, no sources
  list and nothing held back, so storing raw text passed. Now with all
  three, plus each id leading to its own answer.
- Nothing tested what a handoff seed sends the model; passing the raw
  result through, with token, file and sample names, left the suite green.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 84e1134 into main Oct 3, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the review-1a-honest-tests branch October 3, 2026 12:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant