Skip to content

Commit 736c558

Browse files
proof(stage-e): pre-register the claims writer; pin the G2 PoC set as a derived artefact
Before any model writes a claim: claims-writer-spec-v1.md fixes the model (claude-sonnet-5), the prompt (writer_prompt_v1.md, sha on every claim and receipt), the strict output schema, the insertion contract (≥1 overlap-aware unambiguous quote per claim, cross-session citations refused, no partial inserts, post-run substr() re-check proving unbound writes = 0), budgets (pilot 40, population ≤302, hard cap 400), and the G2 audit: blinded (statement, quote)-only entailment by a second family on 100 random claims, seed fixed, with a fixed failure taxonomy. What v1 deliberately omits is listed. The ruler binds G2 to "the 348 PoC sessions" with no committed list. Found its definition (RESULTS-final.md: updated ≥ 2026-08-01, ≥10 messages) and the run's own frozen corpus snapshot, whose sha matches the run's SHA256SUMS. pin_poc_set.py reproduces the rule against that snapshot: 345 (the snapshot was cleaned of empty rows AFTER the 348 was counted; 3 fell below threshold). All 345 exist today; 342 are in the store. Under typed events "≥10 messages" means prose events: 200 sessions. ruler-amendment-003 records both denominators and declares n_prose_ge10 = 200 primary for the G2 clause, without editing the ruler. Population is hash-ordered by session id. 26 of 60 DEV gold sessions lie inside the 342 by design of the gold split (verified, not assumed); the writer is gold-blind by construction (no gold file readable from its environment, asserted by test), not by instruction.
1 parent 86deee4 commit 736c558

6 files changed

Lines changed: 896 additions & 1 deletion

File tree

‎.secrets.baseline‎

Lines changed: 17 additions & 1 deletion
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.
Lines changed: 89 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,89 @@
1+
# Claims writer spec v1 — pre-registered before the first writer run (Stage E.0)
2+
3+
**Declared:** 2026-09-10, before any model has written a claim into `learning-memory.db`.
4+
Ruler binding: G2 ("wind-down produces citation-bound concepts whose quotes entail the
5+
proposition, audited"), the ≤ 400 writer-run cap, writer model `claude-sonnet-5`, $0 external
6+
spend (kiro roster only). ADR-0011 v1.1 Decision 3 ("each with at least one quote-bound citation —
7+
enforced by the database"), Decision 4 (read contract), council dispositions 9, 12, 13, 18.
8+
9+
## What a writer run is
10+
11+
One sub-agent run over **one session**: it receives the session's citable evidence rows
12+
(`visible_evidence(session_id)` — per-event `REPORTED` prose with `evidence_id`, ordered by
13+
turn/seq) plus the derived exchange flags for that session (`derive-v1`: `is_question`,
14+
`had_error`, `resolved`, concept tags), and returns **0–8 claims** as JSON. It never receives:
15+
retrieval code, the ruler, any gold file, any other session, the store path, or the ability to
16+
write. The orchestrator's harness validates and inserts; the model proposes only.
17+
18+
## Model and parameters (frozen for E.1 and E.2)
19+
20+
- Model: `claude-sonnet-5` (ruler). `spawn_run model=claude-sonnet-5`, one session per run.
21+
- Prompt: `scripts/knowledge_proof/writer_prompt_v1.md`; its sha256 is recorded on every receipt
22+
and on every claim row's `writer` field as `sonnet5/writer-v1/<sha8>`.
23+
- No tools for the writer beyond reading the packet it is given. No web. No spawning.
24+
- Output schema (strict JSON, rejected on any deviation):
25+
`{"claims":[{"kind":"Problem|Finding|Decision|Procedure|Preference","title":"≤120","statement":"≤500","tags":["2..5"],"confidence":0.5..1.0,"citations":[{"evidence_id":"…","quote":"exact substring"}]}]}`
26+
27+
## Population rule (gold-blind by construction)
28+
29+
- Population = the 342 ingested PoC sessions (`poc-set-g2.json`), ordered by
30+
`sha256(session_id)` ascending. **E.1 pilot** = the first 40 in that order. **E.2** = the
31+
remainder, in that order, until the writer-run cap or the population is exhausted.
32+
- The writer environment has **no read access to any gold file** — asserted by a test that greps
33+
the writer packet builder and the prompt for `gold`, `sealed`, `receipts/`, and by the packet
34+
being built from the store alone.
35+
- 26 DEV gold sessions lie inside the population (amendment-003). No builder run computes the
36+
SEALED overlap.
37+
38+
## Insertion contract (what the harness enforces, per claim)
39+
40+
1. Schema validity (above); `tags` 2–5 distinct lower-case tokens; `confidence` ∈ [0.5, 1.0].
41+
2. **≥ 1 citation**, each `quote` a non-empty, unambiguous (overlap-aware) substring of the named
42+
evidence row's body — resolved by `Store.add_claim`; refused otherwise. A refused citation
43+
refuses the whole claim (no partial insert).
44+
3. `evidence_id` must belong to the session in the packet (cross-session citation refused).
45+
4. Duplicate claim (same content id) → counted, not re-inserted.
46+
5. Every refusal is recorded with reason in the run receipt; **unbound writes = 0 by
47+
construction**, and the receipt proves it by re-checking every inserted citation with SQLite's
48+
own `substr()` after the run.
49+
50+
## Budgets
51+
52+
- Writer runs: E.1 ≤ 40, E.2 ≤ 302 (population remainder), retries ≤ 1 per session on a
53+
transport/JSON failure only (never on "no claims"). Hard cap 400 (ruler).
54+
- Per run: packet ≤ 48 KiB of evidence text (largest sessions are truncated to the first N
55+
evidence rows that fit; truncation recorded); response ≤ 8 claims.
56+
- Wall: the pilot must finish inside one monitor cycle budget (≤ 40 runs × ~2 min).
57+
58+
## G2 audit design (pre-registered)
59+
60+
- **Yield:** share of population sessions (primary denominator `n_prose_ge10 = 200`; literal
61+
`n_messages_ge10 = 345` also reported) with ≥ 1 inserted claim. Gate: ≥ 90 %.
62+
- **Unbound writes:** 0, proven by post-run `substr()` re-check of every citation.
63+
- **Blinded entailment audit:** a random 100 inserted claims (seed 20260910, drawn after E.2 or
64+
after E.1 if E.2 is not reached), each shown to a **second model family** (deepseek-3.2; gpt-5.6
65+
as tie-break) as `(statement, quote)` **only** — no title, tags, session, or writer identity —
66+
with the question "does the quote entail the statement? yes / partial / no", and a required
67+
one-line reason. Gate: ≥ 95 % `yes`. Every `partial`/`no` is classified into the taxonomy below.
68+
- **Failure taxonomy (fixed now):** `over-claim` (statement exceeds quote), `wrong-subject`
69+
(quote about something else), `hallucinated-detail` (statement adds facts), `procedure-not-
70+
shown` (claims a step the quote does not contain), `preference-inferred` (a preference stated as
71+
fact), `quote-too-thin` (quote true but trivially short), `other`.
72+
- **Mutation tests** (ruler): altered body, stale offsets, wrong evidence id, misaligned code-point
73+
span — already proven at the store level in Stage B/B.1 (`test_claim_citations.py`,
74+
`test_claims_immutable.py`); re-run and cited on the G2 receipt.
75+
76+
## What is deliberately NOT in v1
77+
78+
No supersession (`supersedes` stays NULL: the writer sees one session, so there is nothing to
79+
supersede); no `claim_relations`; no review-item generation; no tool-output citations (archive
80+
holds none); no lineage roll-up. Each is a later spec version.
81+
82+
## Reading the result
83+
84+
- If yield ≥ 90 % and entailment ≥ 95 %: G2 passes on the pilot+population and the claims arm may
85+
be declared in `fusion-spec-v2.md` for DEV look 3.
86+
- If yield < 90 %: report the distribution of "no claims" sessions by harness and prose count;
87+
investigate the *prompt*, never relax the trigger or the denominator.
88+
- If entailment < 95 %: report the taxonomy; the writer is not fit; the claims arm is **not**
89+
declared and Stage E stops with the receipt.

0 commit comments

Comments
 (0)