Skip to content

Commit 2c8fba7

Browse files
proof(stage-d): repair receipt chain; control arm made exact; council record for looks 1–2
Second-family seat (deepseek-3.2, 74685923) on look 2: VALID-WITH-NOTES, one blocking finding — look 2's previous_receipt_sha256 hashed the look 1 file AFTER it had been edited in place to add a VOIDED field, not as committed. The orchestrator's error. Remedy: look 1 restored to its committed bytes (sha ec9d6576…, identical to ca55c65), the void notice moved to a sidecar (stage-d-look1-b1clean.VOIDED.md), look 2 re-run chained to the pristine file. All four arms reproduce identically; chain verified against `git show ca55c65:…`. Rule adopted: a receipt is never mutated after commit. B1_planner now carries the shipped visibility predicate so "identical to B1 except the planner" is literally true (predicate excludes 0/5,879 sessions; per-question hits unchanged). council-stage-d-looks.md records both seats' findings with dispositions and evidence. Standing: look 1 voided (numbers reproduced), look 2 valid; planner +0.142 established, clean +0.184 established, clean-over-planner +0.042 not established; stop rule 2 armed (one more flat look ends the G1 DEV looks). G1 not claimed. SEALED untouched. Council runs 13.
1 parent 543edf4 commit 2c8fba7

6 files changed

Lines changed: 99 additions & 25 deletions

File tree

‎.secrets.baseline‎

Lines changed: 3 additions & 3 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.
Lines changed: 59 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,59 @@
1+
# Receipt council — Stage D DEV looks 1 and 2 (G1 family)
2+
3+
Ruler: "A gate passes only when … a two-family council review of the receipt finds no blocking
4+
objection to its validity. The orchestrator may override a council objection only by citing
5+
artifact evidence that refutes it, recorded in the receipt. Council findings are otherwise leads:
6+
nothing is acted on until verified against the artifact."
7+
8+
| seat | model | run | reviewed | verdict |
9+
|---|---|---|---|---|
10+
| A | gpt-5.6-terra | `4ed818c7` | look 1 (`ca55c653`) | **VOID** — 4 blocking |
11+
| B | deepseek-3.2 | `74685923` | look 2 (`543edf45`) + amendment-002 + r2 | **VALID-WITH-NOTES** — 1 blocking |
12+
13+
Council runs after this record: 13 of 60.
14+
15+
## Seat A (look 1) — dispositions
16+
17+
Every statistic re-derived exactly by the seat (macro recalls, paired cluster bootstrap CI95
18+
[+0.091562, +0.281404] with the receipt's seed, per-stratum non-inferiority); arm conformance to
19+
fusion-spec v1 exact; 91/91 questions returned exactly five distinct ids; no gold-controlled
20+
ingestion or ranking path. The blocking findings were all provenance-record findings.
21+
22+
| # | finding | verified | disposition |
23+
|---|---|---|---|
24+
| A1 | candidate commit equals the spec commit rather than postdating it | true | **not void** — the ruler requires declaration *before the look*; spec and arm were committed at `690a37d4` and the look ran from that HEAD 10 s later; receipt commit `ca55c653` postdates both. Receipts now record `fusion_spec.declared_commit` for a mechanical check. |
25+
| A2 | no fusion-spec field on the receipt | true | **accepted** — `score.py` records `fusion_spec.{path, sha256, declared_commit}`. |
26+
| A3 | gold receipt `dev.sha256` does not identify the DEV file | true, and `sealed.sha256` does not identify the SEALED file either | **accepted; void upheld** — Stage 2 record defect; see amendment-002. |
27+
| A4 | gold receipt `corpus_digest` ≠ result receipts'; ruler voids on mismatch | true; `a0df30bb…` reproduces from no item set | **accepted; void upheld** — cannot be overridden: the refuting evidence does not exist. |
28+
| A5 | B1-vs-B0 non-inferiority not recorded on the aggregate | true | **accepted** — `non_inferiority_macro` on every comparison. |
29+
| A6 | intervention bundles planner + index; scope filter differs | scope: `visibility_sql` excludes 0/5,879 — not a confound; bundling: true | **accepted as a measurement question** → `B1_planner` control, declared in spec v1.1 before look 2. |
30+
| A7 | the 42 errors are not the whole lift (49-question subset 0.494 vs 0.320 correct) | true | note; confirmed by look 2's control. |
31+
| A8–A9 | arm conformance exact; statistics re-derive exactly | — | notes. |
32+
| A10 | store not cryptographically bound; `KNOWLEDGE_PROOF_STORE` can redirect | true | **accepted** — `--store` records sha256 + size on the receipt. |
33+
| A11 | DEV evidence only; not a G1 pass | true | affirmed in every receipt's wording. |
34+
35+
Outcome: **look 1 voided, still counted** (the number was seen). Remedy: ruler-amendment-002,
36+
`gold-v2-receipt-r2.json` via committed `recertify_gold.py`, harness provenance fields.
37+
38+
## Seat B (look 2 + remedy) — dispositions
39+
40+
| # | finding | verified | disposition |
41+
|---|---|---|---|
42+
| B1 | look 2's `previous_receipt_sha256` hashes the look 1 file **after** an in-place `VOIDED` edit, not as committed at `ca55c653` | true (`c4d52cea…` vs committed `ec9d6576…`) | **accepted (blocking)** — the orchestrator's error. Look 1 restored to its committed bytes; the void notice moved to a sidecar (`stage-d-look1-b1clean.VOIDED.md`); look 2 re-run chained to the pristine file; all four arms reproduced identically; chain verified `previous_receipt_sha256 == sha256(git show ca55c653:…)`. **Rule adopted: a receipt is never mutated after commit; annotations live in sidecars.** |
43+
| B2 | `B1_planner` omits `visibility_sql`, so "identical to B1 except the planner" was not literally true (0 sessions excluded, so no recall effect) | true | **accepted** — predicate added; per-question hits identical; spec claim now exact. |
44+
| B3 | `B1_planner` recovers 5 questions B1 threw on; supports planner attribution | true | note. |
45+
| B4 | r2 hashes reproduce from committed code + gold files + DB | reproduced by the seat | the remedy holds. |
46+
| B5 | clean-index increment Δ+0.042 CI95 [+0.009, +0.085] not established | true | recorded as such in the look 2 receipt commit. |
47+
| B6 | look 2 is not an "improvement" (lower bound unchanged at +0.092) | true | **stop rule 2 armed: one more flat DEV look ends Stage D's G1 looks.** |
48+
| B7 | amendment does not weaken SEALED protection (`corpus_digest.sealed` recorded; SEALED receipt must match it) | — | note. |
49+
50+
## Standing of the measurements
51+
52+
- **Look 1** — voided for provenance; numbers re-derived exactly by seat A and reproduced by look 2.
53+
- **Look 2** — valid after B1/B2 remedies: `B1_planner` 0.249 (+0.142 vs B1, CI95 [+0.060, +0.231],
54+
established); `B1_clean` 0.291 (+0.184, CI95 [+0.092, +0.281], established); clean-index
55+
increment +0.042 (CI95 [+0.009, +0.085], **not** established; never loses: 4/0/87).
56+
- **Attribution:** the planner is most of the effect. The shipped AND-first query form is too
57+
strict independent of crashing (0.479 vs 0.320 on the 49 questions it answered). F-B0-1 is
58+
thereby a *measured* defect in the shipped path worth +0.142 macro recall@5 on DEV.
59+
- **G1 is not claimed.** DEV looks used: 2 of 4. SEALED untouched.
Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,14 @@
1+
# stage-d-look1-b1clean.json — VOIDED
2+
3+
Voided by the receipt council (seat gpt-5.6-terra, run `4ed818c7`) for provenance: the Stage 2
4+
gold receipt's hashes are non-reproducible (ruler-amendment-002, F3/F4). The seat re-derived the
5+
receipt's statistics exactly; they were superseded by `stage-d-look2-planner-control.json`, which
6+
reproduces B0/B1/B1_clean identically and adds the `B1_planner` control.
7+
8+
The receipt file itself is **byte-identical to its commit `ca55c653`** (sha256
9+
`ec9d6576028140fa…`). It was briefly edited in place to carry this notice (commit `f5c1597d`),
10+
which the second council seat (deepseek-3.2, run `74685923`) correctly flagged as breaking the
11+
chain's intent; the edit was reverted and the notice moved here. Rule from here on: **a receipt is
12+
never mutated after commit — annotations live in a sidecar.**
13+
14+
Counts as DEV look 1 of ≤ 4 for the G1 family.

‎docs/architecture/session-memory/receipts/stage-d-look1-b1clean.json‎

Lines changed: 1 addition & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -1879,10 +1879,5 @@
18791879
}
18801880
}
18811881
},
1882-
"previous_receipt_sha256": "c66058836c85140882c816dba392fe835a64b804c688285b8f37f4d1285f6b9b",
1883-
"VOIDED": {
1884-
"by": "receipt council seat gpt-5.6-terra run 4ed818c7",
1885-
"reason": "gold receipt provenance hashes non-reproducible (ruler-amendment-002 F3/F4); numbers re-derived exactly by the seat and superseded by look 2",
1886-
"counts_as_dev_look": 1
1887-
}
1882+
"previous_receipt_sha256": "c66058836c85140882c816dba392fe835a64b804c688285b8f37f4d1285f6b9b"
18881883
}

‎docs/architecture/session-memory/receipts/stage-d-look2-planner-control.json‎

Lines changed: 11 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
11
{
22
"receipt": "stage-d-look2-planner-control",
3-
"created_utc": "2026-09-10T03:19:55+00:00",
3+
"created_utc": "2026-09-10T03:32:48+00:00",
44
"ruler_commit": "a98331afd2bb0c864e416c943b4a51962fe659e2",
5-
"candidate_commit": "f5c1597df4c6862cc555882b8efce2f0a600b520",
5+
"candidate_commit": "543edf45847a66a385c0aca9839dc4edd4f193ef",
66
"b0_pin": "031dbab9f72fcc4077da2963d6e60163f4517438",
77
"fusion_spec": {
88
"path": "docs/architecture/session-memory/receipts/fusion-spec-v1.md",
@@ -41,8 +41,8 @@
4141
"macro": 0.08365261813537674
4242
},
4343
"latency_ms": {
44-
"p50": 3.164708847180009,
45-
"p95": 33.22508395649493
44+
"p50": 3.3488329499959946,
45+
"p95": 38.47483289428055
4646
},
4747
"errors": 42
4848
},
@@ -64,8 +64,8 @@
6464
"macro": 0.08365261813537674
6565
},
6666
"latency_ms": {
67-
"p50": 3.077833214774728,
68-
"p95": 33.17750012502074
67+
"p50": 3.4829580690711737,
68+
"p95": 36.87191708013415
6969
},
7070
"errors": 42
7171
},
@@ -87,8 +87,8 @@
8787
"macro": 0.2012771392081737
8888
},
8989
"latency_ms": {
90-
"p50": 30.75541602447629,
91-
"p95": 40.985000086948276
90+
"p50": 31.848042039200664,
91+
"p95": 41.24587494879961
9292
},
9393
"errors": 0
9494
},
@@ -110,8 +110,8 @@
110110
"macro": 0.1741727621037966
111111
},
112112
"latency_ms": {
113-
"p50": 53.07779205031693,
114-
"p95": 66.05166709050536
113+
"p50": 70.3739591408521,
114+
"p95": 86.80554083548486
115115
},
116116
"errors": 0
117117
}
@@ -2510,5 +2510,5 @@
25102510
}
25112511
}
25122512
},
2513-
"previous_receipt_sha256": "c4d52cea52bdf9a794e1659f87930149264ffaae9a938da81db108899263a37b"
2513+
"previous_receipt_sha256": "ec9d6576028140fad744946adf1df1cbf2781ee8e2027994b5692e21292a50c4"
25142514
}

‎scripts/knowledge_proof/proof_arms.py‎

Lines changed: 11 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -77,25 +77,31 @@ def B1_planner() -> Arm: # noqa: N802 - arm names are receipt labels, matched t
7777
Control arm that separates two effects bundled in ``B1_clean``: (i) a planner that
7878
never throws, and (ii) an index that holds prose only. This arm keeps the shipped
7979
``messages_fts`` over every archive row (tool echo, duplicates and all), the shipped
80-
``bm25(messages_fts)`` ranking, the shipped 200-row candidate budget and first-seen
81-
session dedup, and swaps *only* the query text for ``plan_prose_query(question)``.
80+
``bm25(messages_fts)`` ranking, the shipped scope-visibility predicate, the shipped
81+
200-row candidate budget and first-seen session dedup, and swaps *only* the query text
82+
for ``plan_prose_query(question)``.
8283
8384
If ``B1_planner`` ≈ ``B1_clean``, the lift is the planner. If ``B1_planner`` ≈ ``B1``
8485
on the questions ``B1`` answered, the lift is the clean index.
8586
"""
87+
import importlib
88+
8689
from learning_memory.store import plan_prose_query
8790

91+
visibility_sql = importlib.import_module("agent_session_tools.context.public").visibility_sql
92+
8893
def arm(archive: sqlite3.Connection, question: str) -> list[str]:
8994
planned = plan_prose_query(question)
9095
if not planned:
9196
return []
97+
visible, scope_params = visibility_sql(archive, "s.id")
9298
rows = archive.execute(
93-
"SELECT m.session_id FROM messages m "
99+
"SELECT s.id FROM messages m JOIN sessions s ON m.session_id = s.id "
94100
"JOIN messages_fts ON messages_fts.rowid = m.rowid "
95-
"WHERE messages_fts MATCH ? "
101+
f"WHERE messages_fts MATCH ? AND {visible} "
96102
"ORDER BY bm25(messages_fts), m.timestamp DESC "
97103
f"LIMIT {CANDIDATE_ROWS}",
98-
(planned,),
104+
[planned, *scope_params],
99105
).fetchall()
100106
seen: list[str] = []
101107
for (sid,) in rows:

0 commit comments

Comments
 (0)