Skip to content

Commit 690a37d

Browse files
proof(stage-d): pre-declare fusion-spec v1 and the B1_clean arm before the first DEV look
The ruler requires the retrieval configuration on record before the number is seen. B1_clean = prose-only FTS (porter unicode61) over the archive-ingested learning-memory store, phrase-token planner, bm25 over 200 event rows, first 5 distinct session ids. No fusion, no claims, no embeddings — the arm that answers 'what does cleaning buy' and nothing else. Wire-tested on synthetic questions only.
1 parent 430986d commit 690a37d

2 files changed

Lines changed: 117 additions & 0 deletions

File tree

Lines changed: 46 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,46 @@
1+
# Fusion spec v1 — arms pre-declared before the first DEV look
2+
3+
**Declared:** 2026-09-10, before any DEV look on `learning-memory.db`. Ruler clause: "the fused
4+
arm's algorithm, per-source candidate budget, dedup rule and tie-break are versioned in
5+
`receipts/fusion-spec-v<N>.md` before the first DEV look; every arm returns exactly five results
6+
within the same byte budget." No arm below fuses more than one source yet; the spec exists so the
7+
retrieval configuration in force is on record before the number is seen.
8+
9+
## Controls (ruler, unchanged)
10+
11+
- **B0** — shipped `session_search` at pin `031dbab9`: `_session_search_queries` (AND then OR),
12+
`bm25(messages_fts)`, scope visibility, LIMIT 200 message rows → first 5 distinct session ids.
13+
- **B1** — the same path at the candidate commit. Non-inferiority of B1 vs B0 gates everything.
14+
15+
## Arm `B1_clean` (Stage D, "what cleaning buys")
16+
17+
*Question it answers:* does removing tool echo and exporter duplicates from the indexed text —
18+
and planning natural language so the query cannot throw — lift session recall, before any
19+
derivation, claims, embeddings or ontology exist?
20+
21+
- **Store:** `~/.local/share/studyloop/knowledge-proof/learning-memory.db`, schema v2, ingested by
22+
`archive-v1` / `archive-classifier-v1` (receipt `ingest-archive-v1.json`). Opened read-only.
23+
- **Index:** `prose_fts` — external content over `prose_events` (`user`, `assistant_prose` only),
24+
tokenizer `porter unicode61` (the schema default; `unicode61` is a later, separate arm).
25+
- **Query planning:** `Store.search_prose` planner — every whitespace token phrase-quoted, OR-joined,
26+
control/surrogate code points stripped. No AND stage (a deliberate difference from B1: the
27+
AND→OR fallback is the part of the shipped planner that crashes; measured, not assumed, by the
28+
42 DEV errors in the baseline receipt).
29+
- **Ranking:** `bm25(prose_fts)` ascending over event rows; **candidate budget 200 event rows**;
30+
session id = the event's `session_id`; first **5 distinct session ids** in rank order.
31+
- **Dedup rule:** by session id, first occurrence wins. **Tie-break:** bm25 then `events.id`
32+
ascending (ingest order).
33+
- **Byte budget:** identical to B1 — the arm returns ids only; payload budgets apply at Stage G.
34+
- **Latency:** measured cold on a fresh connection per receipt run; p95 ≤ 500 ms and ≤ 2 × B1.
35+
36+
## What is *not* in v1
37+
38+
No claims arm, no embeddings, no metadata filters, no lineage roll-up, no RRF. Each of those is a
39+
later spec version, declared before its own first look. The `unicode61` tokenizer variant is a
40+
separate arm (`B1_clean_u61`) that requires a second store build and is declared here only by name.
41+
42+
## Look accounting for Stage D
43+
44+
This look is the first of ≤ 4 DEV looks for the **G1** family. Improvement = the paired lower
45+
bound vs B1 rose. It is scored on gold **DEV** (`gold-v2-dev.json`, sha `eeca2aaf…`); the SEALED
46+
set is not touched by any Stage D activity.
Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
"""Feature arms for ``score.py`` (``--feature proof_arms:<name>``).
2+
3+
Every arm here is declared in ``receipts/fusion-spec-v<N>.md`` *before* its first DEV
4+
look; the docstring of each factory names the spec version it implements so a receipt
5+
can be checked against the declaration mechanically.
6+
7+
Arms that read ``learning-memory.db`` open their own read-only connection and ignore
8+
the archive connection ``score.py`` passes in: the harness's contract is
9+
``Arm = Callable[[sqlite3.Connection, str], list[str]]`` and the store is a different
10+
database from the gold's corpus. Session ids are identical across the two (ADR-0011 §6),
11+
which is what makes the same gold score both.
12+
"""
13+
14+
from __future__ import annotations
15+
16+
import os
17+
import pathlib
18+
import sqlite3
19+
from collections.abc import Callable
20+
21+
Arm = Callable[[sqlite3.Connection, str], list[str]]
22+
23+
K = 5
24+
CANDIDATE_ROWS = 200
25+
STORE_ENV = "KNOWLEDGE_PROOF_STORE"
26+
DEFAULT_STORE = pathlib.Path.home() / ".local/share/studyloop/knowledge-proof/learning-memory.db"
27+
28+
29+
def _store_path() -> pathlib.Path:
30+
return pathlib.Path(os.environ.get(STORE_ENV, str(DEFAULT_STORE)))
31+
32+
33+
def _open_store_ro(path: pathlib.Path) -> sqlite3.Connection:
34+
if not path.exists():
35+
raise FileNotFoundError(f"learning-memory store not found: {path}")
36+
return sqlite3.connect(f"file:{path}?mode=ro", uri=True)
37+
38+
39+
def B1_clean() -> Arm: # noqa: N802 - arm names are receipt labels, matched to the spec
40+
"""fusion-spec-v1 ``B1_clean``: prose-only FTS over the archive-ingested store.
41+
42+
Planner = ``learning_memory.store.plan_prose_query`` (phrase-quote every token, OR-join,
43+
strip control/surrogate code points); ranking = ``bm25(prose_fts)`` over
44+
``CANDIDATE_ROWS`` event rows; dedup by session id, first occurrence wins;
45+
tie-break bm25 then ``events.id``; exactly the first ``K`` distinct session ids.
46+
"""
47+
from learning_memory.store import plan_prose_query
48+
49+
conn = _open_store_ro(_store_path())
50+
51+
def arm(_archive: sqlite3.Connection, question: str) -> list[str]:
52+
planned = plan_prose_query(question)
53+
if not planned:
54+
return []
55+
rows = conn.execute(
56+
"SELECT e.session_id FROM prose_fts "
57+
"JOIN events AS e ON e.id = prose_fts.rowid "
58+
"WHERE prose_fts MATCH ? "
59+
"ORDER BY bm25(prose_fts), e.id "
60+
f"LIMIT {CANDIDATE_ROWS}",
61+
(planned,),
62+
).fetchall()
63+
seen: list[str] = []
64+
for (sid,) in rows:
65+
if sid not in seen:
66+
seen.append(sid)
67+
if len(seen) == K:
68+
break
69+
return seen
70+
71+
return arm

0 commit comments

Comments
 (0)