Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
59 changes: 49 additions & 10 deletions docs/architecture/learning-tier/plan-2026-09-19.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,10 @@ outstanding:** §7. **Base:** `main` @ `4f8e3e0f`.
**Item 1 — the mentor writes.** Today the mentor agent can *read* learning state on every harness but records
nothing: no MCP teach-back writer exists (CLI only), kiro pre-approves one writer, claude grants none, the
persona's only concrete "record progress" line points at a different tool. Result: zero writer calls in
143,973 historical messages. This item adds the missing writer, grants the additive writers on the harnesses
that have an in-repo grant grammar, gives every harness one machine-readable trigger table, and proves the
pipe is open with a scripted learner on all six installed harnesses — claiming only what was observed.
143,973 historical messages. This item adds the missing writer, names the additive writers in every harness
definition (no new pre-approvals — owner decision 2, §7), gives every harness one machine-readable trigger
table, and proves the pipe is open with a scripted learner on all six installed harnesses — claiming only what
was observed.
Comment on lines +12 to +15

**Item 2 — one honest number for Jev.** A single pre-registered, directional measurement of TypeSafe's Jev
as a *filter* over harness-proposed struggle candidates, on the repo's 13-session human-labelled gold read
Expand All @@ -25,7 +26,9 @@ Jev call before the pre-registration commit.
| 1 | `feat/learning-tier-fed` | `~/code/personal/tools/studyloop-wt/tier` | own PR, mergeable independently |
| 2 | `feat/jev-struggle-eval` | `~/code/personal/tools/studyloop-wt/jev-eval` | own PR (eval code + receipts only) |

Both cut from `main` @ `4f8e3e0f`. No code dependency either way (item 2 scores frozen gold). Each stage:
Item 1's branch is cut from `main` when S1-0 starts; item 2's is **not cut until milestone 0.5.x is closed**
(owner decision 1, 2026-09-23 — item 1 is MVP work, item 2 a post-MVP measurement). No code dependency either
way (item 2 scores frozen gold). Each stage:
commit in coherent steps with why-bodies; full suite diffed against a clean control worktree (only the RED
tests may flip); push only when the owner says so. The `feat/jev-judge` planning branch (Stage 0 receipt,
this plan) merges first or is cherry-picked — owner's call (§7).
Expand All @@ -36,21 +39,28 @@ this plan) merges first or is cherry-picked — owner's call (§7).
Record, from source, the grant mechanism per harness and freeze the writer sets:

- `W_auto` = `log_topic`, `log_struggle`, `record_teachback` (new), `record_plan_learning` — additive, learner-
agreed or learner-stated. Pre-approved.
agreed or learner-stated. **Named in every harness definition; prompt-per-call everywhere** (owner decision 2,
§7). No new pre-approval on any harness this branch; kiro's existing `log_topic` allowlist entry stays as the
one recorded asymmetry (`execute_bash` is pre-approved there, so removing it buys no parity).
- `W_srs` = `record_study_progress`, `log_review_outcome`, `record_topic_progress` — SRS mutators that accept
an unverified `card_hash`/topic id (`mcp/tools.py:114,648`). **Stay prompt-per-call** this branch.
- Grant grammar: kiro `allowedTools` (`agents/kiro/study-mentor.json`); claude `settings.json` permissions +
tool names in `socratic-mentor.md`; opencode already `studyloop *` (recorded as not least-privilege, unchanged);
- Grant grammar: kiro `allowedTools` (`agents/kiro/study-mentor.json`, unchanged); claude `settings.json`
permissions (unchanged — no block) + tool names in `socratic-mentor.md` (**must name each `W_auto` tool**:
its `tools:` line lists `Read, Write, Grep, Bash` and no `mcp__studyloop__*` tool today, so naming is what
makes the writers reachable at all — naming is instruction, not approval); opencode already `studyloop *`
(recorded as not least-privilege, unchanged);
codex/pi/grok have **no in-repo grant grammar** — canonical persona via session-dir `AGENTS.md`, approval is
harness-side (`adapters/{codex,pi,grok}.py`). No `agents/grok/` is created.
harness-side (`adapters/{codex,pi,grok}.py`). No `agents/grok/` is created. Parity is of the *instruction*
(trigger table + `W_auto` set, pinned by test), never of approval — the receipt records each harness's
approval cell in the capability matrix as it is.

Finish: `docs/architecture/learning-tier/receipts/s1-0-capability-lock.md` committed with file:line per harness.

### S1-RED (1 day)
| Test | File | Fails on `main` because |
|---|---|---|
| `record_teachback` MCP tool exists, validates exactly as `cli/_teachback.py:16-46` (cardinality, ints, 1–4, `review_type` enum), lands one row honouring CHECK, no row on failure, `session_id` bound from session state | `tests/test_mcp_teachback.py` (new) | tool absent |
| kiro `allowedTools` ⊇ `W_auto`; claude permissions ⊇ `W_auto` and `socratic-mentor.md` names each; opencode wildcard covers `W_auto`; codex/pi/grok canonical persona names each `W_auto` tool | `tests/test_adapter_parity.py` (extend) | grants/names absent |
| every harness definition names each `W_auto` tool (kiro `tools`, claude `socratic-mentor.md` `tools:`, opencode persona, codex/pi/grok canonical persona); **no new pre-approval**: kiro `allowedTools` ∩ `W_auto` == {`log_topic`} exactly, claude `settings.json` still has no permissions block, opencode wildcard unchanged; each harness's approval cell in the capability matrix equals what its definition file says | `tests/test_adapter_parity.py` (extend) | names absent |
| `agents/shared/recording-protocol.md` exists; fenced YAML trigger table parses; every trigger names a writer in `W_auto`; every projected copy carries identical bytes; `agents/manifest.json` hashes match; `persona.md` no longer routes "record progress" to `tutor-checkpoint` | `tests/test_docs_harness_tier_contract.py` (extend) | file absent; `:32` still names `tutor-checkpoint` |
| Isolation: child process with conflicting `HOME`/`XDG_*` runs each `W_auto` writer → 0 writes outside the sandbox, including Markdown | `tests/test_writer_isolation.py` (new) | passes or fails on `main` — if it passes, keep it as a guard, do not count it as RED |
| No-trigger + duplicate: a replayed sequence with a non-triggering turn writes nothing; a repeated `record_teachback` call produces the documented behaviour (two rows — teach-backs are events) | `tests/test_mcp_teachback.py` | tool absent |
Expand All @@ -60,7 +70,9 @@ Finish: N RED committed; `pytest` on a clean control worktree shows **exactly**
### S1-GREEN (1–2 days)
- `mcp/tools.py`: `record_teachback` delegating to `history/teachback.py:27` through a **shared validator**
extracted from `cli/_teachback.py` (one implementation, both surfaces; CLI kept).
- Grants: kiro `allowedTools` += `W_auto`; claude `settings.json` += `W_auto`, `socratic-mentor.md` names them.
- Names, not grants (owner decision 2, §7): `socratic-mentor.md`'s `tools:` line names each `W_auto` tool and the
readers; kiro `allowedTools` and claude `settings.json` **unchanged**; every other definition names the writers
through the projected protocol.
- `agents/shared/recording-protocol.md`: YAML table (`trigger`, `writer`, `required_ids`, `consent`) + one-line
prose per trigger: after teach-back **and learner agreement** → `record_teachback`; 2+ rounds without
breakthrough → `log_struggle`; end of session → `log_topic` per concept touched; plan wind-down →
Expand Down Expand Up @@ -164,10 +176,37 @@ between J-b-filtered and J-b-unfiltered*, nothing more. Three-seat council reads

1. **Two branches instead of the one requested** (D13) — accept the split, or insist on one branch and accept
that Jev spend/privacy review then gates the agent-definition merge.
- **Owner verdict, 2026-09-23: two branches, and item 2's branch is not cut until milestone 0.5.x is
closed.** Reasoning recorded with the verdict: item 1 is the MVP blocker (the mentor never writes to
the learning tier), item 2 is a post-MVP measurement; and the repository's own history is the argument
against sharing a vehicle — 17 `archive/*` tags whose tips never reached `main`, six parallel branches
opened 7–8 September and archived on the 10th, ~244 commits with no patch-equivalent on `main`. The
owner's self-check ("finished work waiting on unrelated work") was confirmed against that record.
Consequence: decision 4 is deferred to the day S2 starts (see below).
2. **Claude/opencode grants:** approve pre-approving `W_auto` on claude (today: none) and leaving opencode's
wildcard untouched this branch.
- **Owner verdict, 2026-09-23: no pre-approval on Claude — prompt per call; OpenCode's wildcard untouched
this branch.** Stated by the owner ("No pre-approval on Claude; prompt per call"), then confirmed by
delegation after the parity question. Evidence that made it cheap: kiro already pre-approved `log_topic`
and opencode already allowed `studyloop *`, and both still logged zero writer calls — the absence is the
*instruction* (no MCP teach-back writer, the persona's record-progress line pointing elsewhere), not the
prompt. Consequences taken with it: no new pre-approval on *any* harness this branch (kiro's `log_topic`
entry stays as the one recorded asymmetry); `socratic-mentor.md` must still *name* each `W_auto` tool,
because its `tools:` line lists no `mcp__studyloop__*` tool today and an unnamed tool is unreachable —
naming is instruction, not approval; parity is pinned on the instruction (trigger table + `W_auto` set),
and each harness's approval cell is recorded in the capability matrix as it is. The S1-0 bullets and the
test-table parity row above were rewritten to this verdict.
3. **The owner-led noticing episode** — ten minutes, once, after S1-GREEN; scheduled when?
- **Scheduled by rule, 2026-09-23 (owner delegated):** the first real study session the owner runs after
S1-GREEN reaches `main` — ten minutes, no logging asked for. The S1-SIM receipt records the date and the
adjudicated trigger/no-trigger outcomes afterwards. The owner may name a date instead; the rule stands
until he does.
4. **Jev spend cap** for S2 (estimate: 13 sessions × ≤ 8 candidates × 2 arms, under $1 at $0.042/Mtok; the
candidate-generation harness turns cost credits on the owner's subscription — ~13–26 turns).
- **Deferred, 2026-09-23 (consequence of decision 1):** S2 does not start before 0.5.x closes, so the cap
is decided on the day item 2's branch is cut, not before. Nothing in item 1 depends on it.
5. Merge order for the `feat/jev-judge` planning branch (Stage 0 receipt + this plan): first, or cherry-pick
the docs into `feat/learning-tier-fed`.
- **Resolved, 2026-09-23: cherry-picked onto `main`** by the housekeeping PR
([#35](https://github.com/NetDevAutomate/StudyLoop/pull/35)); `feat/learning-tier-fed` branches from
`main` with the plan already on it.
Loading