From 797fd149241194990fb2f1274a99167fc38018c6 Mon Sep 17 00:00:00 2001 From: NetDevAutomate Date: Wed, 23 Sep 2026 09:59:15 +0100 Subject: [PATCH 1/3] =?UTF-8?q?docs(learning-tier):=20owner=20decision=201?= =?UTF-8?q?=20recorded=20=E2=80=94=20two=20branches,=20item=202=20waits=20?= =?UTF-8?q?for=200.5.x?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Decision 1 of the plan §7 queue (#34). The owner accepted the council D13 split and went further: feat/jev-struggle-eval is not cut until milestone 0.5.x is closed, because item 1 (the mentor writes to the learning tier) is the MVP blocker and item 2 is a post-MVP measurement. The reasoning is recorded beside the verdict rather than in chat: the repository history shows 17 archive tags whose tips never reached main and six parallel branches opened 7-8 Sept and archived on the 10th, which is what one shared vehicle risks. Consequence recorded: decision 4 (Jev spend cap) is deferred to the day S2 starts. Decision 5 (merge order) is marked resolved by the housekeeping PR #35 cherry-pick, as issue #34 already states. §2 cut-point sentence updated from the stale 4f8e3e0f. --- .../learning-tier/plan-2026-09-19.md | 16 +++++++++++++++- 1 file changed, 15 insertions(+), 1 deletion(-) diff --git a/docs/architecture/learning-tier/plan-2026-09-19.md b/docs/architecture/learning-tier/plan-2026-09-19.md index 6818505b..8c20ddaa 100644 --- a/docs/architecture/learning-tier/plan-2026-09-19.md +++ b/docs/architecture/learning-tier/plan-2026-09-19.md @@ -25,7 +25,9 @@ Jev call before the pre-registration commit. | 1 | `feat/learning-tier-fed` | `~/code/personal/tools/studyloop-wt/tier` | own PR, mergeable independently | | 2 | `feat/jev-struggle-eval` | `~/code/personal/tools/studyloop-wt/jev-eval` | own PR (eval code + receipts only) | -Both cut from `main` @ `4f8e3e0f`. No code dependency either way (item 2 scores frozen gold). Each stage: +Item 1's branch is cut from `main` when S1-0 starts; item 2's is **not cut until milestone 0.5.x is closed** +(owner decision 1, 2026-09-23 — item 1 is MVP work, item 2 a post-MVP measurement). No code dependency either +way (item 2 scores frozen gold). Each stage: commit in coherent steps with why-bodies; full suite diffed against a clean control worktree (only the RED tests may flip); push only when the owner says so. The `feat/jev-judge` planning branch (Stage 0 receipt, this plan) merges first or is cherry-picked — owner's call (§7). @@ -164,10 +166,22 @@ between J-b-filtered and J-b-unfiltered*, nothing more. Three-seat council reads 1. **Two branches instead of the one requested** (D13) — accept the split, or insist on one branch and accept that Jev spend/privacy review then gates the agent-definition merge. + - **Owner verdict, 2026-09-23: two branches, and item 2's branch is not cut until milestone 0.5.x is + closed.** Reasoning recorded with the verdict: item 1 is the MVP blocker (the mentor never writes to + the learning tier), item 2 is a post-MVP measurement; and the repository's own history is the argument + against sharing a vehicle — 17 `archive/*` tags whose tips never reached `main`, six parallel branches + opened 7–8 September and archived on the 10th, ~244 commits with no patch-equivalent on `main`. The + owner's self-check ("finished work waiting on unrelated work") was confirmed against that record. + Consequence: decision 4 is deferred to the day S2 starts (see below). 2. **Claude/opencode grants:** approve pre-approving `W_auto` on claude (today: none) and leaving opencode's wildcard untouched this branch. 3. **The owner-led noticing episode** — ten minutes, once, after S1-GREEN; scheduled when? 4. **Jev spend cap** for S2 (estimate: 13 sessions × ≤ 8 candidates × 2 arms, under $1 at $0.042/Mtok; the candidate-generation harness turns cost credits on the owner's subscription — ~13–26 turns). + - **Deferred, 2026-09-23 (consequence of decision 1):** S2 does not start before 0.5.x closes, so the cap + is decided on the day item 2's branch is cut, not before. Nothing in item 1 depends on it. 5. Merge order for the `feat/jev-judge` planning branch (Stage 0 receipt + this plan): first, or cherry-pick the docs into `feat/learning-tier-fed`. + - **Resolved, 2026-09-23: cherry-picked onto `main`** by the housekeeping PR + ([#35](https://github.com/NetDevAutomate/StudyLoop/pull/35)); `feat/learning-tier-fed` branches from + `main` with the plan already on it. From 046fa499bb847bd90a99e60a57b16037faae075b Mon Sep 17 00:00:00 2001 From: NetDevAutomate Date: Wed, 23 Sep 2026 10:34:06 +0100 Subject: [PATCH 2/3] =?UTF-8?q?docs(learning-tier):=20owner=20decisions=20?= =?UTF-8?q?2=20and=203=20recorded=20=E2=80=94=20no=20new=20pre-approvals,?= =?UTF-8?q?=20noticing=20episode=20by=20rule?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Decision 2: no pre-approval of W_auto on Claude (prompt per call), OpenCode wildcard untouched this branch; consequence taken: no new pre-approval on any harness, kiro keeps its existing log_topic entry as the one recorded asymmetry. Evidence recorded beside the verdict: kiro and opencode already had grants and still logged zero writer calls, so the defect is the instruction, not the prompt. Naming is distinguished from approving: socratic-mentor.md must name each W_auto tool because its tools: line names none today. The §1 paragraph, S1-0 bullets and the parity test-table row are rewritten to match. Decision 3: the noticing episode is scheduled by rule — the first real study session after S1-GREEN reaches main — with the date recorded in the S1-SIM receipt. --- .../learning-tier/plan-2026-09-19.md | 39 +++++++++++++++---- 1 file changed, 31 insertions(+), 8 deletions(-) diff --git a/docs/architecture/learning-tier/plan-2026-09-19.md b/docs/architecture/learning-tier/plan-2026-09-19.md index 8c20ddaa..09ee6c48 100644 --- a/docs/architecture/learning-tier/plan-2026-09-19.md +++ b/docs/architecture/learning-tier/plan-2026-09-19.md @@ -9,9 +9,10 @@ outstanding:** §7. **Base:** `main` @ `4f8e3e0f`. **Item 1 — the mentor writes.** Today the mentor agent can *read* learning state on every harness but records nothing: no MCP teach-back writer exists (CLI only), kiro pre-approves one writer, claude grants none, the persona's only concrete "record progress" line points at a different tool. Result: zero writer calls in -143,973 historical messages. This item adds the missing writer, grants the additive writers on the harnesses -that have an in-repo grant grammar, gives every harness one machine-readable trigger table, and proves the -pipe is open with a scripted learner on all six installed harnesses — claiming only what was observed. +143,973 historical messages. This item adds the missing writer, names the additive writers in every harness +definition (no new pre-approvals — owner decision 2, §7), gives every harness one machine-readable trigger +table, and proves the pipe is open with a scripted learner on all six installed harnesses — claiming only what +was observed. **Item 2 — one honest number for Jev.** A single pre-registered, directional measurement of TypeSafe's Jev as a *filter* over harness-proposed struggle candidates, on the repo's 13-session human-labelled gold read @@ -38,13 +39,20 @@ this plan) merges first or is cherry-picked — owner's call (§7). Record, from source, the grant mechanism per harness and freeze the writer sets: - `W_auto` = `log_topic`, `log_struggle`, `record_teachback` (new), `record_plan_learning` — additive, learner- - agreed or learner-stated. Pre-approved. + agreed or learner-stated. **Named in every harness definition; prompt-per-call everywhere** (owner decision 2, + §7). No new pre-approval on any harness this branch; kiro's existing `log_topic` allowlist entry stays as the + one recorded asymmetry (`execute_bash` is pre-approved there, so removing it buys no parity). - `W_srs` = `record_study_progress`, `log_review_outcome`, `record_topic_progress` — SRS mutators that accept an unverified `card_hash`/topic id (`mcp/tools.py:114,648`). **Stay prompt-per-call** this branch. -- Grant grammar: kiro `allowedTools` (`agents/kiro/study-mentor.json`); claude `settings.json` permissions + - tool names in `socratic-mentor.md`; opencode already `studyloop *` (recorded as not least-privilege, unchanged); +- Grant grammar: kiro `allowedTools` (`agents/kiro/study-mentor.json`, unchanged); claude `settings.json` + permissions (unchanged — no block) + tool names in `socratic-mentor.md` (**must name each `W_auto` tool**: + its `tools:` line lists `Read, Write, Grep, Bash` and no `mcp__studyloop__*` tool today, so naming is what + makes the writers reachable at all — naming is instruction, not approval); opencode already `studyloop *` + (recorded as not least-privilege, unchanged); codex/pi/grok have **no in-repo grant grammar** — canonical persona via session-dir `AGENTS.md`, approval is - harness-side (`adapters/{codex,pi,grok}.py`). No `agents/grok/` is created. + harness-side (`adapters/{codex,pi,grok}.py`). No `agents/grok/` is created. Parity is of the *instruction* + (trigger table + `W_auto` set, pinned by test), never of approval — the receipt records each harness's + approval cell in the capability matrix as it is. Finish: `docs/architecture/learning-tier/receipts/s1-0-capability-lock.md` committed with file:line per harness. @@ -52,7 +60,7 @@ Finish: `docs/architecture/learning-tier/receipts/s1-0-capability-lock.md` commi | Test | File | Fails on `main` because | |---|---|---| | `record_teachback` MCP tool exists, validates exactly as `cli/_teachback.py:16-46` (cardinality, ints, 1–4, `review_type` enum), lands one row honouring CHECK, no row on failure, `session_id` bound from session state | `tests/test_mcp_teachback.py` (new) | tool absent | -| kiro `allowedTools` ⊇ `W_auto`; claude permissions ⊇ `W_auto` and `socratic-mentor.md` names each; opencode wildcard covers `W_auto`; codex/pi/grok canonical persona names each `W_auto` tool | `tests/test_adapter_parity.py` (extend) | grants/names absent | +| every harness definition names each `W_auto` tool (kiro `tools`, claude `socratic-mentor.md` `tools:`, opencode persona, codex/pi/grok canonical persona); **no new pre-approval**: kiro `allowedTools` ∩ `W_auto` == {`log_topic`} exactly, claude `settings.json` still has no permissions block, opencode wildcard unchanged; each harness's approval cell in the capability matrix equals what its definition file says | `tests/test_adapter_parity.py` (extend) | names absent | | `agents/shared/recording-protocol.md` exists; fenced YAML trigger table parses; every trigger names a writer in `W_auto`; every projected copy carries identical bytes; `agents/manifest.json` hashes match; `persona.md` no longer routes "record progress" to `tutor-checkpoint` | `tests/test_docs_harness_tier_contract.py` (extend) | file absent; `:32` still names `tutor-checkpoint` | | Isolation: child process with conflicting `HOME`/`XDG_*` runs each `W_auto` writer → 0 writes outside the sandbox, including Markdown | `tests/test_writer_isolation.py` (new) | passes or fails on `main` — if it passes, keep it as a guard, do not count it as RED | | No-trigger + duplicate: a replayed sequence with a non-triggering turn writes nothing; a repeated `record_teachback` call produces the documented behaviour (two rows — teach-backs are events) | `tests/test_mcp_teachback.py` | tool absent | @@ -175,7 +183,22 @@ between J-b-filtered and J-b-unfiltered*, nothing more. Three-seat council reads Consequence: decision 4 is deferred to the day S2 starts (see below). 2. **Claude/opencode grants:** approve pre-approving `W_auto` on claude (today: none) and leaving opencode's wildcard untouched this branch. + - **Owner verdict, 2026-09-23: no pre-approval on Claude — prompt per call; OpenCode's wildcard untouched + this branch.** Stated by the owner ("No pre-approval on Claude; prompt per call"), then confirmed by + delegation after the parity question. Evidence that made it cheap: kiro already pre-approved `log_topic` + and opencode already allowed `studyloop *`, and both still logged zero writer calls — the absence is the + *instruction* (no MCP teach-back writer, the persona's record-progress line pointing elsewhere), not the + prompt. Consequences taken with it: no new pre-approval on *any* harness this branch (kiro's `log_topic` + entry stays as the one recorded asymmetry); `socratic-mentor.md` must still *name* each `W_auto` tool, + because its `tools:` line lists no `mcp__studyloop__*` tool today and an unnamed tool is unreachable — + naming is instruction, not approval; parity is pinned on the instruction (trigger table + `W_auto` set), + and each harness's approval cell is recorded in the capability matrix as it is. The S1-0 bullets and the + test-table parity row above were rewritten to this verdict. 3. **The owner-led noticing episode** — ten minutes, once, after S1-GREEN; scheduled when? + - **Scheduled by rule, 2026-09-23 (owner delegated):** the first real study session the owner runs after + S1-GREEN reaches `main` — ten minutes, no logging asked for. The S1-SIM receipt records the date and the + adjudicated trigger/no-trigger outcomes afterwards. The owner may name a date instead; the rule stands + until he does. 4. **Jev spend cap** for S2 (estimate: 13 sessions × ≤ 8 candidates × 2 arms, under $1 at $0.042/Mtok; the candidate-generation harness turns cost credits on the owner's subscription — ~13–26 turns). - **Deferred, 2026-09-23 (consequence of decision 1):** S2 does not start before 0.5.x closes, so the cap From 3b8a7537559de70920fb784fea07fa33252ddf19 Mon Sep 17 00:00:00 2001 From: NetDevAutomate Date: Wed, 23 Sep 2026 10:49:29 +0100 Subject: [PATCH 3/3] docs(learning-tier): S1-GREEN grants bullet rewritten to decision 2 (names, not grants) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The stage list still said kiro allowedTools += W_auto and claude settings.json += W_auto, contradicting the §7 verdict recorded two commits earlier. Now: names in every definition, both grant files unchanged. --- docs/architecture/learning-tier/plan-2026-09-19.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/architecture/learning-tier/plan-2026-09-19.md b/docs/architecture/learning-tier/plan-2026-09-19.md index 09ee6c48..4e17fad2 100644 --- a/docs/architecture/learning-tier/plan-2026-09-19.md +++ b/docs/architecture/learning-tier/plan-2026-09-19.md @@ -70,7 +70,9 @@ Finish: N RED committed; `pytest` on a clean control worktree shows **exactly** ### S1-GREEN (1–2 days) - `mcp/tools.py`: `record_teachback` delegating to `history/teachback.py:27` through a **shared validator** extracted from `cli/_teachback.py` (one implementation, both surfaces; CLI kept). -- Grants: kiro `allowedTools` += `W_auto`; claude `settings.json` += `W_auto`, `socratic-mentor.md` names them. +- Names, not grants (owner decision 2, §7): `socratic-mentor.md`'s `tools:` line names each `W_auto` tool and the + readers; kiro `allowedTools` and claude `settings.json` **unchanged**; every other definition names the writers + through the projected protocol. - `agents/shared/recording-protocol.md`: YAML table (`trigger`, `writer`, `required_ids`, `consent`) + one-line prose per trigger: after teach-back **and learner agreement** → `record_teachback`; 2+ rounds without breakthrough → `log_struggle`; end of session → `log_topic` per concept touched; plan wind-down →