Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
32 commits
Select commit Hold shift + click to select a range
ee43863
test(ci): RED — the nightly installer job must plant a harness before…
NetDevAutomate Sep 19, 2026
57b3883
fix(ci): plant harness markers in the nightly installer job's isolate…
NetDevAutomate Sep 19, 2026
4f8e3e0
docs(tasks): item 7 step 4 done — #25 (D-D), #26 (D-E) filed; #21 sta…
NetDevAutomate Sep 19, 2026
ef319a7
test(now): RED — item 5 (D-F): energy demand for repair and the body-…
NetDevAutomate Sep 19, 2026
326abcf
feat(now): item 5 (D-F) — energy demand for repair, and the body-doub…
NetDevAutomate Sep 19, 2026
8e9cbdf
fix(now): the body double names ready plans only
NetDevAutomate Sep 19, 2026
6d5a2d0
docs(spec): the no-plan payload is byte-identical only when nothing i…
NetDevAutomate Sep 19, 2026
0fdda5c
docs(council): brief for review 7 — item 5 (D-F), reviewed tree 6d5a2d0e
NetDevAutomate Sep 19, 2026
d193525
fix(now): quote every offered command as one literal shell argument (…
NetDevAutomate Sep 19, 2026
c1d7f2f
fix(now): the milestone-deferral line no longer promises repair (revi…
NetDevAutomate Sep 19, 2026
3347567
feat(web): the Today card hands a body-double proposal's plan to the …
NetDevAutomate Sep 19, 2026
f814dc3
fix(web): the Today notes block is not "Your plans" for a learner wit…
NetDevAutomate Sep 19, 2026
7194b66
test(now): pin the item-5 boundaries the council asked for; tolerate …
NetDevAutomate Sep 19, 2026
02e282e
docs(now): the body double's ordering claim now matches the engine (r…
NetDevAutomate Sep 19, 2026
bdf4d6c
docs: the no-plan contract names both plan-independent changes (revie…
NetDevAutomate Sep 19, 2026
3eb31f2
docs(council): review 7 seats, arbitration (GATE ACCEPT), rubric row …
NetDevAutomate Sep 19, 2026
849c78a
docs(receipts): full-suite matched control re-run on the review-7 tree
NetDevAutomate Sep 19, 2026
a094637
test(now): RED -- the body-double reason leads with recorded progress…
NetDevAutomate Sep 20, 2026
c519bc2
feat(now): the body-double reason leads with recorded progress and de…
NetDevAutomate Sep 20, 2026
e770733
docs(rubric-3b): council decision brief for the owner; receipt amende…
NetDevAutomate Sep 20, 2026
0760254
test(now): RED — the low-demand teach-back door is the micro form; a …
NetDevAutomate Sep 20, 2026
f2cb572
feat(now): the low-demand teach-back door is the micro form; a learni…
NetDevAutomate Sep 20, 2026
c7cc1b0
test(agents): RED — the teach-back protocol carries the low-energy gu…
NetDevAutomate Sep 20, 2026
214df37
docs(agents): the micro teach-back carries a low-energy guided-explan…
NetDevAutomate Sep 20, 2026
b161291
docs(rubric-3b): owner verdicts (a)–(d) recorded; (c)'s three fixes l…
NetDevAutomate Sep 20, 2026
8f851d3
test(now): RED — the due collector's copy of a live struggle defers w…
NetDevAutomate Sep 20, 2026
fb63fb6
feat(now): the due collector's copy of a live struggle defers with th…
NetDevAutomate Sep 20, 2026
291e869
docs(rubric-3b): (e) yes; (f) resolved by fb63fb6b; every reading re-…
NetDevAutomate Sep 20, 2026
a4503a4
docs(rubric-3b): (c) confirmed by the owner on the both-collectors sc…
NetDevAutomate Sep 20, 2026
90d749c
docs(rubric-3b): the two 0.6.0 findings are issues #30 (concrete firs…
NetDevAutomate Sep 20, 2026
60d5f94
test(web): RED — the timer's two init() runs make two state reads; a …
NetDevAutomate Sep 20, 2026
96806fe
fix(web): the session timer's init() runs once per page load
NetDevAutomate Sep 20, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion .github/workflows/nightly-install.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,15 +72,40 @@ jobs:
# (~/.kiro, ~/.claude, ~/.codex, ~/.config/opencode). With a fresh HOME
# there is no Python 3.12 either, so this also exercises the script's
# `uv python install` fallback every night.
#
# The script refuses, by design, when no supported AI tool is present
# ("No supported AI tools detected"), and the runner has none: no
# harness binary on PATH and nothing in the fresh HOME. So the fixture
# plants the four harness markers `detect_available_agent_tools()`
# reads without a binary -- ~/.kiro, ~/.claude, ~/.pi, ~/.grok -- and
# the install-agents step then exercises all four link sets for real.
# (Without this the job failed every night from 2026-09-15, before its
# verify steps ever ran; pinned by test_ci_workflow_contract.py.)
env:
HOME: ${{ runner.temp }}/home
UV_TOOL_DIR: ${{ runner.temp }}/tools
UV_TOOL_BIN_DIR: ${{ runner.temp }}/bin
run: mkdir -p "$HOME" && ./scripts/install.sh --non-interactive --no-smoke
run: |
mkdir -p "$HOME/.kiro" "$HOME/.claude" "$HOME/.pi" "$HOME/.grok"
./scripts/install.sh --non-interactive --no-smoke

- name: Verify installed CLI entry points
env:
UV_TOOL_BIN_DIR: ${{ runner.temp }}/bin
run: |
"$UV_TOOL_BIN_DIR/studyloop" --version
"$UV_TOOL_BIN_DIR/session-export" --help

- name: Verify installed agent definitions
# One artefact per planted harness, each a path `install agents`
# writes for that tool (installers.py _TOOL_LINKS / _configure_*).
env:
HOME: ${{ runner.temp }}/home
run: |
test -e "$HOME/.kiro/agents/study-mentor.json"
test -e "$HOME/.kiro/agents/study-plan-architect.json"
test -e "$HOME/.claude/agents/socratic-mentor.md"
test -e "$HOME/.claude/agents/study-plan-architect.md"
test -e "$HOME/.pi/agent/AGENTS.md"
test -e "$HOME/.grok/hooks/studyloop.json"
test -e "$HOME/.grok/rules/session-db.md"
4 changes: 2 additions & 2 deletions .secrets.baseline

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

4 changes: 2 additions & 2 deletions agents/manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -70,8 +70,8 @@
"updated": "2026-09-14"
},
"shared/teach-back-protocol.md": {
"hash": "9bbe8831f1c74837",
"updated": "2026-09-14"
"hash": "09ee534106322244",
"updated": "2026-09-20"
},
"shared/wind-down-protocol.md": {
"hash": "5b1ec3303b8d1086",
Expand Down
2 changes: 2 additions & 0 deletions agents/shared/teach-back-protocol.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,8 @@ Only score 2 dimensions: Accuracy and Own Words (max 8 points).
| 6-7 | Good recall in own words | On track |
| 8 | Strong | Consider accelerating interval |

**Low-energy fallback.** When `studyloop now` offers a micro teach-back on a low-energy day (a concept recorded as `learning`, energy demand `low`), the ask is one sentence and nothing more. If the student cannot produce it — a blank, "I don't know", or the source's own words — move straight to a **guided explanation**: give the explanation in two or three plain sentences with a networking analogy, then ask the student to say back ONE phrase of it in their own words. This is the fallback, not the four-round Stuck ladder in `socratic-engine.md`: at low energy that ladder is the productive struggle the day cannot carry, and a stall there is the confidence damage (RSD) the deferral of live struggles exists to prevent. Do not score the blank as a teach-back — at low energy it measures the day, not the concept. Record the phrase that came back (`studyloop progress "<concept>" -t <topic> -c learning`) and let the structured review at the 7-day mark measure the concept.

## Detecting Understanding vs Memorisation

### Red Flags (Surface Learning)
Expand Down
1,734 changes: 1,734 additions & 0 deletions docs/architecture/plan-integration/council/brief-review7-2026-09-19.md

Large diffs are not rendered by default.

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# Arbitration — council review 7 (item 5 of the plan-integration follow-on programme: D-F)

**Date:** 2026-09-19 · **Arbiter:** coordinating agent (owner present; the owner chose the seats' key and asked
for the run) · **Reviewed tree:** `feat/energy-demand-body-double` @ `6d5a2d0e` (four commits on `main`
`4f8e3e0f`; range `4f8e3e0f..6d5a2d0e`, 15 files, +1,042/−25). Seats ran against
`brief-review7-2026-09-19.md` (`review7/manifest.json`, run 12:10:45Z; astra and qwen `finish_reason=stop`,
**grok `length`** — its 16,000-token cap cut its answer inside §3 Refutations, so its verdict, findings and the
first refutation are complete and its §4 Gate is absent; not re-run, since every grok finding is 🔵/💡 and its
verdict is ACCEPT). **Corrections landed at:** `d1935256` (F1), `c1d7f2f2` (F5), `3347567e` (F7), `f814dc36`
(grok's label), `7194b66d` (F4 pins + a collector crash they surfaced), `02e282e6` (F2 claim), `bdf4d6c6`
(F3/F6 wording). Two corrections the agent found on its own re-read *before* the brief was frozen are part of
the reviewed tree (`8e9cbdf5` ready plans only; `6d5a2d0e` wording) and are listed in the brief §1.

**Brief size, recorded:** 117.6 KB, 1,734 lines (~31k prompt tokens per seat): D-F and rubric row 3 verbatim,
design §5 with its amendments and GREEN decisions, every diff in the range in full, row 3b as written, the
control receipt, reference facts, ten numbered check questions — including the scope of the plan-independent
deferral asked outright as (a).

## Seats and verdicts

| Seat | Verdict | 🔴 | 🟡 | 🔵/💡 |
| --- | --- | --- | --- | --- |
| `openai.gpt-6-astra` | ACCEPT-WITH-CORRECTIONS | F1 shell quoting; F2 ordering claim | F3 compatibility promise; F4 unpinned boundaries; F5 contradictory copy | F6 match semantics; F7 Web door context |
| `qwen3-coder` | ACCEPT-WITH-CORRECTIONS | scope (gate deferral on a plan); ordering claim | medium/high partition; docs; gentle-recall fallback; Web door; substring tests; row 3b ask | `evidence_command` reuse |
| `grok-4.6` | ACCEPT | — | — | 14-day boundary + unparseable path untested; decision 5 untested; two-plan case untested; Web pre-fill; "Your plans" label; medium=high behaviourally; `evidence_command` fine |

### Method

Each 🔴/🟡 was reproduced before acceptance — F1 with a real `/bin/sh` and a stub `studyloop` (the title's
`$(touch pwned)` ran), F2 through the engine at two modalities (42 vs 118 / 58 / 34; 60 vs 100 / 76 / 34), F5 by
reading row 3b's own rendered lines — then fixed one commit each with a RED test named by the seat where one was
named, and stash-proved where a fix could be reverted alone (F1: 3 of 7 titles fail without it).

### Findings and dispositions

| # | Finding (seat) | Disposition | Commit / test |
| --- | --- | --- | --- |
| F1 | The offered command replaced `"` with `\"` — presentation, not quoting; a plan title with `$(…)` executed when pasted (astra 🔴). | **Accepted, reproduced.** One `_shell_word` helper: plain text keeps the golden's double-quoted form, shell-special text is `shlex.quote`d; both command builders use it (`_evidence_command` had the same shape for every concept and topic). | `d1935256`; `test_body_double_command_preserves_title_as_one_literal_shell_argument` (7 titles through `/bin/sh`; stash-proved 3/7 red without the fix) |
| F2 | "Base below `MILESTONE_BASE_SCORE` so every real candidate outranks it" is false after adjustments: at low energy the proposal (42) outranks a hands-on task that energy penalises (34) (astra 🔴, qwen 🔴; grok 💡 "acceptable, do not lower the base"). | **Claim corrected, behaviour kept.** Reproduced exactly as computed. Every due and conversation candidate still outranks it at every modality; only a hands-on task the low-energy rule already penalises sits beneath — the energy rule doing what the finding asked ("instead of the least-bad task"), not a filter: nothing is removed from the ranking. Astra's proposed post-scoring floor would put a penalised hands-on task above the proposal at low energy, i.e. re-recommend the class of work the finding objected to. The *claim* was the defect: corrected in the constant's comment, the docstring, the spec delta's rule-5 clause and design §5 decision 4; the judgement is the owner's — row 3b reading (e). | `02e282e6`; `test_body_double_ordering_after_adjustments_follows_the_energy_rule` (recall and conversation modality) |
| F3 | "No plan and nothing deferred → byte-identical" still false: every struggle candidate gains `metadata.energy_demand` at every energy; older struggles and weak teach-backs defer too (astra 🟡). | **Accepted.** The contract now names exactly the two plan-independent changes, in the spec delta, both docs and design decision 1. | `bdf4d6c6` (docs) + `02e282e6` (spec paragraph, same file as F2's edit); `test_no_plan_eligible_repair_exposes_demand_without_deferral` (in `7194b66d`) |
| F4 | Five unpinned invariants: the 14-day boundary and unparseable dates; confidence/teach-back precedence; the capability matrix; decision 5 (deferred sole representative → eligible milestone conversation); two ready plans (astra 🟡; grok 🔵 on three of them). | **Accepted, all five written.** The boundary pin surfaced a **pre-existing collector crash**: `_struggle_candidates` sorted by `row["last_seen"]`, so a legacy `study_progress` row with a NULL `last_seen` raised and lost every struggle candidate; it now sorts as the oldest and derives as live. | `7194b66d`; `test_energy_demand_recency_boundaries_and_unknown_dates`, `test_energy_demand_confidence_and_teachback_precedence`, `test_repair_demand_capability_matrix`, `test_deferred_repair_allows_only_eligible_milestone_conversation`, `test_body_double_two_ready_plans_has_deterministic_context` |
| F5 | The milestone-deferral sentence ended "plan-related review and repair stay available" one line above the line deferring the repair (astra 🟡). | **Accepted.** Engine reason and CLI line now promise only due recall and gentle review; pinned as whole sentences. | `c1d7f2f2`; `test_cli_milestone_deferral_does_not_promise_live_repair` |
| F6 | Rule 7 can attach a husk reference to the proposal through its topic, so "names ready plans only" is true of the proposal's text, not of every rendered reference (astra 🔵). | **Accepted as a spec sentence**, no code change: the two-stage behaviour is now described in the rule-5 clause; `test_body_double_is_never_synthesised_for_an_unready_plan` already permits the husk ref and pins the proposal's text. | `02e282e6` |
| F7 | The Web door navigated only; the Body Double picker opened blank while the CLI door carried the plan title (astra 🔵, grok 🔵, qwen 🟡). | **Accepted, built** — the same event-not-storage handoff `today-resume` uses: `startAction` dispatches `body-double-request` {activity, energy}; `bodyDoubleSession.init` adopts it. Nothing starts on its own. | `3347567e`; JS `starting a body-double primary hands its plan to the Body Double view…`, `bodyDoubleActivity…` |
| G1 | `hasPlanContext` is true for a no-plan deferred repair and the block is headed "Your plans" (grok 🔵). | **Accepted.** `planNotesLabel()`: "Your plans" once any note involves a plan, else "Set aside today". | `f814dc36`; JS `planNotesLabel…` |
| Q1 | Gate the deferral on an active plan (qwen 🔴). | **Rejected** — see below. | design §5 decision 1; row 3b reading (d) |
| Q2 | Collapse the medium/high partition (qwen 🟡). | **Rejected** — see below. | design §5 decision 7; spec sentence |
| Q3 | Synthesise a same-concept gentle recall as the no-plan floor (qwen 🟡; grok 🔵 "follow-on"). | **Not taken; recorded as a follow-on** and put to the owner. | design §5 decision 8; row 3b reading (d) |
| Q4 | Row 3b should ask the owner about the no-plan consequence (qwen 🟡; grok's tail). | **Accepted.** Readings (d) and (e) added, printed from the real engine on `bdf4d6c6`. | `receipts/now-rubric-2026-09-16.md` |

### Rejected, with reasons

- **Q1 — gate repair deferral on an active plan (qwen 🔴).** Two seats and the owner's own words go the other
way: the finding is "hands-on repair of a *live* struggle on a low-energy day compounds the struggle (RSD)" —
a fact about the learner's day, not about plans — and D-F's item (1) says "at low energy a live struggle
defers like new work" with no plan condition. Gating it would leave the original "no" standing for every
learner without a plan. Astra: "reintroducing an unsafe recommendation merely because the learner lacks a plan
would contradict the rationale for D-F". Grok: the same, and "listed, not recommended, is the honest state".
Kept, with the compatibility contract re-worded to the truth (F3) and the no-plan floor put to the owner as
row 3b reading (d) — the one person who can say whether the starter is the floor they want.
- **Q2 — collapse `medium` into `high` (qwen 🟡).** True that nothing in `ENERGY_CAPABILITY` sits between 3 and 6,
so the two classes never differ in eligibility today. Astra and grok both keep the class as explanatory state
— the payload says *why* a repair asks for 4/10 rather than 6/10, and the next energy scale change would
otherwise need the derivation rebuilt. Kept; the spec now says both classes need at least medium energy today.
- **Astra's F2 fix (a post-scoring floor for `body_double`).** Rejected in favour of correcting the claim: the
floor would rank a hands-on task the low-energy rule penalises above the proposal, which is the shape of
recommendation the finding objected to. The behaviour is pinned exactly as it is and goes to the owner.
- **Q3 — a same-concept gentle recall as the no-plan floor.** Astra: "do not synthesise recall on the deferred
struggle merely to maintain topical relevance; that would invent an unvalidated lower-demand action." Grok:
a follow-on, not a merge gate. Recorded as decision 8 and asked in row 3b (d).

### Refutations, weighed

- Astra 1 ("42 < any real candidate" false) — **true**; F2. Astra 2–3 (byte-for-byte still false; "live struggle"
too narrow) — **true**; F3. Astra 4 ("each GREEN decision is a test" not established: decisions 5 and the
two-plan case untested) — **true**; F4 now pins both. Astra 5 ("the Today card starts it" — navigation only) —
**true at the reviewed tree**; F7 built the handoff. Astra 6 ("zero regressions" is about failing ids at GREEN,
not the reviewed tree) — **accepted as stated**; a full suite runs on the final tree below. Astra 7
("`learning` means recovered" is the chosen proxy, not established) — **accepted**; row 3b reading (c) asks
the owner exactly that. Grok's one refutation (the same F2 claim) — **true**. Qwen: none.

### Verification after fixes (tree `bdf4d6c6` + the rubric edit)

- `test_now_plan_guidance.py` **71 passed** (47 at the reviewed tree, +24 test cases from F1–F5 including
parametrisations), golden byte-identical; with `test_learning_decision.py` and the docs contract **96 passed**;
JS **147/147** (+3); `openspec validate` valid;
`mkdocs --strict` exit 0; ruff / ruff format / pyright clean on every touched file. Full suite and matched
control on the final tree `3eb31f2d`: 30 failed / **5146 passed** / 14 errors vs control 30 / 5120 / 14; item5 ∖
control = ∅, control ∖ item5 = ∅, item5 ∖ committed environmental set = ∅
(`receipts/full-suite-control-item5-2026-09-19.md`, second table).

### Process findings

- The brief's check question (a) put the arbiter's own hardest decision to the seats directly and got a 2–1
split with reasons on both sides — more useful than a unanimous nod. Keep doing that.
- Grok's 16k cap cut its Gate section. Its findings were complete and all 🔵/💡, so no re-run; a future brief
of this size should either raise `--max-tokens` for that seat or ask for the Gate before the Refutations.
- One commit carried two logical changes: `3347567e` (F7) also contains the `planNotesLabel()` function that
`f814dc36` (G1) relies on, because both edits were in `today-panel.js` when F7 was staged. Recorded rather than
rewritten; both are unpushed and tested.
- F4's boundary pin finding a real, pre-existing crash (`None` `last_seen`) is the argument for writing boundary
tests even when the design says the boundary is "obvious".

## Gate decision

**GATE: ACCEPT** — for the tree at `bdf4d6c6` (with the rubric-row edit), not the reviewed tree. Every 🔴 and 🟡
is either landed with a discriminating test (F1, F3, F4, F5, F7, G1, Q4) or rejected here with the reason and
the owner's question that replaces it (Q1, Q2, Q3, astra's F2 remedy). Item 5 is ready to merge once CI is green
on the branch; the change is **not** archived until the owner scores row 3b.

## Still open for the owner

1. **Rubric row 3b, five readings** — (a) sit with the plan rather than repair a live struggle; (b) the proposal
beneath an unrelated due recall; (c) the gentle teach-back on a `learning` concept; **(d) the no-plan floor**
(starter + deferred line, where the hands-on repair used to be); **(e) the proposal above an unrelated
hands-on drill at low energy**. (d) and (e) are the two places the seats split; only the owner closes them.
2. Whether a same-concept gentle recall should be synthesised for a no-plan learner (decision 8) — after (d).
47 changes: 47 additions & 0 deletions docs/architecture/plan-integration/council/review7/manifest.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
{
"run_at": "2026-09-19T12:10:45+00:00",
"brief": "docs/architecture/plan-integration/council/brief-review7-2026-09-19.md",
"brief_sha256": "6fde40f067eccee3c0fb46af74ea448f8e02c8eab1eb078c6229e1cc3b45a15b",
"system_sha256": "5f21e273f399cdb0455a9fadb8575cc3e231ce0d7931f517206335c98c3f934b",
"seats": [
{
"model": "openai.gpt-6-astra",
"ok": true,
"reasoning_chars": 0,
"finish_reason": "stop",
"elapsed_s": 66.1,
"usage": {
"prompt_tokens": 30920,
"completion_tokens": 3939,
"total_tokens": 34859
},
"error": null
},
{
"model": "grok-4.6",
"ok": true,
"reasoning_chars": 0,
"finish_reason": "length",
"elapsed_s": 107.3,
"usage": {
"prompt_tokens": 32856,
"completion_tokens": 16000,
"total_tokens": 48856
},
"error": null
},
{
"model": "qwen3-coder",
"ok": true,
"reasoning_chars": 0,
"finish_reason": "stop",
"elapsed_s": 26.1,
"usage": {
"prompt_tokens": 31758,
"completion_tokens": 1845,
"total_tokens": 33603
},
"error": null
}
]
}
Loading
Loading