diff --git a/CHANGELOG.md b/CHANGELOG.md index e6fc6d54a..9ef0fd185 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,10 @@ experience may change before `1.0.0`. ## [Unreleased] +_Nothing yet._ + +## [0.5.0] - 2026-09-21 + ### Added - An acceptance-test tier modelled on the Terraform provider's `make testacc`: @@ -57,6 +61,60 @@ experience may change before `1.0.0`. messages past 300 ms ("last load: X s" from per-machine cold/warm records — never a percentage bar), and a web encoder-warm status chip fed by `GET /api/retrieval/health`. +- **Study Plans reach the surfaces you decide on.** One `PlanApplication` + seam (`studyloop.planning.application`) now carries every plan use case — + browse, inspect, prepare planning, active guidance, apply a change, assess — + and the CLI, the Web UI and the MCP server all go through it. Two bugs that + seam exists because of are closed: activation readiness is judged once on + the resulting document for every entry path (create-with-status and + whole-document replacement included), and a checkpoint whose database write + failed is reported as partial instead of as a clean record. +- **`studyloop now` is plan-aware.** An active plan biases the recommendation + — matching due work and a synthesised next milestone carry explicit plan and + milestone references, urgent reviews and fresh struggles can still outrank + new milestone work, several active plans are considered deterministically, + and a milestone beyond today's energy is deferred with a reason rather than + dropped. CLI `now`, Web Today, the recap and MCP `get_next_action` consume + one additive result; with no active plan the JSON is byte-identical to the + pinned golden. `get_next_action` gains `interleave` parity with the CLI. +- **Repair carries an energy demand of its own, and a body-doubling floor.** + A live struggle (`struggling`, seen within 14 days) asks for 6/10, an older + struggle or a weak teach-back for 4/10, a concept still `learning` for none; + below the day's capability a repair is deferred like new work into + `energy_deferred_repairs` — never ranked — while due recall and teach-back + reviews are never deferred. The due-review reader emits every `struggling` + row as a hands-on "guided repair" as well, so that copy carries the same + demand and defers with the repair, named once — without it the repair the + learner had just been spared came back as the primary. When nothing + plan-related fits the day and an active, ready plan exists, one low-scoring + proposal is synthesised: sit with the plan in a body-double session + (`studyloop study "" --mode co-study`), leading with the progress the + plan records and naming what is deferred. Every offered command quotes + learner-authored text as one shell word. +- **The low-energy teach-back is the one-sentence kind, with a way out.** A + concept still `learning` is offered as a *micro* teach-back + (`studyloop teachback … --type micro`) and its reason reads as the gentle + review it is, not "repair now"; the teach-back protocol gains a low-energy + fallback — when the sentence will not come, the mentor gives a short guided + explanation and asks for one phrase back, instead of walking the four-round + stuck ladder, and the blank is not scored. +- **Plan with the architect from the Web UI.** *Plans → Plan with architect* + launches the Study Plan Architect through the ordinary session machinery + (one-session authority, reconnect, the same console) with a `planning` + purpose; starting a conversation creates no draft. The manual form remains. + `studyloop plan repair ` opens the architect on an active plan the + readiness gate refuses to write, and `studyloop plan close ` reviews a + fully-checked plan against its evidence and proposes *extend* or *close* — + status never changes by itself, and a partial review never proposes a clean + close. A fully-checked plan that is not active can be closed or deleted; the + architect asks which. +- **MCP lifecycle parity for Study Plans**: `list_study_plans`, + `get_study_plan`, `get_planning_interview`, `create_study_plan`, + `update_study_plan`, `set_study_plan_status`, `set_study_plan_milestone`, + `evaluate_study_plan` and `delete_study_plan` (confirmation required), thin + adapters over the same seam; the mission is revisable through + `update_study_plan`. `studyloop doctor` reports a plan document it cannot + read by name instead of hiding it behind "all ready". ### Changed @@ -82,6 +140,39 @@ experience may change before `1.0.0`. ranked lists, cosine ≥ 0.999 per query); encoder load falls from ~2.9 s to ~0.2 s and a one-shot CLI hybrid search from ~2.9 s to ~0.33 s p50. +### Fixed + +- The Study Plan Architect's launch from the Web UI waits at most 8 s for the + session picker's options, so a stalled `/api/session/options` cannot hold a + click forever; leaving the page before a requested launch lands leaves a + session like any other (reattachable, ended with *End*) rather than a + half-state. +- A legacy `study_progress` row with no `last_seen` no longer crashes the + struggle collector — and with it `studyloop now` — on the way to a + recommendation. +- The release gate (`scripts/check-release-consistency.py --release`) read a + folded `deferred: >-` reason in an openspec change's `.openspec.yaml` as the + literal marker `>-`, so it printed no reason and would have accepted an + empty one — the unexplained deferral the guard exists to refuse. It now + reads the indented text and refuses an empty block. +- The Study Session view's timer ran its `init()` twice per page load (Alpine's + own call plus the markup's `x-init`), each run reading `/api/session/state`. + A tab sitting on the session picker could adopt a session started in + another tab when the slower read landed — Start hidden under the live + layout, no 409, no reattach offer. `init()` now runs once per page load; + a second tab's Start reaches the server and is offered reattach. Named by + the state captured at the click in CI (third occurrence, first with + evidence) and pinned by a `node --test` replay of that race. + +### Security + +- The offered commands in `studyloop now` (`evidence_command`, the body-double + door) quote every learner-authored value — plan titles, concepts, topics — + as a single shell word. A plan titled `SQL $(touch pwned) Windows` used to + run its substitution when the offered line was pasted into a shell. +- `anyio` 4.12.1 → 4.15.1 for two published advisories (CVE-2026-63374, + CVE-2026-64847); both dependency audits are clean again. + ### Removed - The forbidden-token test that failed the suite whenever "KiroCrew" appeared diff --git a/openspec/changes/archive/2026-09-21-plan-integration-followons/.openspec.yaml b/openspec/changes/archive/2026-09-21-plan-integration-followons/.openspec.yaml new file mode 100644 index 000000000..f08077847 --- /dev/null +++ b/openspec/changes/archive/2026-09-21-plan-integration-followons/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-09-16 diff --git a/openspec/changes/plan-integration-followons/design.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/design.md similarity index 100% rename from openspec/changes/plan-integration-followons/design.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/design.md diff --git a/openspec/changes/plan-integration-followons/proposal.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/proposal.md similarity index 100% rename from openspec/changes/plan-integration-followons/proposal.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/proposal.md diff --git a/openspec/changes/plan-integration-followons/specs/active-learning-decisions/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/active-learning-decisions/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/active-learning-decisions/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/active-learning-decisions/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/agent-adapters/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/agent-adapters/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/agent-adapters/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/agent-adapters/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/cli-surface/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/cli-surface/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/cli-surface/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/cli-surface/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/health-and-diagnostics/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/health-and-diagnostics/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/health-and-diagnostics/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/health-and-diagnostics/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/live-session-orchestration/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/live-session-orchestration/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/live-session-orchestration/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/live-session-orchestration/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/mcp-server/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/mcp-server/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/mcp-server/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/mcp-server/spec.md diff --git a/openspec/changes/plan-integration-followons/specs/web-ui/spec.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/specs/web-ui/spec.md similarity index 100% rename from openspec/changes/plan-integration-followons/specs/web-ui/spec.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/specs/web-ui/spec.md diff --git a/openspec/changes/plan-integration-followons/tasks.md b/openspec/changes/archive/2026-09-21-plan-integration-followons/tasks.md similarity index 93% rename from openspec/changes/plan-integration-followons/tasks.md rename to openspec/changes/archive/2026-09-21-plan-integration-followons/tasks.md index b568e0a97..9302c80f3 100644 --- a/openspec/changes/plan-integration-followons/tasks.md +++ b/openspec/changes/archive/2026-09-21-plan-integration-followons/tasks.md @@ -258,7 +258,7 @@ writer, through the existing gate, closes that. Kept out of item 3 so item 3's f ## Item 7 — push step (owner present; HANDOFF §3 item 7) -- [ ] **T7.1** Not run unattended. Ruleset D-I; tokens D-J. **State 2026-09-18** (HANDOFF §3 item 7's seven +- [x] **T7.1** Not run unattended. Ruleset D-I; tokens D-J. **State 2026-09-18** (HANDOFF §3 item 7's seven steps): (1) pushes — done, owner pushed `main` three times (#20 `46262d23`, #22 `4bba58b6`, #23 `a5b9f903`), each a fast-forward after CI green on the PR; (2) ruleset D-I — done three times (2026-09-17 `feat/knowledge-proof` + `fix/plan-integration-bugs`; 2026-09-18 `feat/plan-close` + @@ -270,7 +270,21 @@ writer, through the existing gate, closes that. Kept out of item 3 so item 3's f item-6 issues — done 2026-09-19: **#25** (D-D derived plan bias) and **#26** (D-E nudge + retire/snooze), both `ready-for-agent`, bodies drawn from the proposals with every cited test id, symbol and path checked against `main` `a03fc9bd`; #21 given a status comment from the harness receipt and **left open** (its DoD names OpenCode - in core; OpenCode stays preview on the `opencode.db` exporter and the no-completed-reply gaps). **Open:** - (5) the local tag - `archive/feat-clean-start-2026-09-15` — owner: push or discard; (6) the GitHub Support ticket text — owner; - (7) revoke both tokens and delete `~/tmp/.env` — owner (D-J). + in core; OpenCode stays preview on the `opencode.db` exporter and the no-completed-reply gaps). + **Steps 5–7 closed 2026-09-21** (owner walkthrough, each fact re-checked against the live system first; #10, + #15 and #7 closed the same morning after the owner pushed `main` to `96806feb`): (5) the tag — **pushed** by + the owner (`9ec38e72` on origin). Checked before asking: its tip `daf46c81` is on no branch, only the dotenv fix + was ever cherry-picked (`bfe0695c`), and `clean_rebuild.py`, `corpus.py`, their tests and the clean-start + cutover receipt exist nowhere else — the tooling that produced the live `sessions.db`; it was the only one of + seventeen archive tags not on origin. (6) the Support ticket — the symptom it was written for is **gone**: the + repo page's contributors fragment lists only `@claude` and `@NetDevAutomate`, and the contributors API shows + the owner alone. `refs/pull/1..6/head` still exist, all six orphaned (unreachable from any branch or tag) and + fetchable, their histories carrying 85 commits under the owner's work address that the consolidation removed + from every branch. Owner is sending the ticket **reworded to that reason** (delete the cached head refs; the + merged PRs stay). (7) the tokens — `~/tmp/.env` **no longer holds them**: rewritten 2026-09-19 into a Fabric + configuration (fourteen variables, neither `GITHUB_TOKEN` nor `AWS_BEARER_TOKEN_BEDROCK`), so the "delete + `~/tmp/.env`" step would have deleted the wrong file and revoked nothing — **not done, deliberately**. Owner + revoked both D-J tokens at source 2026-09-21. Found alongside: `gh` on this machine ran on a classic PAT with + `admin:enterprise`, `admin:org`, `delete_repo`, `repo`, `workflow` scopes — every agent shell acted with it; + replaced by a browser OAuth login (`repo`, `read:org`, `gist`, `admin:public_key`), verified via + `gh auth status`, before the PAT was revoked. diff --git a/openspec/changes/plan-integration-followons/.openspec.yaml b/openspec/changes/plan-integration-followons/.openspec.yaml deleted file mode 100644 index 9e71f9379..000000000 --- a/openspec/changes/plan-integration-followons/.openspec.yaml +++ /dev/null @@ -1,9 +0,0 @@ -schema: spec-driven -created: 2026-09-16 - -deferred: >- - Follow-on programme to the archived plan-application-seam change (owner - decisions D-A..D-J, docs/architecture/plan-integration/HANDOFF-2026-09-16.md). - Items 1-4 land under one council review, item 5 under its own review round - with a rubric row the owner scores before it ships; the change is archived - when item 5's row is scored, not before. Not part of the current release tag. diff --git a/openspec/specs/active-learning-decisions/spec.md b/openspec/specs/active-learning-decisions/spec.md index dd1d79f52..42bff139a 100644 --- a/openspec/specs/active-learning-decisions/spec.md +++ b/openspec/specs/active-learning-decisions/spec.md @@ -439,7 +439,8 @@ study actions and SHALL consume active plans through exactly one call to `PlanApplication().get_active_guidance(today=…)`, where `today` is the date of the same instant `generated_at` records. It SHALL apply these rules, in this order (the plan-application-seam design, §3; decision D-5 of its -council plan): +council plan; item 5 / D-F of the follow-ons for rule 2's repair half and +the body-doubling floor): 1. Candidates are collected as before; a failure to read plans at all SHALL degrade to a `warnings` entry, never a failed recommendation, and SHALL be @@ -447,8 +448,21 @@ council plan): programming error cannot hide behind the learner-facing warning. 2. The energy capability is `low|medium|high → 3|6|10`. For an active plan whose `energy_floor` exceeds it, the next milestone SHALL be listed in - `energy_deferred` and SHALL NOT become a candidate; plan-related due recall - and struggle repair stay eligible and plan-related. + `energy_deferred` and SHALL NOT become a candidate. Plan-related due recall + stays eligible and plan-related whatever its recorded confidence. A + struggle **repair** carries an energy demand of its own, derived once in + the struggle collector from its own row classes and carried in the + candidate's `metadata["energy_demand"]`: `struggling` seen within 14 days → + `high` (asks for 6/10); `struggling` older than that, or a row whose only + signal is a weak teach-back → `medium` (4/10); `learning` → `low` (0/10). + A repair whose demand exceeds the capability SHALL be deferred exactly + like new milestone work — listed in `energy_deferred_repairs` as a + `DeferredRepair` (`plan_id`/`plan_title` when plan-related, else `None`, + `concept`, `topic`, `confidence`, `energy_demand`, `required_capability`, + `energy_capability`, `reason` naming the struggle) and never ranked; the + deferral does not depend on a plan existing. A `low`-demand repair is + always carried. `energy_deferred` stays milestone-shaped; a repair is never + folded into it. 3. A candidate is plan-related when `normalise_match_key` of its concept, topic or course **equals** one of the plan's `match_keys`; no substring test. It names the plan's next milestone (`milestone_index`) only when the @@ -466,7 +480,28 @@ council plan): for it (source `study_plan::`, concept = the milestone's first concept or its title, topic = the plan's first topic), scored below every due and repair class. A learner with an active plan and no evidence - is therefore sent to the plan, and `starter` is `false`. + is therefore sent to the plan, and `starter` is `false`. A deferred repair + (rule 2) does not "represent" a milestone. When, after rules 2 and 5, **no + candidate is plan-related** and at least one matchable active plan exists, + one **body-double** candidate SHALL be synthesised instead of leaving the + plan to the least-bad task: `source = "body_double"`, `action_type = + "conversation"`, base score below every real candidate's base and then + scored like any other candidate — so every due and conversation candidate + outranks it at every energy while a hands-on task the low-energy rule + penalises may not; a proposal, never a filter: nothing is removed from the + ranking — `plan_refs` `(plan_id, None)` for every **ready** matchable plan + (rule 7 may add a reference to an unready plan whose topic the proposal + shares; the proposal itself names ready plans only), a reason that opens with + the progress the plan records (`N of M milestones of done.`, omitted when + none is done — never invented) before naming the deferred milestones and + repairs it stands in for, and describes the door by what the co-study persona + guarantees (the learner drives; the companion stays quiet unless asked), and + `evidence_command` the + co-study session door — `studyloop study "" --mode co-study` — + set explicitly, never a progress write. No active plan (a draft is not + one) → no body-double candidate; when every real candidate was deferred + and no plan exists, the starter stands in and its reason says the energy + deferred the repair work, not that no evidence exists. 6. After de-duplication every matching `PlanRef(plan_id, milestone_index)` SHALL be attached to each ranked action, ordered by target urgency (`overdue`, `soon`, `later`, `undated`) → most recent `updated` → `plan_id`, @@ -474,6 +509,9 @@ council plan): 7. When primary + alternates hold no plan-backed action and an eligible one whose estimate fits the requested time exists further down, it SHALL replace the last alternate only; the primary is never re-ranked by plans. + Below a plan's floor that plan-backed action is the body-double proposal + (it advertises no work the energy cannot carry), never the deferred + milestone. 8. A fully-checked active plan SHALL appear in `completion_actions` and SHALL be neither matched nor synthesised. An active-but-unready plan SHALL be listed and matched (bias and a `milestone_index = None` reference) but @@ -481,25 +519,38 @@ council plan): its blockers. `NowPlan` gains `active_plans` (ordered as rule 6), `energy_deferred`, -`completion_actions` and `warnings`; `LearningRecommendation` gains -`plan_refs: tuple[PlanRef, ...] = ()`. `to_json_dict()` SHALL omit each of -these when empty, so a learner with no active plan receives the pre-#10 -payload **byte for byte** — pinned by `tests/golden/now_plan_no_active.json`, -captured before any of this shipped. Renderers (`studyloop now`, `GET -/api/now`, the Today card, the daily recap in its JSON, spoken and Rich-panel -forms) SHALL show plan relevance, energy deferral and the engine's warnings -from these fields, SHALL escape learner-authored text before any markup -(Rich or HTML), and SHALL NOT re-rank. Ranking tests prove -ranking compliance, not learner benefit (D-16); a five-scenario human rubric -receipt accompanies the change. +`energy_deferred_repairs`, `completion_actions` and `warnings`; +`LearningRecommendation` gains `plan_refs: tuple[PlanRef, ...] = ()`. +`to_json_dict()` SHALL omit each of these when empty, so a learner with no +active plan, no struggle candidate and nothing deferred receives the pre-#10 +payload **byte for byte** — the golden world, pinned by +`tests/golden/now_plan_no_active.json`, captured before any of this shipped. +Exactly two changes are plan-independent (rule 2's repair half): every +struggle-collector candidate's `metadata` carries `energy_demand` at every +energy, and repair above the day's capability — a live struggle, an older +struggle or a weak teach-back at low energy — is deferred into +`energy_deferred_repairs` (with the starter, if nothing else was collected) +where the learner used to receive the repair itself. `medium` and `high` +demand both need at least medium self-reported energy today; the class is +carried so the payload says why. +Renderers (`studyloop now`, `GET /api/now`, the Today card, the daily recap in +its JSON, spoken and Rich-panel forms) SHALL show plan relevance, energy +deferral — one line per deferred milestone **and** one per deferred repair — +and the engine's warnings from these fields, SHALL label a body-double +primary's command as the session door it is ("Sit with the plan", and the +Today card starts it in the Body Double view) rather than as evidence to +record, SHALL escape learner-authored text before any markup (Rich or HTML), +and SHALL NOT re-rank. Ranking tests prove ranking compliance, not learner +benefit (D-16); a five-scenario human rubric receipt accompanies the change, +with row 3b re-run after this requirement's repair half. #### Scenario: No active plan is byte-identical to the golden - **WHEN** no active plan exists (an empty plans directory, or only a draft) and `build_now_plan()` runs with a frozen clock in an empty world - **THEN** the serialised `to_json_dict()` equals `tests/golden/now_plan_no_active.json` byte for byte, and no - `active_plans`, `energy_deferred`, `completion_actions`, `warnings` or - `plan_refs` key is present + `active_plans`, `energy_deferred`, `energy_deferred_repairs`, + `completion_actions`, `warnings` or `plan_refs` key is present #### Scenario: Matching due concept outranks unrelated of the same urgency - **WHEN** an active plan's milestone names `window function` and two due @@ -515,12 +566,64 @@ receipt accompanies the change. #### Scenario: Energy below the floor defers the milestone, keeps repair - **WHEN** energy is `low` (3/10), the plan's `energy_floor` is 5, its next - milestone is `Frames` and a struggle repair on a finished milestone's - concept is collected -- **THEN** the repair is primary with `PlanRef(plan, None)`, - `energy_deferred` names `(plan, 1, 5, 3)`, and no `study_plan:` candidate - exists; at `medium` energy nothing is deferred and the milestone is - synthesised + milestone is `Frames` and a `learning` row on a finished milestone's + concept is collected by the struggle collector (gentle repair) +- **THEN** that repair is primary (`teachback`, `energy_demand == "low"`) with + `PlanRef(plan, None)`, `energy_deferred` names `(plan, 1, 5, 3)`, + `energy_deferred_repairs` is empty, and no `study_plan:` or `body_double` + candidate exists; at `medium` energy nothing is deferred and the milestone + is synthesised + +#### Scenario: A live struggle's repair defers at low energy like new work +- **WHEN** energy is `low` and the struggle collector holds a `struggling` row + seen 3 days ago on a plan concept, a `struggling` row seen 20 days ago on + another, a `struggling` row seen 1 day ago unrelated to any plan, and a + `confident` row kept only for a teach-back score of 9 +- **THEN** none of the four is ranked; `energy_deferred_repairs` names all + four — the live plan-related one `("high", 6, 3)` with the plan's id and + title, the 20-day one `medium` (4), the weak-teach-back one `medium`, the + unrelated one with `plan_id` and `plan_title` `None` — `energy_deferred` + still names the milestone alone, and at `medium` energy the key is absent + and the live repair is ranked again + +#### Scenario: Due recall is never deferred +- **WHEN** energy is `low`, a due `recall` item on a plan concept is collected + (whatever its metadata says about confidence) and a live struggle repair is + also collected +- **THEN** the due item is primary with `PlanRef(plan, None)`, the repair is + in `energy_deferred_repairs`, and no `body_double` candidate exists + +#### Scenario: A struggling row's due item is the repair, collected twice +- **WHEN** energy is `low` and one `struggling` row seen 3 days ago on a + finished milestone's concept is read by both the due-progress collector + (which labels every `struggling` row due — "Guided repair + tiny practice", + `hands-on`) and the struggle collector +- **THEN** both copies are deferred, `energy_deferred_repairs` names the + concept once (`high`), the concept is ranked nowhere, and the primary is the + `body_double` proposal (the starter when no plan exists); at `medium` energy + the due copy is primary as before and the key is absent + +#### Scenario: Nothing plan-related fits, so the engine proposes sitting with the plan +- **WHEN** energy is `low`, the plan's `energy_floor` is 5 (milestone + deferred) and its only repair is a live struggle (deferred) +- **THEN** the primary is `source == "body_double"`, `action_type == + "conversation"`, `plan_refs == (PlanRef(plan, None),)`, `evidence_command == + 'studyloop study "" --mode co-study'`, its reason names the + deferred milestone and the deferred repair, `starter` is `false`, and the + JSON keys are the golden's then `active_plans`, `energy_deferred`, + `energy_deferred_repairs` + +#### Scenario: The body-double candidate is a proposal, not a filter +- **WHEN** the same world also collects an unrelated due item +- **THEN** the due item is primary with no refs and the body-double + candidate is the only alternate, with a lower score + +#### Scenario: No active plan, no body double +- **WHEN** no plan document exists (or only a draft) and a live unrelated + struggle is collected at `low` energy +- **THEN** no `body_double` candidate exists, `energy_deferred_repairs` names + the struggle with `plan_id None`, `starter` is `true` and the starter's + reason says the energy deferred the repair work #### Scenario: A deferred milestone is never named by a reference - **WHEN** energy is `low`, the plan's `energy_floor` is 5 and the only @@ -555,8 +658,9 @@ receipt accompanies the change. `energy_floor` is 5 - **THEN** at `medium` energy the synthesised milestone replaces the second alternate (the primary and first alternate are unchanged); at `low` energy - the alternates are the unrelated items and `energy_deferred` names the - milestone + the primary and first alternate are the two best unrelated items, the + second alternate is the body-double proposal with `PlanRef(plan, None)`, + no `study_plan:` candidate exists and `energy_deferred` names the milestone #### Scenario: Fully-checked plan emits a completion action - **WHEN** an active plan's every milestone is done and an unrelated due item @@ -567,11 +671,16 @@ receipt accompanies the change. #### Scenario: Renderers show, never re-rank - **WHEN** `studyloop now --energy low`, `GET /api/now?energy=low` and the - daily recap run against the energy-deferral fixture -- **THEN** each names the primary the engine chose, the plan it advances, and - the deferred milestone; with no plan the CLI panel prints no plan lines, - `GET /api/now` equals the golden, and the recap's `plan_context` is absent - from its JSON, its spoken text and the `recap today` panel + daily recap run against the energy-deferral fixture, and against the + live-struggle fixture +- **THEN** each names the primary the engine chose, the plan it advances, the + deferred milestone and — for the live-struggle fixture — one line per + deferred repair with its demand and the day's capability; a body-double + primary is labelled "Sit with the plan" with its `--mode co-study` door + and never "Record evidence"; with no plan and no struggle candidate the CLI + panel prints no plan lines, `GET /api/now` equals the golden, and the + recap's `plan_context` is absent from its JSON, its spoken text and the + `recap today` panel #### Scenario: Learner-authored text is data to every renderer - **WHEN** an active plan's title, topic or milestone text contains Rich @@ -621,3 +730,100 @@ field change SHALL still be one save, and an empty revision remains a is applied - **THEN** `InvalidField` carrying that message is raised and no record is added + +### Requirement: The completion action is a closing review, never a verdict +Rule 9's completion action for a fully-checked active plan (item 4 / D-G) +SHALL be composed from the plan's **end assessment**, read through the preview +path — `PlanApplication().assess(AssessPlan(plan_id, phase="end", +record=False))` — exactly once per fully-checked plan per `build_now_plan`. The +read SHALL write nothing: the document's bytes and status, the plans directory +and the checkpoint log are unchanged, and the recording writers +(`evaluate_and_record`, `record_checkpoint`) are never called. The engine +proposes; the architect asks; the learner decides; `set_study_plan_status` +remains the only door to `complete`. + +`CompletionAction` SHALL gain `due_reviews: int`, `struggles: int`, +`unverified_milestones: int`, `proposal: Literal["extend", "close"] | None` +and `evidence: tuple[str, ...]`, and SHALL keep `action`, the sentence every +renderer prints — now naming the proposal and the three counts and the one +door to acting on them, `studyloop plan close `; it SHALL differ from the +pre-change either-way sentence. The counts and lines SHALL come from one +definition, `planning.views.CompletionReview.from_evaluation`, consumed by both +this action and the `plan close` brief so the two surfaces never disagree: +`proposal == "extend"` iff any count is above zero, else `"close"`; one +evidence line per counted item, capped at `COMPLETION_EVIDENCE_CAP` (8) with a +final `… and N more` line. **Due reviews SHALL count only rows that name a +concept** (owner decision, 2026-09-17): the scheduler's `New topic -- start +fresh` row (`concept: None`, `evidence: configured_topic`) is a cold-start hint +for "what should I review now", not a lapsed review, and SHALL NOT be counted; +`plan evaluate` keeps the row, the exclusion is the completion review's. +`CompletionAction` and `CompletionReview` SHALL carry `partial: bool`. + +**A partial read SHALL NOT propose** (council review 6, F1). `evaluate_plan` +turns a reader that fails into a warning ending `unavailable — evaluation is +partial` (`evaluation.PARTIAL_READ_MARKER`, one definition) and an empty +default, so a count read while that reader was down is unread, not zero. When +the evaluation carries such a warning the review SHALL keep the counts it did +read, set `partial` true, set `proposal` `None` — neither `close` (a clean +slate is a fact about evidence, not its absence) nor `extend` — and name each +gap among its evidence lines (`Not read: unavailable — …`); the +sentence SHALL say the review is partial and could not propose, never "clean"; +the `plan close` brief's proposal line SHALL read `unassessed — the review is +partial` and its status line SHALL NOT say the review proposes. + +When the assessment fails, the recommendation SHALL NOT fail: the action +SHALL keep the plan-static sentence with `proposal` `None`, the counts `0` and +`evidence` empty, and `NowPlan.warnings` SHALL carry one entry naming the plan +and the failure, logged with its traceback first — so no renderer reads a +clean slate or outstanding work into a failure. The evaluation's own data-gap +warnings SHALL travel back into `warnings` prefixed with the plan id. + +The new keys SHALL appear only inside `completion_actions` entries, which +exist only when a fully-checked active plan exists; the no-plan payload stays +byte-identical to `tests/golden/now_plan_no_active.json`. Renderers SHALL +show the sentence (CLI `now`, the Today card, the daily recap), the CLI SHALL +print each evidence line beneath it, and none SHALL re-rank. + +#### Scenario: Due work on the plan's concepts proposes extend +- **WHEN** an active plan's every milestone is done and the end assessment + finds one due review on one of its concepts +- **THEN** `completion_actions[0]` carries `(due_reviews, struggles, + unverified_milestones) == (1, 0, 0)`, `proposal == "extend"`, an evidence + line naming the concept, and a sentence naming the plan and `extend`; the + JSON entry carries all five keys; no `study_plan:` candidate exists + +#### Scenario: A clean assessment proposes close +- **WHEN** the end assessment finds no due reviews, no struggles and every + done milestone backed by evidence +- **THEN** the counts are `(0, 0, 0)`, `proposal == "close"`, `evidence` is + empty and the sentence names `close` + +#### Scenario: New-topic rows are not due +- **WHEN** `spaced_repetition_due` returns only the `New topic -- start + fresh` row (`concept: None`) for the plan's topic and the concepts have + session mentions +- **THEN** `due_reviews == 0` and `proposal == "close"` + +#### Scenario: The ranker never changes a status +- **WHEN** `build_now_plan` runs against a fully-checked active plan with the + recording writers patched to raise +- **THEN** exactly one `AssessPlan(plan_id, "end", record=False)` intent is + assessed, the document's bytes are unchanged, the status is still `active` + and the checkpoint history is empty + +#### Scenario: A failed assessment keeps the sentence and warns +- **WHEN** `assess` raises for the fully-checked plan +- **THEN** `completion_actions[0].action` equals the pre-change sentence, + `proposal is None`, `warnings` names the plan and the failure, and the + primary is still the collected due item + +#### Scenario: A partial assessment never proposes a clean close +- **WHEN** one of the end assessment's history readers raises inside the + evaluation and every other reader finds nothing outstanding +- **THEN** `completion_actions[0]` carries `proposal is None`, + `partial is True`, the counts `(0, 0, 0)`, an evidence line beginning + `Not read:`, a sentence that says the review is partial and never "clean" or + "closing the plan", and `warnings` names the plan and the unavailable reader; + `plan close ` still launches the architect, its brief's fourth line is + `Proposal: unassessed — the review is partial`, the gap is among the first + section's lines, and its status line does not say the review proposes diff --git a/openspec/specs/agent-adapters/spec.md b/openspec/specs/agent-adapters/spec.md index 12c0f4318..687b8268a 100644 --- a/openspec/specs/agent-adapters/spec.md +++ b/openspec/specs/agent-adapters/spec.md @@ -273,3 +273,47 @@ and the `focus` persona SHALL be byte-identical before and after this change. - **THEN** each body after its harness header equals the canonical persona byte-for-byte, and `agents/manifest.json` carries the generator's own hash for every architect projection it tracks + +### Requirement: Harness-launched architects carry the plan tools +The `study-plan-architect` definitions for Kiro CLI (`agents/kiro/study-plan-architect.json`) and Claude +Code (`agents/claude/study-plan-architect.md`) SHALL attach the `studyloop` MCP server and allow exactly the +nine plan lifecycle tools named by `studyloop.mcp.inventory.PLAN_TOOL_NAMES` plus +`studyloop.mcp.inventory.LEARNING_RECORD_TOOL` — and no other tool of that server — in the spelling the +harness honours (owner decision D-A, 2026-09-16: no harness-launched architect falls back to the CLI with full +shell permissions; receipt `docs/architecture/plan-integration/receipts/kiro-agent-tools-probe-2026-09-16.md`). + +For Kiro, visibility and trust are two arrays: `tools` SHALL contain `@builtin`, `@studyloop` and +`@session-db`; `mcpServers` SHALL declare `studyloop` (`studyloop-mcp`) and `session-db` (`session-db-mcp`); +`allowedTools` SHALL carry `@studyloop/` for exactly the ten and SHALL NOT carry a bare `@studyloop` +(which would trust every tool of the server) nor any `mcp__` entry (inert in an agent config). +The `session-db` server is visible for the loaded `shared/session-protocol.md` session-start step and is not +trusted: its tools prompt. For Claude Code, the frontmatter `tools:` allow-list SHALL carry +`mcp__studyloop__` for exactly the ten and no other `mcp__` entry, alongside its built-in tools. + +`agents/kiro/study-mentor.json` SHALL use the same spelling: `@studyloop` in `tools` and every MCP grant in +`allowedTools` as `@/`. `agents/manifest.json` SHALL carry the generator's hash for each +edited definition, and the persona body of every architect projection SHALL remain byte-identical to the +canonical persona (the frontmatter and JSON header are the only edits). Learner confirmation before +`delete_study_plan` remains a persona rule: tool permission is not user authorisation. + +#### Scenario: Kiro architect definition carries the server and exactly the plan tools +- **WHEN** `agents/kiro/study-plan-architect.json` is parsed +- **THEN** `mcpServers["studyloop"]["command"] == "studyloop-mcp"` and `mcpServers["session-db"]["command"] + == "session-db-mcp"`, `tools` contains `@builtin`, `@studyloop` and `@session-db`, and the set of + `allowedTools` entries starting with `@studyloop` equals `{f"@studyloop/{name}" for name in + PLAN_TOOL_NAMES + (LEARNING_RECORD_TOOL,)}`, with no bare `@studyloop` and no entry starting with `mcp_` + +#### Scenario: Claude architect allow-list is exactly the plan tools +- **WHEN** the frontmatter of `agents/claude/study-plan-architect.md` is parsed +- **THEN** the `mcp__` entries of `tools:` equal `{f"mcp__studyloop__{name}" ...}` for exactly the ten names, + and the body after the frontmatter is byte-identical to `agents/shared/personas/plan-architect.md` + +#### Scenario: Installed Kiro architect resolves its prompt and its servers +- **WHEN** `studyloop install agents --tool kiro` has run into a sandboxed home +- **THEN** `~/.kiro/agents/study-plan-architect.json` is a symlink whose `prompt` `file://` URI resolves to a + file, whose `mcpServers` names `studyloop`, and whose stop hook is `session-export --kiro-only` + +#### Scenario: Mentor grants use the spelling the CLI honours +- **WHEN** `agents/kiro/study-mentor.json` is parsed +- **THEN** `tools` contains `@studyloop`, no `allowedTools` entry starts with `mcp_`, and each MCP grant is + `@/` for a server declared in `mcpServers` diff --git a/openspec/specs/cli-surface/spec.md b/openspec/specs/cli-surface/spec.md index 9e4c4dcdd..2099ab594 100644 --- a/openspec/specs/cli-surface/spec.md +++ b/openspec/specs/cli-surface/spec.md @@ -341,3 +341,138 @@ seam's `Invalid value: …` refusal, exit `1`. the command's `RevisePlan` runs - **THEN** the command exits `0` with `created: false, number: 1`, made no `inspect` call, and the plan holds one record + +### Requirement: The plan CLI discovers husks and repairs them through the one launch chain +An active plan that is not ready — a "husk" (item 3 / D-C; deviation 12 kept) +— refuses every write until it is repaired or paused, and the CLI SHALL let +the learner find one before they trip over the refusal. `studyloop plan list` +SHALL mark a husk with `!` after its status in the Rich table and nothing +after any other status; `--husks` SHALL list only husks (`PlanApplication.husks()`, +read-only, storage-pinned identity, `browse` order); every `--json` row SHALL +carry `ready` as its eighteenth key. `PlanSummary.ready` and +`StudyPlan.summary()["ready"]` SHALL agree, so the D-3 legacy-dict pin holds +with the contract grown by one key on both sides. + +`studyloop plan repair ` SHALL be the architect launch and never a second +path: it SHALL `_inspect(id)` (unknown id → the seam's not-found through +`_fail_for`, exit `1`), SHALL exit `0` with `Nothing to repair on ''` and +no launch for a ready plan, SHALL exit `0` with no launch and a pointer to +`studyloop plan architect` for a plan that is not active (a draft or a paused +plan is unready by nature, not a husk), and for a husk SHALL `ctx.invoke(study, +…, mode="plan-architect", topic=, brief=…, brief_intro=…)`. +`brief` and `brief_intro` SHALL be plain keywords on `study()` — not click +options — threaded `study → _handle_start → start_session → +build_canonical_persona`. The brief's first section SHALL be +`### Repair: what this plan is missing` listing exactly `readiness.blockers` +as `- ` lines and nothing else, followed by the plan as it stands (title, id, +status, topics, milestones done/total, created) and one provenance sentence: +`creation stamp predates the readiness gate` (and `cannot tell when it became +incomplete`) only when `created` parses as a date before +`READINESS_GATE_DATE`; otherwise `cannot tell how it got that way`; never +`never judged` (council review 6 F4). The +sentence SHALL never claim a hand edit. The `brief_intro` SHALL say `PLAN +REPAIR` and `ask the learner only for what is missing`; the default intro +(`None`) SHALL keep the planning sentence byte-for-byte so the Web door's +`persona_hash` does not move. The command itself SHALL write nothing: the +document, the plans directory and the checkpoint log are unchanged after it +returns. + +The refusal a husk write meets (`_refuse_activation(already_active=True)`) +SHALL name both exits: `studyloop plan repair ` and `studyloop plan status + paused`. + +#### Scenario: plan list marks the husk, filters to it, and every JSON row carries ready +- **WHEN** one ready active plan, one draft and one active document with no + mission exist and `plan list`, `plan list --json`, `plan list --husks` and + `plan list --husks --json` are run +- **THEN** the table's Status cell reads `active !` for the husk and `active` / + `draft` for the others; every JSON row has 18 keys with `ready` `true` / + `false` / `false`; `--husks` lists only the husk in both forms + +#### Scenario: plan repair on a husk launches once with the blockers first and creates nothing +- **WHEN** `plan repair husk` is run on an active document with no mission, + created before the gate date +- **THEN** exactly one `start_session` call is made with `mode="plan-architect"` + and `topic` equal to the plan's title; the brief's first section lists + exactly the two mission blockers; the brief names the title, status, + topics and `0/1` milestones and says `predates the readiness gate`; the + intro says `PLAN REPAIR` and not `build a study plan`; the plans directory + and the checkpoint history are unchanged + +#### Scenario: plan repair is honest when provenance is unknown +- **WHEN** `plan repair husk` is run on a husk whose `created` is after the + gate date +- **THEN** the brief says `cannot tell how it got that way`, does not say + `predates the readiness gate`, and does not say `hand edit` + +#### Scenario: Nothing to repair, unknown id, refusal names both exits +- **WHEN** `plan repair glue-etl` is run on a ready active plan; `plan repair + nope` on no such plan; and `plan evaluate husk --record` on a husk +- **THEN** the first exits `0` with `Nothing to repair on 'glue-etl'` and no + launch; the second exits `1` naming `nope` with no traceback; the third + exits `1` and names both `studyloop plan status husk paused` and `studyloop + plan repair husk` + +### Requirement: The plan CLI closes a fully-checked plan through the one launch chain, consensually +`studyloop plan close ` (item 4 / D-G) SHALL be the architect launch and +never a second path — the sibling of `plan repair`, through the same +`ctx.invoke(study, …, mode="plan-architect", topic=, +brief=…, brief_intro=…)`. It SHALL `_inspect(id)` (unknown id → the seam's +not-found through `_fail_for`, exit `1`); SHALL exit `1` with `'' still +has N open milestone(s)` and no launch while any milestone is open; SHALL +exit `1` with a pointer to `studyloop plan architect` for a plan with no +milestones; SHALL exit `0` with no launch for a plan that is already +`complete`; and for a fully-checked plan of any other status SHALL run the end +assessment as a **preview** (`AssessPlan(phase="end", record=False)`) and +launch once. A fully-checked `draft`, `paused` or `abandoned` plan is reviewed +like an active one (owner decision 2026-09-18, council review 6 open item 2): +the learner may close it or delete it, and the architect asks which — the +brief's closing section SHALL end with a `Status: — not active; ask +the learner whether to close it (…) or delete it (…)` line naming both doors, +`set_study_plan_status(plan_id, "complete")` and `delete_study_plan(plan_id, +confirmed=True)`, and the status line SHALL name the status. The +command itself SHALL write nothing: the document, the plans directory, the +plan's status and the checkpoint log are unchanged after it returns; the +status moves to `complete` only when the learner agrees in the launched +session and the architect calls `set_study_plan_status`. + +The brief's first section SHALL be `### Closing review`, whose first four +`- ` lines are `Due reviews on plan concepts: N`, `Struggles on plan +concepts: N`, `Unverified milestones: N` and `Proposal: extend|close` — or +`Proposal: unassessed — the review is partial` when a reader was unavailable +(council review 6 F1) — followed by one `- ` evidence line per counted item +and one `Not read: …` line per unavailable reader — the same +`CompletionReview` the `now` engine puts on its completion action, so the two +never disagree on a count (new-topic rows excluded) — then, for a non-active +plan, the `Status:` line above, then the plan as it stands (title, id, status, +topics, milestones done/total, created), every learner-authored value one +line (`planning.one_line`, council review 6 F2). The `brief_intro` SHALL say +`CLOSING REVIEW` and `only when the learner agrees`, and SHALL NOT say `build a +study plan`. + +#### Scenario: plan close on a fully-checked plan launches once with the review first and writes nothing +- **WHEN** `plan close glue-etl` is run on an active plan whose two milestones + are both done, with the due reader returning one real due row on a plan + concept and one `New topic -- start fresh` row (`concept: None`) +- **THEN** exactly one `start_session` call is made with `mode="plan-architect"` + and `topic` equal to the plan's title; the `### Closing review` section's + first four lines are `Due reviews on plan concepts: 1`, `Struggles on plan + concepts: 0`, `Unverified milestones: 0`, `Proposal: extend`, followed by an + evidence line naming the due concept; the brief says `2/2`; the intro says + `CLOSING REVIEW` and `only when the learner agrees` and not `build a study + plan`; the plans directory, the plan's `active` status and the checkpoint + history are unchanged + +#### Scenario: plan close on a fully-checked non-active plan launches and names both doors +- **WHEN** `plan close glue-etl` is run on a plan whose two milestones are both + done and whose status is `abandoned`, `paused` or `draft` +- **THEN** exactly one launch is made; the status line names the status; the + `### Closing review` section's last line begins `Status: ` and names + `set_study_plan_status`, `delete_study_plan` and asking the learner; the + plans directory and the plan's status are unchanged + +#### Scenario: plan close on an unfinished plan refuses without launching +- **WHEN** `plan close glue-etl` is run on an active plan with two open + milestones +- **THEN** it exits `1` with `'glue-etl' still has 2 open milestone(s)`, no + traceback, no launch, and the plans directory unchanged diff --git a/openspec/specs/health-and-diagnostics/spec.md b/openspec/specs/health-and-diagnostics/spec.md index 6c0c31fe7..a07eb7f99 100644 --- a/openspec/specs/health-and-diagnostics/spec.md +++ b/openspec/specs/health-and-diagnostics/spec.md @@ -190,3 +190,58 @@ acceptable), and `studyloop doctor --json` emits valid JSON with exit code 2 - **THEN** `install.sh` prints an error and exits non-zero, halting the installation + +### Requirement: doctor names each active-but-unready study plan +`check_study_plans()` (`cli/_doctor.py`, beside `check_unknown_config_keys`) +SHALL be registered under the existing `config` category — the category set +is enumerated verbatim elsewhere in this spec and gains none here — and SHALL +report on the plans `PlanApplication.husks()` returns: active plans that are +not ready and therefore refuse every write (item 3 / D-C; deviation 12 kept). +It SHALL emit one `warn` row per husk with `name="study_plans"`, +`fix_auto=False` (the repair is a conversation with the architect, not a +script), a message naming the plan id, its title, the exact +`ReadinessView.blockers`, and the shared provenance sentence +(`husk_provenance`: `This plan's creation stamp predates the readiness gate +(); the seam cannot tell when it became incomplete` only +for a `created` that parses as a date before `READINESS_GATE_DATE`, else +`cannot tell how it got that way`; never `hand edit`, never `never judged` — +a stamp establishes when the document was created, not what judged it or +when it became incomplete; council review 6 F4), and a `fix_hint` naming both +exits: `studyloop plan repair (or: studyloop plan status paused)`. +When every active plan is ready it SHALL emit one `pass` row counting the +active plans and saying they are ready *as they stand* — never a promise +about writes that have not happened (`will pass`, `every write`); when no +plan is active it SHALL emit one `info` row, not a warning. A draft is +unready by nature and is never reported. A plans directory that cannot be +read SHALL be one `warn` row, not a crash of doctor; a single document the +seam cannot read (`PlanApplication.survey_husks().unreadable`: a filename +that is not a valid id, a file that is not UTF-8) SHALL be its own `warn` row +naming the id beside the readiness rows, so a parse failure never reads as +health (F4). + +#### Scenario: Two husks, one ready active plan, one draft +- **WHEN** `check_study_plans()` runs over a ready active plan, a draft with + no mission, a husk created before the gate date and a husk created after it +- **THEN** exactly two `warn` rows are returned, both `config` / + `study_plans` / `fix_auto=False`; each names its plan id and title and + both mission blockers; the older one says `predates the readiness gate` + and the newer says `cannot tell how it got that way` and not `hand edit`; + each `fix_hint` names `studyloop plan repair ` and `studyloop plan + status paused`; neither the ready plan nor the draft is named + +#### Scenario: All active plans ready is one pass row; no plans is info +- **WHEN** `check_study_plans()` runs with one ready active plan, and again + with no plans at all +- **THEN** the first returns one `pass` row saying `1 active plan` and + `ready` and neither `will pass` nor `every write`; the second returns one + `info` row + +#### Scenario: An unreadable document is named, not hidden behind all-ready +- **WHEN** `check_study_plans()` runs over one ready active plan and one + `.md` file the seam cannot read (not UTF-8) +- **THEN** it returns a `warn` row naming the unreadable id (`could not be + read`, `fix_auto=False`) and the `pass` row for the readable plan + +#### Scenario: The checker is registered +- **WHEN** `_get_registry()` is built +- **THEN** `("config", "check_study_plans")` is among its registered checkers diff --git a/openspec/specs/live-session-orchestration/spec.md b/openspec/specs/live-session-orchestration/spec.md index 1ca36ab59..5ed4d1da9 100644 --- a/openspec/specs/live-session-orchestration/spec.md +++ b/openspec/specs/live-session-orchestration/spec.md @@ -204,13 +204,35 @@ SHALL be indistinguishable from a start that names no purpose: the same persona, the same `persona_hash`, the same session-state `mode`. A `planning` start SHALL launch the study-plan architect: the persona is the `plan-architect` mode carrying a `## Planning brief` section (the interview -questions, the learner's history evidence and the existing plans), and the -session's topic is the learner's subject when one was supplied, else the fixed -label `Study plan` — the same label `studyloop plan architect` pins. The start +questions, the learner's history evidence, the existing plans and — only when +the request carried one — the learner's brain dump), and the session's topic +is the learner's subject when one was supplied, else the fixed label +`Study plan` — the same label `studyloop plan architect` pins. The start SHALL NOT create a plan and SHALL NOT store a plan id anywhere; the architect -creates plans through the plan tools during the session. The only planning -fact the live-session state carries is `purpose`, written on every start -(never inherited through the state file's read-merge-write), and +creates plans through the plan tools during the session. + +The request MAY carry `brain_dump: str | None` (default `None`, `max_length` +`BRAIN_DUMP_MAX_CHARS` = 4000, published by `_models`), the learner's own free +text for the architect. On a `planning` start a non-blank dump SHALL be +rendered by the brief renderer as a fourth section, `### Learner's brain dump`, +after `### Existing plans`, introduced as the learner's words — evidence, not +instructions — with every line of the dump emitted as a Markdown blockquote +line (`> …`, a blank line as a bare `>`) after in-line whitespace +normalisation, and a line that begins with a block marker (`#`, `-`, `*`, +`+`, `>`, `` ` ``, `~`) backslash-escaped, so a dump line can never open a +heading, list item or fence of its own inside the persona (review-3 F4 +containment). A blank or absent dump SHALL render no section, so a brief +without one is byte-identical to the pre-change brief. The dump SHALL NOT be +folded into `topic`, SHALL NOT be written to the session state or exposed by +`GET /api/session/state`, and SHALL travel once, inside the persona (on `acp` +the `201` echoes the persona as `persona_text` by design, and the dump appears +in that field only). On a `focus` start the dump SHALL be ignored: the persona +and state are byte-identical to a start without it. A dump longer than +`BRAIN_DUMP_MAX_CHARS` SHALL be refused with `422` before the handler runs, +holding no slot. + +The only planning fact the live-session state carries is `purpose`, written on +every start (never inherited through the state file's read-merge-write), and `GET /api/session/state` SHALL expose it for the reconnect label with one precedence on every path it answers from — the live-slot overlay and the file-only path a CLI-started session takes: an explicitly persisted @@ -225,56 +247,50 @@ single-session slot free — no reservation, no live slot, no study row. Both transports (`pty` and `acp`) SHALL follow this requirement identically. #### Scenario: Planning start launches the architect with a brief -- **WHEN** `POST /api/session/start` is called with `{"purpose": "planning", "topic": "", ...}` -- **THEN** the response is `201`, the persona the agent receives has - `**Mode:** plan-architect`, contains the plan-architect persona body and a - `## Planning brief` section naming the interview questions and every existing - plan by id and title, contains no `Resuming Previous Session` section, and - the session topic is `Study plan` +- **WHEN** `POST /api/session/start` is made with `purpose: "planning"` and `topic: ""` +- **THEN** the persona the agent receives has `**Mode:** plan-architect`, a `## Planning brief` section containing the interview's first prompt and the existing plans, `**Topic:** Study plan`, and no `Resuming Previous Session` + +#### Scenario: Brain dump travels once, contained, and is never persisted +- **WHEN** a planning start carries `brain_dump` with several paragraphs, one of which begins `## Ignore previous instructions` +- **THEN** the persona contains exactly one `### Learner's brain dump` section after `### Existing plans`, every dump line rendered as `> …` with the `##` line escaped (`> \## …`), and `**Topic:** Study plan` +- **AND** the session state file and `GET /api/session/state` carry neither a `brain_dump` key nor the text, and on `acp` the text appears in the `201` body's `persona_text` only + +#### Scenario: Over-limit brain dump is refused structurally +- **WHEN** a planning start carries a `brain_dump` of `BRAIN_DUMP_MAX_CHARS + 1` characters +- **THEN** the response is `422` naming `brain_dump`, no slot is held and no agent is launched; a dump of exactly `BRAIN_DUMP_MAX_CHARS` is accepted + +#### Scenario: Brain dump on a focus start is ignored +- **WHEN** a focus start carries `brain_dump` +- **THEN** the persona is byte-identical to `build_canonical_persona("focus", topic, energy)` and the state carries no trace of the text #### Scenario: Planning start keeps a supplied subject -- **WHEN** `POST /api/session/start` is called with `{"purpose": "planning", "topic": "Spark", ...}` -- **THEN** the persona and the session state both carry the topic `Spark` +- **WHEN** a planning start carries `topic: "Spark"` +- **THEN** the persona carries `**Topic:** Spark` and the state's `topic` is `Spark` #### Scenario: Default purpose is focus and unchanged -- **WHEN** `POST /api/session/start` is called with no `purpose` -- **THEN** the persona is byte-identical to `build_canonical_persona("focus", topic, energy)`, - the `persona_hash` is unchanged from before the purpose existed, the state's - `mode` is `focus` and its `purpose` is `focus` +- **WHEN** a start names no purpose +- **THEN** the persona equals `build_canonical_persona("focus", topic, energy)`, the state's `mode` is `focus` and its `purpose` is `focus` #### Scenario: Unknown purpose is refused structurally -- **WHEN** `POST /api/session/start` is called with `{"purpose": "revision", ...}` -- **THEN** the response is `422` and no session slot is held +- **WHEN** a start carries `purpose: "revision"` +- **THEN** the response is `422` and no slot is held #### Scenario: Planning start creates no plan and stores no plan id -- **WHEN** one plan exists and `POST /api/session/start` is called with `purpose: planning` -- **THEN** the set of plan ids on disk is unchanged, the `201` body has no - `plan_id`, and the session state has no `plan_id` key +- **WHEN** a planning start succeeds +- **THEN** the plans directory is unchanged, the `201` body has no `plan_id`, and the state carries `purpose == "planning"` and no `plan_id` #### Scenario: Purpose is persisted for the reconnect label -- **WHEN** a `planning` session has started -- **THEN** the session state's `purpose` is `planning` and - `GET /api/session/state` reports `purpose == "planning"` alongside the live - session's id and topic +- **WHEN** a planning start succeeds +- **THEN** the state file's `purpose` is `planning` and `GET /api/session/state` reports it with the fixed topic `Study plan` #### Scenario: A CLI-started architect is labelled from its persisted mode -- **WHEN** the state file was written by `studyloop plan architect` (`mode == - "plan-architect"`, no `purpose` key) and `GET /api/session/state` is called -- **THEN** the body reports `purpose == "planning"`; a file with `mode == - "focus"` and topic `Study plan` reports `focus`; a file carrying `purpose == - "focus"` beside `mode == "plan-architect"` reports `focus` +- **WHEN** the state file carries `mode == "plan-architect"` and no `purpose` +- **THEN** `GET /api/session/state` reports `purpose == "planning"`; an explicit persisted `purpose` wins over the mode; `mode == "focus"` with topic `Study plan` reports `focus` #### Scenario: Brief failure releases the session claim -- **WHEN** `PlanApplication.prepare_planning` raises during a `planning` start -- **THEN** the response is `500` with an `error` naming the brief and - `purpose == "planning"`, the session state file is empty, no in-process - session is held, no study row was created, and a following `focus` start - succeeds with `201` +- **WHEN** the planning brief cannot be built on either transport +- **THEN** the response is a structured `500` with `error`, `purpose` and `repair`, and the active slot is free #### Scenario: PTY and ACP resolve the mode through one resolver -- **WHEN** a `planning` start is made over `transport: pty` and, separately, - over `transport: acp` -- **THEN** each start calls `agent_launcher.persona_mode_for` exactly once - with `planning`, each persona has `**Mode:** plan-architect` and a - `## Planning brief` section, and each state records its own `transport` - with `purpose == "planning"` +- **WHEN** a planning start is made over `pty` and over `acp` +- **THEN** both personas carry the same mode and brief section and neither route names a persona mode as a literal diff --git a/openspec/specs/mcp-server/spec.md b/openspec/specs/mcp-server/spec.md index 9465c32ea..3a16912ee 100644 --- a/openspec/specs/mcp-server/spec.md +++ b/openspec/specs/mcp-server/spec.md @@ -453,3 +453,41 @@ carry a second copy of the inventory or a tool count. - **THEN** they find every registered tool in one table (32 rows), the stated count, and the registration snippet for their harness — and the contract test fails if the table or the count ever disagrees with `register_tools` + +### Requirement: update_study_plan revises the mission through the one gate +`update_study_plan` SHALL expose `why: str | None`, `success: list[str] | None`, +`constraints: list[str] | None` and `out_of_scope: list[str] | None` beside its +existing fields (item 3b, design §3b), forwarded unchanged onto the same one +`RevisePlan` — so the schema's property set is exactly `plan_id, title, topics, +target_date, energy_floor, review_cadence_days, notes, milestones, status, +why, success, constraints, out_of_scope`, and an omitted mission field SHALL +reach the seam as `None` ("leave as is"), never as `""` or `[]`. The seam SHALL +strip `why`, treat each list as a whole-list replacement stripped of blanks, +and refuse a bare string where a list belongs as `InvalidField` (`invalid: …`) +before any write. The resulting document is judged by the single readiness +gate exactly as for every other field: on an active plan a write that leaves +any blocker standing is `not_ready: …` and nothing is saved; a write that +clears every blocker in one call is saved once. With this, every blocker +`readiness()` can name — mission `why`, success criteria, milestones — is +repairable with the tool the architect already holds, so `plan repair` is a +repair rather than dictation. + +#### Scenario: Mission fields reach the seam on one intent +- **WHEN** `update_study_plan("decorators", why="Own the nightly pipeline", + success=["Deploy unaided"], constraints=["Evenings only"], + out_of_scope=["Spark"])` is called +- **THEN** exactly one `RevisePlan` is applied carrying those four values and + `None` for every other field, and the response is the seam's + `PlanDetail.to_json_dict()` + +#### Scenario: Omitted mission fields are None +- **WHEN** `update_study_plan("decorators", notes="Only this.")` is called +- **THEN** the applied intent's `why`, `success`, `constraints` and + `out_of_scope` are all `None` + +#### Scenario: Partial mission repair on a husk is refused; the whole repair lands once +- **WHEN** `update_study_plan("husk", why="…")` is called on an active plan with + no mission, and then `update_study_plan("husk", why="…", success=["…"])` +- **THEN** the first is `not_ready: … No observable success criteria.` with the + document unchanged, and the second saves once, leaves the plan `active` + and `ready`, and `list_study_plans` no longer reports it as a husk diff --git a/openspec/specs/web-ui/spec.md b/openspec/specs/web-ui/spec.md index 97103084b..bd3224569 100644 --- a/openspec/specs/web-ui/spec.md +++ b/openspec/specs/web-ui/spec.md @@ -471,21 +471,33 @@ phase check of its own: an unknown phase on `POST` is the seam's The Study Plans view SHALL offer a **Plan with architect** control beside **New plan** (button name `Plan with architect`, `data-testid="plan-architect"`) with an optional subject field (`data-testid="plan-architect-subject"`, labelled -for assistive technology) and a live status region +for assistive technology), an optional brain-dump textarea +(`data-testid="plan-architect-braindump"`, labelled, `maxlength` equal to the +server's `BRAIN_DUMP_MAX_CHARS`) and a live status region (`data-testid="plan-architect-status"`, `role="status"`, `aria-live="polite"`). One activation SHALL cause exactly one `POST /api/session/start` carrying `purpose: "planning"`, `topic: ` (never omitted; -the server resolves `""` to the fixed label `Study plan`), `origin: "study"` -and the start picker's own `energy`, `agent` and `transport`. The Plans view -SHALL NOT post, open a WebSocket, mount a terminal or listen for the console's -`study-session-start` event: it dispatches one `plan-architect-request` window -event and the Study Session view's session timer — the one owner of the start -POST, the 409 handling and the `study-session-start` event the live console -mounts on — starts the session and navigates the learner to the existing -console (`#study-session`), then reports the outcome back with exactly one +the server resolves `""` to the fixed label `Study plan`), `brain_dump: ` **only when it is non-blank** (a blank dump sends no +key, so the server's "no dump" and "empty dump" are one case), `origin: +"study"` and the start picker's own `energy`, `agent` and `transport`. The +brain dump SHALL never be folded into `topic` and SHALL never ride a `focus` +start. The Plans view SHALL NOT post, open a WebSocket, mount a terminal or +listen for the console's `study-session-start` event: it dispatches one +`plan-architect-request` window event (detail `{purpose, topic, brainDump}`) +and the Study Session view's session timer — the one owner of the start POST, +the 409 handling and the `study-session-start` event the live console mounts on +— starts the session and navigates the learner to the existing console +(`#study-session`), then reports the outcome back with exactly one `plan-architect-result` event. A second activation while a launch is in flight SHALL be a no-op. The Study Session view's `init()` SHALL register its window -listeners once even when called twice (Alpine auto-init plus `x-init`). +listeners once even when called twice (Alpine auto-init plus `x-init`). A +planning launch that arrives before `init()`'s `/api/session/options` fetch has +settled SHALL wait for it before judging whether an agent exists (a cold server +is not a missing agent), and that wait SHALL be bounded (`optionsWaitMs`, 8 s): +past the bound the launch judges the agent as it stands and refuses with the +picker's own `Select an agent to continue.`; a settlement that arrives later +SHALL launch nothing on its own (council review 6). The live console SHALL carry a purpose label (`data-testid="console-purpose-label"`, `role="status"`, `aria-live="polite"`) that is rendered only for a planning @@ -502,59 +514,125 @@ launch refused with the existing `409` conflict shape (`error`, the learner on the picker's recovery block with the reattach lever, not on a second console. The manual **New plan** path SHALL be unchanged. +Abandoning a launch is the console's existing End control (the ■ button, then +the in-page confirmation — never a native dialog): fired as soon as the launch +has been accepted, it SHALL leave no live slot, no plan document, no +`plan_id`, no planning label, and at most one WebSocket ever opened. +Navigating away is **not** the abandon path: a closed socket detaches with a +grace period by design, so an accidental reload cannot kill a live session — +and a launch that lands *after* the learner has left the console is a session +like any other (owner decision 2026-09-18, council review 6 F3b): it exists, +it created no plan, the console reattaches to it on return, and the End +control abandons it. There is no cancel for a pending launch. + +The architect's one-question-at-a-time protocol is asserted as persona text +(owner decision D-B): the browser and unit tests prove the brief — including +the brain dump — is delivered to the agent process; no test asserts how a live +model behaves with it, and the docs say so. + #### Scenario: One click, one POST, the existing console -- **WHEN** the learner types `SQL window functions` into the subject field and - activates **Plan with architect** -- **THEN** exactly one `POST /api/session/start` is made with `purpose == - "planning"`, `topic == "SQL window functions"`, `origin == "study"`; the - response is `201` with `purpose == "planning"` and a `ws_url`; the page - navigates to `#study-session`; exactly one `study-session-start` event with - `purpose == "planning"` is dispatched; exactly one WebSocket to the `ws_url` - opens; exactly one console is visible and nothing is mounted in the Plans view +- **WHEN** the learner types `SQL window functions` into the subject field and activates **Plan with architect** +- **THEN** exactly one `POST /api/session/start` is made, with `purpose == "planning"`, `topic == "SQL window functions"`, `origin == "study"` and no `brain_dump` key +- **AND** the response is `201` with `purpose == "planning"` and a `ws_url` +- **AND** the page navigates to `#study-session`, exactly one `study-session-start` event fires with `purpose == "planning"`, exactly one WebSocket is opened, one console is visible and nothing is mounted in the Plans view + +#### Scenario: The brain dump reaches the architect +- **WHEN** the learner types a multi-line brain dump into the textarea, leaves the subject blank and activates the control +- **THEN** the one POST carries `brain_dump` equal to the trimmed text and `topic == ""`, the `201` names the topic `Study plan` +- **AND** the persona the fake agent received contains `### Learner's brain dump` with every dump line quoted (`> …`), `**Topic:** Study plan`, and `GET /api/session/state` carries neither the key nor the text #### Scenario: No subject -- **WHEN** the subject field is empty and **Plan with architect** is activated -- **THEN** the request carries `topic == ""` and the `201` body's `topic` is - `Study plan` +- **WHEN** the learner activates the control with an empty subject +- **THEN** the request carries `topic == ""` and the `201` body's `topic` is `Study plan` #### Scenario: Label survives a reload -- **WHEN** a planning session is live and the page is reloaded -- **THEN** the console re-adopts the session from `GET /api/session/state`, - whose body has `purpose == "planning"`, and exactly one visible - `console-purpose-label` reads as a planning session, before and after the - reload +- **WHEN** a planning session is running and the learner reloads the page +- **THEN** the console shows exactly one visible purpose label reading "planning" both before and after the reload #### Scenario: CLI-started architect is labelled from its persisted mode -- **WHEN** the session state file was written by `studyloop plan architect` - (`mode == "plan-architect"`, no `purpose` key) and `GET /api/session/state` - is called -- **THEN** the body has `purpose == "planning"`; a file with `mode == "focus"` - and `topic == "Study plan"` reports `focus`; a file carrying `purpose == - "focus"` beside `mode == "plan-architect"` reports `focus` +- **WHEN** the session state file carries `mode == "plan-architect"` and no `purpose` key +- **THEN** `GET /api/session/state` reports `purpose == "planning"`; a file with `mode == "focus"` and topic `Study plan` reports `focus` #### Scenario: The brief's structure, never its wording -- **WHEN** the fake PTY agent receives the persona of a Web-launched planning - session -- **THEN** it contains, in order, `## Planning brief`, `### Interview`, - `### Evidence from the learner's history`, `### Existing plans` and the - architect body's `## Tooling` section, with `**Mode:** plan-architect` and no - `Resuming Previous Session` section +- **WHEN** a planning launch reaches the fake PTY agent +- **THEN** the persona it received contains, in order, `## Planning brief`, `### Interview`, `### Evidence from the learner's history`, `### Existing plans`, `## Tooling`, with `**Mode:** plan-architect` and no `Resuming Previous Session` #### Scenario: Nothing is created by the launch -- **WHEN** **Plan with architect** is activated -- **THEN** `GET /api/plans` is unchanged, the plans directory holds no new - document, and the session state carries no `plan_id` +- **WHEN** the control is activated and the console mounts +- **THEN** `GET /api/plans` is unchanged, the plans directory holds no new document, and the session state carries no `plan_id` + +#### Scenario: Abandoning a launch mid-flight leaves no session and no plan +- **WHEN** the learner activates the control and, as soon as the `201` arrives, uses the console's End control and confirms +- **THEN** `GET /api/session/state` has no `study_session_id` and no planning purpose, `GET /api/plans` and the plans directory are unchanged, exactly one `study-session-start` fired, at most one WebSocket was opened, and no purpose label is visible + +#### Scenario: Leaving before the launch lands leaves a session like any other +- **WHEN** the learner clicks "Plan with architect", the start POST is held, + the learner navigates to the Plans view before it is answered, and the POST + is then released with `201` +- **THEN** `GET /api/session/state` reports that `study_session_id` with + `purpose: "planning"` and not `ended` (and still does half a second later), + `GET /api/plans` and the plans directory are unchanged, returning to the + Study Session view reattaches the console to the same session with one + planning label and one `study-session-start` event in total, and the End + control releases the slot #### Scenario: Conflict is the existing shape with a reattach lever -- **WHEN** a session is live and `POST /api/session/start` is called again with - `purpose: "planning"` -- **THEN** the response is `409` with `error`, `study_session_id`, `topic`, - `agent`, `detached` and `reattach_url == `; and a - Plans-view launch that meets that `409` shows the picker's `.picker-error` - with the body's `error` and the `Reattach to this session` control, with no - terminal mounted +- **WHEN** a session is already running and the learner activates the control +- **THEN** the POST returns the existing `409` body and the picker's recovery block appears with the reattach lever; no second console mounts #### Scenario: Manual New plan is unchanged -- **WHEN** the learner uses **New plan**, fills the form and creates the plan -- **THEN** the reader shows the plan, `GET /api/plans` counts one more, and no - session was started +- **WHEN** the learner uses **New plan**, fills the form and submits +- **THEN** the plan is created and listed and no session is started + +### Requirement: Plan list rows carry readiness and the sidebar marks a husk +Every row of `GET /api/plans` SHALL be `PlanSummary.to_json_dict()` and SHALL +carry `ready` — the same verdict the readiness gate judges every write by — +as its eighteenth key, so a client can tell an active-but-unready plan (a +"husk", item 3 / D-C) from the list alone, with no per-row round-trip. The +Plans sidebar SHALL render one mark (`data-testid="sidebar-plan-husk"`, the +glyph `!`, `role="img"` with an `aria-label` and a `title` naming the +condition and both exits) inside `.sidebar-plan-meta` for a row whose +`status` is `active` and whose `ready` is `false`, and SHALL render it for no +other row. The mark is computed from the row's own `ready`; the sidebar +SHALL make no further request to decide it. Colour SHALL come from the theme's +own tokens and SHALL not be the only carrier of the information. + +#### Scenario: List payload carries ready +- **WHEN** `GET /api/plans` is served for one ready active plan and one active + document with no mission +- **THEN** every row has 18 keys, the ready plan's row has `ready: true`, the + husk's row has `ready: false`, and `?status=active` returns both + +#### Scenario: Sidebar marks the husk and only the husk +- **WHEN** the Plans sidebar renders those rows +- **THEN** exactly one `sidebar-plan-husk` mark is present, on the husk's row, + and the ready plan's row has none + +### Requirement: PATCH carries the mission to the one RevisePlan +`PATCH /api/plans/{id}` SHALL accept `why`, `success`, `constraints` and +`out_of_scope` beside the fields it already carries (item 3b, design §3b), and +SHALL translate them onto the same single `RevisePlan` — the route validates +and writes nothing itself. An absent key is `None` ("leave as is"); a +supplied list replaces the whole list. The seam's refusals map as they always +have: a bare string where a list belongs is the seam's `InvalidField` → `400` +naming the field; a write whose resulting document would be active but not +ready is `422` with the `readiness` body and nothing written. The `PATCH` +response is the write receipt (`updated`, `plan`, `readiness`); the mission is +read back from `GET /api/plans/{id}`. With this the Web UI's whole-document +`PATCH markdown` is no longer the only mission writer. + +#### Scenario: Partial mission on a husk is 422; the whole mission lands and the row flips to ready +- **WHEN** `PATCH /api/plans/husk` is sent `{"why": "…"}` on an active plan + with no mission, and then `{"why": "…", "success": ["…"], "constraints": + ["…"], "out_of_scope": ["…"]}` +- **THEN** the first is `422` with `detail.ready == false` and + `detail.blockers == ["No observable success criteria."]` and the document + is unchanged; the second is `200` with `plan.status == "active"`, + `plan.ready == true` and `readiness.blockers == []`, `GET /api/plans/husk` + returns the four mission values, and the `GET /api/plans` row for `husk` + has `ready: true` + +#### Scenario: A string where a list belongs is the seam's 400 +- **WHEN** `PATCH /api/plans/{id}` is sent `{"success": "one string"}` +- **THEN** the response is `400` naming `success` and the document is + unchanged diff --git a/packages/studyloop/pyproject.toml b/packages/studyloop/pyproject.toml index d43edbafc..c70e83580 100644 --- a/packages/studyloop/pyproject.toml +++ b/packages/studyloop/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "studyloop" -version = "0.4.0" +version = "0.5.0" description = "AuDHD-aware study tool with AI Socratic mentoring, spaced repetition, and content pipeline" requires-python = ">=3.12" readme = "README.md" diff --git a/packages/studyloop/tests/test_release_consistency_script.py b/packages/studyloop/tests/test_release_consistency_script.py index 213b885ba..14a41b999 100644 --- a/packages/studyloop/tests/test_release_consistency_script.py +++ b/packages/studyloop/tests/test_release_consistency_script.py @@ -307,3 +307,79 @@ def test_release_note_that_is_still_the_prepare_release_skeleton_fails(tmp_path: assert result.returncode == 1 assert "skeleton" in result.stderr.lower() or "release summary" in result.stderr.lower() + + +# --------------------------------------------------------------------------- +# The openspec-shipped guard: a change with commits since the last tag must be +# archived or carry an EXPLAINED `deferred:`. `.openspec.yaml` writes the reason +# as a YAML folded scalar (`deferred: >-` + indented lines), and the guard's +# line-match read the fold marker `>-` as the reason -- so an empty fold passed +# and the printed reason was `>-`, defeating the "unexplained deferral is +# indistinguishable from a forgotten one" rule the guard exists for. +# --------------------------------------------------------------------------- + + +def _write_changes_fixture(tmp_path: Path, openspec_yaml: str | None) -> None: + """A tagged repo, then a post-tag commit touching one openspec change.""" + _write_release_fixture(tmp_path, "1.2.3", "2026-09-06") + init_git_repo_with_tag(tmp_path, tag="v1.2.3", tag_date="2026-09-06") + change_dir = tmp_path / "openspec" / "changes" / "some-change" + change_dir.mkdir(parents=True) + (change_dir / "proposal.md").write_text("# proposal\n", encoding="utf-8") + if openspec_yaml is not None: + (change_dir / ".openspec.yaml").write_text(openspec_yaml, encoding="utf-8") + subprocess.run(["git", "add", "-A"], cwd=tmp_path, check=True, capture_output=True) + subprocess.run( + ["git", "commit", "-q", "-m", "touch change"], + cwd=tmp_path, + check=True, + capture_output=True, + text=True, + ) + + +def test_release_mode_names_an_undeferred_change_with_commits_since_the_tag( + tmp_path: Path, +) -> None: + _write_changes_fixture(tmp_path, openspec_yaml="schema: spec-driven\n") + + result = run_release_check(tmp_path) + + assert result.returncode == 1 + assert "neither archived nor deferred" in result.stderr + assert "some-change" in result.stderr + + +def test_release_mode_accepts_a_folded_deferral_and_prints_its_text_not_the_fold_marker( + tmp_path: Path, +) -> None: + _write_changes_fixture( + tmp_path, + openspec_yaml=( + "schema: spec-driven\n\n" + "deferred: >-\n" + " Follow-on programme; archived when the owner\n" + " scores the last rubric row.\n" + ), + ) + + result = run_release_check(tmp_path) + + assert result.returncode == 0, result.stderr + assert "deferred: Follow-on programme; archived when the owner scores the last rubric row." in ( + result.stdout + ) + assert "deferred: >-" not in result.stdout + + +def test_release_mode_refuses_an_empty_folded_deferral(tmp_path: Path) -> None: + """`deferred: >-` with nothing under it is the forgotten-deferral case.""" + _write_changes_fixture( + tmp_path, openspec_yaml="schema: spec-driven\n\ndeferred: >-\n\ncreated: 2026-09-16\n" + ) + + result = run_release_check(tmp_path) + + assert result.returncode == 1 + assert "neither archived nor deferred" in result.stderr + assert "some-change" in result.stderr diff --git a/pyproject.toml b/pyproject.toml index 3ad4720ec..672dfd4fd 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -3,7 +3,7 @@ name = "studyloop-workspace" # Must equal packages/studyloop/pyproject.toml's version -- enforced by # scripts/check-release-consistency.py (R-39; this drifted to 1.0.0 while the # package sat at 0.1.0, unnoticed because the script never read this file). -version = "0.4.0" +version = "0.5.0" description = "StudyLoop: AuDHD-aware Socratic mentoring with AI session management" requires-python = ">=3.12" readme = "README.md" diff --git a/releases/v0.5.0.md b/releases/v0.5.0.md new file mode 100644 index 000000000..c1a1ab9ad --- /dev/null +++ b/releases/v0.5.0.md @@ -0,0 +1,129 @@ +# v0.5.0 + +The plan-integration release: a Study Plan is no longer a document that sits +beside the product. Activating one changes what `studyloop now`, the Web Today +card, the recap and MCP `get_next_action` recommend; the Study Plan Architect +can be launched from the Web UI through the ordinary session machinery; MCP +clients get lifecycle parity over the same application seam the CLI and Web +use; and a low-energy day is handled honestly — repair work the day cannot +carry is deferred with a reason, and when nothing plan-related fits, the +recommendation is to sit with the plan rather than take the least-bad task. +Every change in this release was built test-first, reviewed by a three-seat +model council against the diff, and — where it changes what a learner is told +to do — scored by the owner on a human rubric before it shipped. + +StudyLoop 0.5.0 follows the documented source-checkout installation flow. CI +builds and installs wheel/sdist artifacts as validation evidence; this release +does not advertise GitHub or PyPI binary distribution. + +## Changes + +### One application seam for Study Plans + +- `studyloop.planning.application` (`PlanApplication`) carries every plan use + case — browse, inspect, prepare planning, active guidance, apply a change, + assess — behind immutable views and transport-neutral domain errors. The + CLI, the Web routes and the MCP tools are thin adapters over it; an AST + guard keeps them that way. +- Two bugs closed by construction: activation readiness is judged once, on + the resulting document, for every entry path (the create-with-status and + whole-document-replacement doors used to bypass it); a checkpoint whose + database write failed is reported as partial, never as a clean record. +- Markdown stays the source of truth and SQLite stays derived; plan identity + and creation time survive revision and document replacement. + +### Plan-aware `now` + +- An active plan biases the recommendation and never filters it: matching due + work and a synthesised next milestone carry explicit `plan_refs`; urgent + reviews and fresh struggles can still outrank new milestone work; several + active plans are considered deterministically; a milestone beyond today's + energy is deferred with a reason. With no active plan the JSON is + byte-identical to the pinned golden. +- Repair has an energy demand of its own: a live struggle asks for 6/10, an + older one or a weak teach-back for 4/10, a concept still `learning` for + none. Below the day's capability a repair is deferred like new work + (`energy_deferred_repairs`); due recall and teach-back reviews are never + deferred. When nothing plan-related fits and an active, ready plan exists, + one low-scoring proposal is synthesised — sit with the plan in a body-double + session — that leads with the progress the plan records and names what is + deferred. +- A concept still `learning` is offered as the one-sentence *micro* + teach-back and described as the gentle review it is; the teach-back + protocol gains a low-energy fallback (a short guided explanation and one + phrase back when the sentence will not come, instead of the four-round + stuck ladder; the blank is not scored). +- A fully-checked active plan yields a completion action, not more study + work: `studyloop plan close ` reviews the plan against its evidence and + proposes *extend* or *close*; status never changes by itself, and a partial + review never proposes a clean close. + +### The architect on every surface + +- **Web UI:** *Plans → Plan with architect* launches the Study Plan Architect + through the existing session start, one-session authority, reconnect flow + and console, with a `planning` purpose; starting a conversation creates no + draft. The manual form is retained. Leaving the page before a launch lands + leaves a session like any other. +- **CLI:** `studyloop plan repair ` opens the architect on an active plan + the readiness gate refuses to write; `plan close ` as above. Both + contain learner-authored text through the seam. +- **MCP:** `list_study_plans`, `get_study_plan`, `get_planning_interview`, + `create_study_plan`, `update_study_plan` (mission included), + `set_study_plan_status`, `set_study_plan_milestone`, `evaluate_study_plan` + and `delete_study_plan` (confirmation required) join the existing + `record_plan_learning`, all thin adapters over the seam. + `get_next_action` gains `interleave` parity with the CLI and Web. +- Nothing in the planning flow is specific to one harness: one canonical + persona is delivered to Kiro CLI, Claude Code, Codex, pi, OpenCode and Grok + Build through each harness's own door. + +### Safety and correctness + +- Every command `studyloop now` offers quotes learner-authored values as one + shell word; a plan title carrying `$(…)` no longer executes when pasted. +- A legacy `study_progress` row with no `last_seen` no longer crashes the + struggle collector. +- The due-review reader emits every `struggling` row as a hands-on "guided + repair" too — the same repair the struggle collector defers, collected a + second time. That copy now carries the repair's demand and defers with it + (named once), so a low-energy day no longer shows the repair the learner + was just spared as the primary with its own deferral printed beneath it. + Found by emitting the owner's rubric readings through both readers; the + earlier fixture had silenced one. +- The Study Session timer's `init()` ran twice per page load, each run + reading the live-session state; the slower read could adopt a session + started in another tab into a tab that was sitting on the picker. It now + runs once, so a second tab's Start is refused with the reattach offer it + was designed to get. +- `anyio` 4.15.1 (two published advisories); both dependency audits clean. + +### Release tooling + +- The release gate reads a folded `deferred: >-` reason in an openspec + change's metadata as its text and refuses an empty block. +- Seam tests seed their per-test `sessions.db` from one migrated template per + process (241 production bootstraps → 1 on the same test set), removing the + step a slow CI disk had stalled in. +- The nightly install check plants a harness marker before running the + installer in its isolated HOME; the job had been red every night since it + was added. + +## Verification + +- Every PR in this release was merged as a local fast-forward of `main` after + CI green on the PR's own head (#20, #22, #23, #24, #27, #28) — 15/15 jobs + including the 3.12/3.13 matrix, e2e and install-smoke; the nightly install + check dispatched against the installer fix: 7/7 green (run 35498047718). +- Full unit suite on each item's final tree against a clean-`main` control + worktree: zero regressions each time (the 44 sandbox-environmental ids are + identical on both sides and recorded by name in + `docs/architecture/plan-integration/receipts/`). +- Council records: `docs/architecture/plan-integration/council/` (reviews 1–7 + and the rubric-3b decision brief); owner rubric: + `docs/architecture/plan-integration/receipts/now-rubric-2026-09-16.md` — + every row scored by the owner, row 3b (six readings) on 2026-09-20 with each + reading re-emitted through both real readers before its verdict was read as + a verdict on shipped behaviour. +- `just release-check` (pre-tag) on the release commit; `just release-verify` + after the tag. diff --git a/scripts/check-release-consistency.py b/scripts/check-release-consistency.py index fe8de453b..ad53bd2d8 100755 --- a/scripts/check-release-consistency.py +++ b/scripts/check-release-consistency.py @@ -132,17 +132,37 @@ def _deferred_reason(change_dir: Path) -> str | None: stdlib-only property. A bare ``deferred:`` with no reason does NOT count — an unexplained deferral is indistinguishable from a forgotten one, which is the state this guard exists to catch. + + The reason is usually written as a YAML block scalar (``deferred: >-`` with + the text on the indented lines below). The block indicator is not the + reason: the indented continuation is, joined into one line, and an + indicator with nothing under it is an unexplained deferral like any other. """ meta = change_dir / ".openspec.yaml" if not meta.is_file(): return None - for line in meta.read_text(encoding="utf-8").splitlines(): - match = re.match(r"\s*deferred:\s*(\S.*)$", line) - if match: - return match.group(1).strip() + lines = meta.read_text(encoding="utf-8").splitlines() + for index, line in enumerate(lines): + match = re.match(r"\s*deferred:\s*(.*)$", line) + if not match: + continue + value = match.group(1).strip() + if value and value not in _YAML_BLOCK_INDICATORS: + return value.strip("\"'") or None + continuation: list[str] = [] + for following in lines[index + 1 :]: + if not following.strip(): + continue + if not following[0].isspace(): + break + continuation.append(following.strip()) + return " ".join(continuation) or None return None +_YAML_BLOCK_INDICATORS = frozenset({">", ">-", ">+", "|", "|-", "|+"}) + + def validate_openspec_changes_shipped(repo_root: Path) -> None: """Release mode only: a change that shipped work must be archived. diff --git a/uv.lock b/uv.lock index 698eef577..eac54115e 100644 --- a/uv.lock +++ b/uv.lock @@ -3003,7 +3003,7 @@ wheels = [ [[package]] name = "studyloop" -version = "0.4.0" +version = "0.5.0" source = { editable = "packages/studyloop" } dependencies = [ { name = "agent-session-tools" }, @@ -3097,7 +3097,7 @@ dev = [ [[package]] name = "studyloop-workspace" -version = "0.4.0" +version = "0.5.0" source = { editable = "." } dependencies = [ { name = "fastapi" },