Item 4 plan close (D-G), council review 6 corrections, verify checks - #22
Merged
Merged
Conversation
Seven tests pin scenario 4's redesign before any production code changes: the completion action rule 9 emits for a fully-checked active plan must carry the end assessment (due reviews, struggles and unverified milestones counted on the plan's own concepts), propose `extend` while any count is above zero and `close` when all are zero, compose its sentence from that proposal, and never change a status — the read is `assess(AssessPlan(phase="end", record=False))`, the preview path, so the document, its status and the checkpoint log are untouched. A failed assessment keeps the pre-change sentence and adds a warning; `now` never fails on it. `plan close <id>` is the launch sibling of `plan repair`: the same architect chain with a `### Closing review` first section (three counts, proposal, evidence lines), refusing a plan with open milestones with their count. Owner decision 2026-09-17: the completion review counts only due rows that name a concept. The scheduler's "New topic -- start fresh" row (`concept: None`, `evidence: configured_topic`) is a cold-start hint for "what should I review now", not a lapsed review; the evaluator already ignores it at concept level (it never matches a milestone concept, so it contributes nothing to `unverified_milestones`), and counting it would tell a learner who has just ticked every milestone to "start fresh". `plan evaluate` keeps the row — the exclusion is the completion review's, applied by one definition in both the engine and the brief. Evidence is planted on the `studyloop.history` package attributes the evaluation resolves at call time, so the relevance filter and `has_evidence` logic stay live and no sessions database is involved. All seven fail for the intended reason (missing attributes, `assess` never called, no warning, no `close` command); the 57 pre-existing tests in both files pass and the no-plan golden sha is unchanged. The first test's T4.1 name lost one redundant word (`plan_concepts` → `concepts`): no def line in the repo exceeds 100 chars and none carries a `noqa`. The `pyright: ignore[reportAttributeAccessIssue]` tags on the not-yet-existing attributes are the 3b RED's pattern; GREEN strips them.
…tem-4 RED pins Design §4 now states the closing-review brief format, intro and refusal text the RED fixed; records the owner's 2026-09-17 decision that the completion review counts only due rows naming a concept (the scheduler's "new topic" row is excluded, with the reasoning and the measured cost of the preview); and keeps the one decision deliberately left open — what `proposal` holds when the end assessment fails — with the default GREEN will take unless vetoed (nullable proposal, counts stay int).
…(D-G)
A fully-checked active plan used to get an either-way sentence ("close the
plan or extend it") that the owner scored "no as phrased" on rubric row 4:
it proposed nothing and rested on nothing. The completion action is now a
closing review read from the plan's own end assessment, and `plan close <id>`
is the door to acting on it — with the learner, never automatically.
One definition, two surfaces. `planning.views.CompletionReview` reads the
`end` evaluation into three counts on the plan's concepts (due reviews,
struggles, milestones marked done with no evidence), a proposal (`extend`
while any count is above zero, else `close`) and one evidence line per
counted item (capped at 8). Both the `now` engine's `CompletionAction` and
the `plan close` brief consume it, so they cannot disagree on a count.
Due reviews count only rows that name a concept (owner decision,
2026-09-17): the scheduler's `New topic -- start fresh` row is a cold-start
hint for "what should I review now", not a lapsed review, and counting it
would tell a learner who just ticked every milestone to start fresh.
`plan evaluate` keeps the row; the exclusion is the completion review's.
The engine reads through the preview path — `assess(AssessPlan(phase="end",
record=False))`, one call per fully-checked plan per `build_now_plan` — so
the document, its status and the checkpoint log are untouched. A failed
assessment keeps the plain sentence with `proposal` None and the counts 0,
plus a warning naming the plan: the counts are then unknown, not zero, and no
renderer reads a clean slate or outstanding work into a failure. Measured on
the live 877 MB sessions.db the preview costs ~320 ms per fully-checked plan,
a transient state by design. New JSON keys appear only inside
`completion_actions` entries; the no-plan golden is byte-identical.
`plan close <id>` mirrors `plan repair`: `_inspect`, refuse with the open
count (exit 1) while any milestone is open, leave a `complete` plan alone,
and for a fully-checked plan launch the architect through the one
`study --mode plan-architect` chain with `### Closing review` as the brief's
first section (the four count/proposal lines, then the evidence). The shared
"plan as it stands" section is refactored out of the repair brief
byte-identically. The persona gains "Extend or close" (three projections
re-projected, manifest regenerated with `updated` restored on unmoved
entries, secrets baseline refreshed whole-repo: 72 -> 72 files, exactly the
two manifest hashes). CLI `now` prints the evidence lines beneath the
sentence; the Today card gains `completionEvidence()` for the same lines
(JS 136/136); the recap speaks the composed sentence. Docs, two spec deltas,
and rubric row 4b (re-run primary and both proposals, verdict PENDING for
the owner; row 4 kept as the record of the finding).
Verification: 7 RED -> green (64/64 across the two files); golden sha
ec451ce8 unchanged; `just lint`, `just typecheck`, `mkdocs build --strict`,
`openspec validate` clean. Full suite 30 failed / 7233 passed / 14 errors,
diffed against a clean f1c52ce control run in parallel: item4 - control =
empty, control - item4 = exactly the seven REDs, 44 shared environmental ids
committed by name in receipts/full-suite-control-item4-2026-09-18.md.
…em 4's GREEN-time decisions T4.2 GREEN is 8229329 (7 RED -> green, matched-control full suite with zero regressions, 44 environmental ids now committed by name); T4.3's rubric row 4b is in the same commit with the verdict pending the owner. Design section 4 closes its one open decision - `proposal` is nullable, None on a failed assessment, represented once at the action - and records the two decisions taken in GREEN: the Today card renders the review's evidence lines, and the docs' pinned six-boundary list stays at six with the consensual close stated beside it.
…oth readings The owner scored rubric row 4b on 2026-09-18: (a) the evidence-backed 'extend' proposal is one they would walk, (b) a clean review's 'close' is one they would agree to. This closes row 4's 'no as phrased' finding from 2026-09-16; row 4 stays as the record of the original finding. Recorded in the receipt's status line and verdict cell exactly as given (no rationale was supplied beyond the two answers, and none is invented), in tasks.md (T4.3 note; T5.5's archive condition now names 3b as the one row still outstanding), and on the public study-plans page, which said the re-run row 'awaits the maintainer's score'. The contract pin keys on 'owner verdicts RECORDED' and stays satisfied.
…e repair/close refusals; a red suite names its nodes (T6.3) Design §6 of plan-integration-followons asks the verification receipt for two more registered checks beside the golden: the two architect grants (the nine plan tools + record_plan_learning, derived from the inventory, in the spelling each harness honours) and the plan repair / plan close refusal texts as the CLI seam tests pin them. Five tests pin their names, shapes and one negative case (a copy of the tree with one grant removed and one stray grant added must fail naming both). The fifth test is receipt honesty: the first real verify run on 9d10fee recorded full-suite-studyloop as ok:false with a 12-line output_tail, so the receipt could not say WHICH of the sandbox's 30 failures + 14 errors occurred and could not be reconciled against the 44 named environmental ids. A failed pytest check now keeps every FAILED/ERROR short-summary line on its row. RED: 5 failed / 21 passed, each for the intended missing name, function or key.
… checks; a red pytest check names its failed nodes (T6.3) — GREEN Design §6 of plan-integration-followons: two registered checks beside the golden. `architect-grants` is in-process — it derives the ten names from studyloop.mcp.inventory (PLAN_TOOL_NAMES + LEARNING_RECORD_TOOL) and judges agents/kiro/study-plan-architect.json (`@studyloop` visible in `tools`; allowedTools carries exactly the ten as `@studyloop/<tool>`, no bare server grant, no inert mcp_studyloop_* spelling) and the Claude architect's `tools:` frontmatter (exactly the ten as `mcp__studyloop__<tool>`, no other server), measuring what it found so the receipt reads without the files. `repair-close-refusals` runs exactly the four CLI seam node ids that pin the husk refusal (both exits), `plan repair` on a ready plan and an unknown id, and `plan close` on an unfinished plan. Receipt honesty: a failed pytest check now keeps every FAILED/ERROR short-summary node id on its row (`failed_nodes`), so a red full suite can be reconciled against the named environmental set from the receipt alone — the first real run on 9d10fee could not be. Registry 29 -> 31. Verify-script tests 5 RED -> 26 passed; both new checks pass on this tree; ruff, format, pyright clean.
…ncil review 6, F1)
evaluate_plan turns a history reader that fails into one warning
('<reader> unavailable — evaluation is partial') and an empty default, so a
count read while that reader was down is unread, not zero. CompletionReview
read those zeros as a clean slate: the now engine said 'the closing review is
clean — it proposes closing the plan' and plan close printed 'proposes:
close' with the gap relegated to a separate section. GPT-Astra 🔴 F1, Grok 🔵
(v); reproduced with two RED tests before this change.
The fix lives in the one definition. CompletionReview keys on the evaluator's
own marker (PARTIAL_READ_MARKER, defined once in evaluation.py and used by
_safe), keeps the counts it did read, sets partial=True, proposes None —
neither close nor extend — and names each gap among its evidence lines
('Not read: …'), so both surfaces carry it without a second path.
CompletionAction gains partial; the sentence has a third branch that says the
review is partial and could not propose; plan close's proposal line reads
'unassessed — the review is partial' and its status line no longer says the
review proposes. The separate '### Data gaps' brief section is gone: the gaps
are the review's own lines now.
Persona 'Closing a Plan' tells the architect what an unassessed proposal
means (walk what was read, prefer re-running the review, never infer a clean
slate); three projections re-projected, manifest regenerated (updated
restored on the 20 unmoved entries), baseline refreshed whole-repo with the
pinned detect-secrets 1.5.0 (exactly the two manifest hashes, 72 -> 72
files). Spec delta states the rule and adds the scenario.
66/66 across the two item-4 files (2 RED -> green), 614 across the plan
suites + pins, JS 136/136, golden sha unchanged, ruff/pyright/mkdocs/openspec
clean.
…council review 6) 626ea12 made a planning click await init()'s /api/session/options fetch before judging whether an agent exists, so a cold server no longer refused with 'Select an agent'. That put an unbounded network wait on the click's critical path: a request that never settles (a server that accepts the connection and stalls) held the learner's click forever — no POST, no refusal, no message. qwen3-coder 🔴, GPT-Astra F3, Grok 🔵 (i); reproduced by a JS test that never settles the fetch (timed out at 5 s). The wait now races the options promise against optionsWaitMs (8 s, on the timer's state so tests can shorten it); past the bound the launch judges the agent as it stands and gives the picker's own refusal, and a settlement that arrives later launches nothing on its own. The two existing wait tests (resolve-then-launch once; empty picker still refuses) are unchanged. JS 137/137 (+1), browser journey 11/11 -m e2e; web-ui spec delta states the wait and its bound.
…ument is named (council review 6, F4) Three sentences claimed more than the evidence. husk_provenance read a creation stamp before the gate as 'was never judged by it' — but a plan created before 2026-09-15 can be saved ready after it and hand-edited into a husk later; the stamp establishes when the document was created, not what judged it or when it became incomplete. It now says 'This plan's creation stamp predates the readiness gate (…); the seam cannot tell when it became incomplete.' doctor's healthy row promised 'every write the gate judges will pass' — a future write can remove a required field — and now says 'all ready as they stand.' And husks() skipped a document it could not load, so doctor could report all-ready over a directory holding a file no listing can read. GPT-Astra 🔴 F4; reproduced with a parametrised provenance test and two doctor tests before this change (the first fixture was not unreadable at all — the parser is lenient with malformed front matter — the real class is a non-UTF-8 file or a filename that is not a valid id). PlanApplication.survey_husks() is the one pass that returns both facts (HuskSurvey: husks + unreadable ids); husks() is now a view over it, so plan list/repair are unchanged. doctor emits one warn row per unreadable document beside the readiness rows, fix_auto=False. Health and cli-surface deltas state the honest wording and the new row. 312 across the pinning files, guard 30/30, ruff/format/pyright clean, openspec valid.
… mentor grants went live, and when to re-probe kiro-cli (council review 6, F5)
The Claude architect's frontmatter names ten mcp__studyloop__ tools, and no
diff in the reviewed range showed where the server itself is declared for
Claude Code — so three seats asked whether the grant was inert the way the
mentor's had been (GPT F5 🟡, Grok 🔵 c). It is not: installers._MCP_HARNESSES
includes claude and `studyloop install agents` merges both servers into
~/.claude.json's mcpServers. The install doc now says so, and the pin takes
the path from installers._mcp_config_path('claude') rather than a remembered
string, so the sentence cannot drift from the code.
Two more disclosures the seats asked for: existing Kiro study-mentor installs
start seeing the twelve MCP tools their file always named once re-installed
(f5c2057 fixed the inert spelling; nothing a user read said so — Grok 💡 d,
GPT 🔵), and the grant spelling is evidence pinned to kiro-cli 2.21.4/2.22.0,
so the doc and the probe receipt's header both say to re-run probes A and B
on a newer binary. The receipt also states that the session-db
'visible, prompts' reading is an expectation, not a measurement (Grok 🔵 a).
Persona/install/docs/prompt-contract pins 137 passed; mkdocs --strict clean.
…s plan; the label says review, not complete (council review 6, F6) The card rendered every completion sentence, then every evidence line flattened beneath them, so with two finished plans a line lost the plan it belonged to — the contextual review D-G asks for, weakened at the surface most learners read. And both the card and CLI now labelled the note 'Plan complete' over a plan whose status is still active until the learner agrees with the architect. GPT-Astra F6 🟡; reproduced from the markup. today-panel.js gains completionReviews(): one block per action — sentence and its own evidence lines — keyed by plan_id in the engine's order; the flat completionNotes()/completionEvidence() helpers now derive from it. index.html renders one .today-plan-review block per plan (data-plan-id) with the lines nested inside, labelled 'Closing review'; cli/_now.py prints the same label. A JS test pins the grouping (an extend review, a partial review carrying its 'Not read:' line, and a pre-D-G entry) and a markup test pins the keyed block, the nesting and the absence of both a flat evidence loop and the word 'Plan complete'. No test had pinned either old label. JS 139/139 (+2), CLI now/guidance/seam 67/67, e2e Today/plan subset 29 passed; design §4 records the correction.
…ues the way the Web brief does (council review 6, F2) The Web door one-lines every value the planning brief quotes (review-3 F4); the CLI plan repair / plan close briefs interpolated the plan's title, id, topics, created stamp and the review's evidence lines raw. A YAML-quoted front-matter title holding '\n## Forged section' survives parse_plan, and _render_plan_as_it_stands printed it as a real heading inside the brief the architect reads. GPT-Astra F2 🟡; reproduced by rendering such a document (the '##' line was a heading of the brief). The containment is now one definition on the seam, studyloop.planning.one_line (whitespace runs, newlines included, collapse to one space, so no value can start a line): the CLI briefs quote every learner-authored value through it, and the Web door's _one_line delegates to it, so every adapter's brief is contained the same way. The repair brief's blocker lines are seam-authored sentences and are left as they are. The architecture guard passes: one_line is a studyloop.planning import (D-6). RED test pins both briefs against a hostile title, topic and evidence line; guard + CLI seam + Web brief tests 119 passed; ruff/pyright clean.
…only branches, the evidence cap, and plan close's three no-launch exits (council review 6, F7) The seven item-4 REDs established due-work extension, clean closure, new-topic exclusion, the exception fallback, no plan writes and the two principal CLI paths. GPT-Astra F7 🟡 named what they did not: a struggle alone or an unverified milestone alone carrying 'extend', the evidence cap's overflow line, and plan close on an already-complete plan, on a plan with no milestones, and on an unknown id. Six tests now pin those branches; the zero-milestone test also proves no assessment is read on that exit. Discrimination proved by mutation: with the proposal rule changed to read the due count alone, exactly the struggle-only and unverified-only tests fail and the rest stay green; the source was restored byte-identical. 73 passed across the two files; ruff/pyright clean.
Brief over items 1-4 of plan-integration-followons, reviewed tree 9d10fee (range 1565234..9d10fee: 29 commits, 73 files): the owner's decisions and the four taken during the batch, design §1-§4 verbatim, the agent's own T-notes verbatim, every diff grouped by item, the seven outside-the-items commits classified, reference facts each verified on the tree, and ten numbered deliverables. 6,082 lines, ~95k tokens; the manifest records the sha256 of the bytes the seats received (7ef91328…) — the pre-commit whitespace hook then stripped trailing spaces from the committed copy and from one seat file, so the committed brief's digest differs from the manifest's by whitespace only (recorded in the arbitration). Seats (scripts/council/run_council.py, max-tokens 40000, timeout 1700 s, run 09:59:01Z): openai.gpt-6-astra ACCEPT-WITH-CORRECTIONS (2 🔴 / 5 🟡 / 2 🔵), grok-4.6 ACCEPT (0 / 0 / 5 🔵 / 6 💡), qwen3-coder ACCEPT (1 🔴 / 1 🟡 already addressed / 1 🔵). All three finish_reason=stop with content; no re-run. Arbitration and the reproductions follow in their own commits.
…/31 with the 44 sandbox ids named; tick T6.1–T6.3 Arbitration over the three seats: every red and yellow reproduced before acceptance (a RED test that failed for the seat's reason, or a probe of the renderer/markup), seven corrections landed one commit each, one claim refuted by a named test (qwen's empty-agent launch), three questions carried to the owner as decisions rather than invented (the abandonment contract, plan close on checked non-active plans, the session-db prompt probe), and the arbiter's own error recorded (brief §9's reconnect-purpose sentence contradicted the code; the code and its pin are right). Verify on the clean head d0251fd: 31 registered checks, 30 ok; the one red is full-suite-studyloop (30 failed / 5108 passed / 14 errors) whose new failed_nodes field lists exactly the 44 sandbox-environmental ids named in the item-4 control receipt — the receipt reconciles itself now, which is what failed_nodes was added for. plan-suites 518, browser journey 11, agent-session-tools 2146, both new checks green. Also: the active-learning delta said "Rule 8" where the code and rubric say rule 9 (GPT spec review); fixed.
…leted; the architect asks (owner decision 2026-09-18) Council review 6 left open whether plan close should refuse a fully-checked abandoned/paused/draft plan. The owner decided: it should be closable or deletable, and the agent asks the learner which. Four tests pin it: plan close on each of the three statuses launches once with the closing section's last line naming the status and both doors (set_study_plan_status to complete; delete_study_plan with confirmed=True) and 'ask', the status line naming the status, nothing written and the status unchanged; and the persona's Closing a Plan section names the three statuses, both tools, and keeps deletion behind the learner's explicit word. RED: 4 failed for the intended reasons (no Status line in the brief; the persona does not name the statuses).
…he learner's word — GREEN (owner decision 2026-09-18) plan close reviews a fully-checked draft, paused or abandoned plan like an active one (it never refused them; D-G did not restrict the command). What changes: the brief's closing section ends with a 'Status: <status> — not active; ask the learner whether to close it (set_study_plan_status(plan_id, "complete")) or delete it (delete_study_plan(plan_id, confirmed=True), only after they say so in this conversation)' line, and the command's status line names the state the plan is in. The command still writes nothing. Persona 'Closing a Plan' gains the rule: say what each door means (closing keeps the document as complete; deleting leaves only the checkpoint log), ask, delete only after the learner has said in so many words that this plan goes — the standing deletion rule holds, never because the status was abandoned — and offer closing as the reversible choice when they are unsure (verified: the seam accepts complete -> active). Three projections re-projected; manifest regenerated (updated restored on 20 unmoved entries); baseline refreshed whole-repo with the pinned detect-secrets 1.5.0 (Claude's entry replaced; OpenCode's new 16-hex hash is below the entropy threshold and has no entry; 72 -> 72 files). The cli-surface delta also catches up with review-6 F1/F2: the closing section has no separate Data gaps section any more, the proposal line may read 'unassessed — the review is partial', and learner-authored values are one-lined. Docs: study-plans.md and cli-reference.md say a paused, abandoned or never-activated plan can be closed or deleted the same way. 4 RED -> green; 157 across the seam/persona/install/docs/guidance files; ruff/format/pyright/mkdocs/openspec clean.
…s open items; the cross-harness rule Open item 2 (plan close on a checked non-active plan) is decided and landed (dfc75d6 RED, 983aa72 GREEN): close or delete, the architect asks. Open item 3 is reframed by the owner's standing rule — nothing in StudyLoop may be kiro-cli specific; every process and steering rule is stated for all six supported harnesses — so the question is no longer 'probe kiro-cli' but 'what can the harness-launched architect reach, per harness', and the arbitration now states that per harness as verified on the tree: Kiro (visible + ten trusted; session-db trust unmeasured), Claude (ten allow-listed, server in ~/.claude.json, session-db unreachable), OpenCode (both servers global, permission-block grants), Codex/Grok/pi (no named-agent feature; reached through StudyLoop's own harness-neutral launch chain — six adapters, one canonical persona, --agent threaded through plan repair and plan close; pi CLI-fallback by design). What remains owner-side is a per-harness reachability receipt on a real install, one row per harness. The rule is recorded in design.md's preamble so item 5 is written under it.
…ore it lands is a session like any other (owner decision 2026-09-18, review 6 F3b: option b) Design §2 promised 'navigate away / cancel before the console attaches → no live slot'. The web-ui spec and the landed test refused that (navigating away detaches with a grace period by design; End is the abandon path), and council review 6 asked the owner to settle which was the contract. The owner chose (b): a requested launch is a session like any other — leaving never destroys, End abandons, and there is no cancel for a pending launch, because a second meaning for 'leave' that depends on sub-second timing is the less predictable rule for the learner this design is for. (c) — recording cancellation as debt — was rejected: it would keep a wrong promise alive as a TODO. The barrier test holds the start POST while the learner navigates to Plans, releases it, and proves: the 201's session exists with purpose planning (and still does half a second later), no plan was created, returning to the Study Session view reattaches to the same session with one label and one start event, and End releases it. Discrimination proved by mutating the client to end any launch that lands after the learner left (option a): the test fails at the existence assertion; source restored byte-identical. Design §2 rewritten with the retraction and the reason; web-ui delta gains the scenario; the arbitration's open item 1 is closed. Journey module 12/12 -m e2e in natural order; ruff/format/pyright clean; openspec valid.
…settings journeys (CI run 35341660469) The e2e job on 795ba6b failed two pre-existing tests while 567 passed; both read a loading state as the answer on a runner 25 minutes into the job, and neither touches a path the branch changed (the two previous heads ran e2e green on identical product code). test_second_brain_ui::test_settings_shows_the_destination_command_pattern_when_it_is_missing counted `.brain-active` cards the instant the section was visible; the active class arrives with the launch-target response, so it read 0. Its sibling was given `active.first.wait_for(state="attached")` in 46262d2; this one now has the same wait. test_session_recovery_journey::TestStudyPickerRecovery::test_409_from_a_second_tab_offers_reattach_that_adopts_the_session created tab A's session over HTTP 300 ms after tab B's picker rendered, while tab B's init() /api/session/state fetch was still in flight; on the loaded runner that fetch landed after the session existed, tab B adopted it, and the Start button the test then clicked was hidden ("waiting for element to be visible" for 30 s). The test now waits for the timer's own settled signal (topic 'No active session', sessionActive false) before the session exists, so the 409 path is the one under test. Both modules 13 passed / 2 skipped locally in natural order; ruff/format clean.
…(CI run 35341660469 attempt 2) The e2e rerun on 795ba6b failed exactly one test, and a different one from attempt 1: test_abandoning_a_launch_mid_flight_leaves_no_session_and_no_plan read `assert [''] == []` — a purpose label still visible with empty text the instant /api/session/state reported the session gone. The label is cleared by the console's stop() on the study-session-stop event, an async path separate from the state poll the test gates on, so a bare DOM read can land one frame early. The test had passed on this branch's three earlier e2e runs; the assertion is right, its timing was not. It now waits for every label to be hidden before asserting, as the two sibling fixes in c8b832f do. Journey module 12/12 -m e2e locally; ruff/format clean.
…where the journey reads it (CI run 35347509248) Third e2e run, third head, fourth different test — one failure each time with 567-568 passing. All four are one class: a DOM read the frame the console's `connected` flips, while the purpose label's x-show is applied on the next frame (and its clearing arrives on the study-session-stop event). This run: test_console_is_labelled_planning_and_label_survives_reconnect read [] one frame early after the reload. One helper, _wait_for_planning_label, waits for exactly one rendered label and returns the labels; the three post-connect reads in the module (the reconnect test before and after reload, the barrier test's return to the console) go through it. The abandon test's cleared-label wait from 725b9e0 is the mirror case and stays. The two failures in attempt 1 were in modules that run before this one and were fixed in c8b832f; nothing leaks from the new barrier test into its neighbours. Journey module 12/12 -m e2e locally in natural order; ruff/format/pyright clean.
This was referenced Sep 18, 2026
This was referenced Sep 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Items 1–4 of the plan-integration follow-on programme: item 4 (
plan close <id>, evidence-based consensual completion, owner decision D-G) plus the council-review-6 corrections over the whole batch and the two verify checks design §6 asked for. Items 1–3b are already onmainvia #20.Item 4 —
CompletionReview(one definition inplanning/views.py) reads the end assessment as a preview (assess(phase="end", record=False), one call per fully-checked plan, no write, no status change) and proposesextendwhile any count on the plan's own concepts is above zero,closewhen all three are zero; the scheduler's "New topic — start fresh" rows are not counted as due (owner decision 2026-09-17).plan close <id>launches the architect with the review as the first section of its brief through the onestudychain; status moves only when the learner agrees. CLInow, the Today card and MCPget_next_actioncarry the review; recap prints the sentence. Rubric row 4b scored yes / yes by the owner.Council review 6 (
council/review-6-arbitration-2026-09-18.md,GATE: ACCEPT; three seats, every 🔴/🟡 reproduced before acceptance), seven corrections each RED-before-GREEN in its own commit:proposal: None,partial: true, the gap named as aNot read:evidence line on both surfaces;/api/session/optionscannot hold a click forever;studyloopserver is registered (~/.claude.json, from the installer's own constant), that the mentor's grants went live, and when to re-probe kiro-cli;one_line, shared with the Web door;plan close's three no-launch exits (mutation-proved).Verify —
scripts/verify/plan_integration.pygainsarchitect-grants(in-process, inventory-derived) andrepair-close-refusals, and a red pytest check now records itsfailed_nodes. Receiptverify-d0251fd1.json: 30/31, the one red being the sandbox's 44 known environmental ids, named on the receipt and identical to the item-4 control set.Tested
Full suite with matched control at item 4's GREEN (zero regressions); per-correction runs recorded in each commit; golden
now_plan_no_active.jsonsha unchanged; JS 139/139; browser journey 11/11; ruff / format / pyright / mkdocs --strict / openspec validate clean; all pre-commit hooks. CI on this PR is the first run to see the item-4 commits and the corrections.Open for the owner (arbitration §"Still open")
The abandonment contract for a pending planning launch (three alternatives recorded); whether
plan closeshould refuse a fully-checkedabandoned/draft/pausedplan; thesession-db"visible, prompts" probe.Merge as a local fast-forward once green (receipts cite SHAs by name), as with #20.