Skip to content

Item 4 plan close (D-G), council review 6 corrections, verify checks - #22

Merged
NetDevAutomate merged 23 commits into
mainfrom
feat/plan-close
Sep 18, 2026
Merged

NetDevAutomate merged 23 commits into
mainfrom
feat/plan-close

Conversation

@NetDevAutomate

Copy link
Copy Markdown
Owner

Summary

Items 1–4 of the plan-integration follow-on programme: item 4 (plan close <id>, evidence-based consensual completion, owner decision D-G) plus the council-review-6 corrections over the whole batch and the two verify checks design §6 asked for. Items 1–3b are already on main via #20.

Item 4 — CompletionReview (one definition in planning/views.py) reads the end assessment as a preview (assess(phase="end", record=False), one call per fully-checked plan, no write, no status change) and proposes extend while any count on the plan's own concepts is above zero, close when all three are zero; the scheduler's "New topic — start fresh" rows are not counted as due (owner decision 2026-09-17). plan close <id> launches the architect with the review as the first section of its brief through the one study chain; status moves only when the learner agrees. CLI now, the Today card and MCP get_next_action carry the review; recap prints the sentence. Rubric row 4b scored yes / yes by the owner.

Council review 6 (council/review-6-arbitration-2026-09-18.md, GATE: ACCEPT; three seats, every 🔴/🟡 reproduced before acceptance), seven corrections each RED-before-GREEN in its own commit:

  • a partial end assessment (a history reader down) never proposes a clean close — proposal: None, partial: true, the gap named as a Not read: evidence line on both surfaces;
  • the planning launch's wait for the picker's options is bounded (8 s) so a stalled /api/session/options cannot hold a click forever;
  • discovery text says only what the seam knows (provenance, doctor's healthy row) and a document doctor cannot read is named, not hidden behind "all ready";
  • the install doc names where Claude's studyloop server is registered (~/.claude.json, from the installer's own constant), that the mentor's grants went live, and when to re-probe kiro-cli;
  • the Today card keeps each closing review's evidence with its plan and labels it "Closing review", not "Plan complete";
  • the CLI repair/closing briefs contain learner-authored values through the seam's one_line, shared with the Web door;
  • six tests pin the struggle-only / unverified-only branches, the evidence cap, and plan close's three no-launch exits (mutation-proved).

Verify — scripts/verify/plan_integration.py gains architect-grants (in-process, inventory-derived) and repair-close-refusals, and a red pytest check now records its failed_nodes. Receipt verify-d0251fd1.json: 30/31, the one red being the sandbox's 44 known environmental ids, named on the receipt and identical to the item-4 control set.

Tested

Full suite with matched control at item 4's GREEN (zero regressions); per-correction runs recorded in each commit; golden now_plan_no_active.json sha unchanged; JS 139/139; browser journey 11/11; ruff / format / pyright / mkdocs --strict / openspec validate clean; all pre-commit hooks. CI on this PR is the first run to see the item-4 commits and the corrections.

Open for the owner (arbitration §"Still open")

The abandonment contract for a pending planning launch (three alternatives recorded); whether plan close should refuse a fully-checked abandoned/draft/paused plan; the session-db "visible, prompts" probe.

Merge as a local fast-forward once green (receipts cite SHAs by name), as with #20.

Seven tests pin scenario 4's redesign before any production code changes:
the completion action rule 9 emits for a fully-checked active plan must
carry the end assessment (due reviews, struggles and unverified
milestones counted on the plan's own concepts), propose `extend` while
any count is above zero and `close` when all are zero, compose its
sentence from that proposal, and never change a status — the read is
`assess(AssessPlan(phase="end", record=False))`, the preview path, so
the document, its status and the checkpoint log are untouched. A failed
assessment keeps the pre-change sentence and adds a warning; `now` never
fails on it. `plan close <id>` is the launch sibling of `plan repair`:
the same architect chain with a `### Closing review` first section
(three counts, proposal, evidence lines), refusing a plan with open
milestones with their count.

Owner decision 2026-09-17: the completion review counts only due rows
that name a concept. The scheduler's "New topic -- start fresh" row
(`concept: None`, `evidence: configured_topic`) is a cold-start hint for
"what should I review now", not a lapsed review; the evaluator already
ignores it at concept level (it never matches a milestone concept, so it
contributes nothing to `unverified_milestones`), and counting it would
tell a learner who has just ticked every milestone to "start fresh".
`plan evaluate` keeps the row — the exclusion is the completion
review's, applied by one definition in both the engine and the brief.

Evidence is planted on the `studyloop.history` package attributes the
evaluation resolves at call time, so the relevance filter and
`has_evidence` logic stay live and no sessions database is involved.
All seven fail for the intended reason (missing attributes, `assess`
never called, no warning, no `close` command); the 57 pre-existing
tests in both files pass and the no-plan golden sha is unchanged.

The first test's T4.1 name lost one redundant word (`plan_concepts` →
`concepts`): no def line in the repo exceeds 100 chars and none carries
a `noqa`. The `pyright: ignore[reportAttributeAccessIssue]` tags on the
not-yet-existing attributes are the 3b RED's pattern; GREEN strips them.
…tem-4 RED pins

Design §4 now states the closing-review brief format, intro and refusal
text the RED fixed; records the owner's 2026-09-17 decision that the
completion review counts only due rows naming a concept (the scheduler's
"new topic" row is excluded, with the reasoning and the measured cost of
the preview); and keeps the one decision deliberately left open — what
`proposal` holds when the end assessment fails — with the default GREEN
will take unless vetoed (nullable proposal, counts stay int).
…(D-G)

A fully-checked active plan used to get an either-way sentence ("close the
plan or extend it") that the owner scored "no as phrased" on rubric row 4:
it proposed nothing and rested on nothing. The completion action is now a
closing review read from the plan's own end assessment, and `plan close <id>`
is the door to acting on it — with the learner, never automatically.

One definition, two surfaces. `planning.views.CompletionReview` reads the
`end` evaluation into three counts on the plan's concepts (due reviews,
struggles, milestones marked done with no evidence), a proposal (`extend`
while any count is above zero, else `close`) and one evidence line per
counted item (capped at 8). Both the `now` engine's `CompletionAction` and
the `plan close` brief consume it, so they cannot disagree on a count.

Due reviews count only rows that name a concept (owner decision,
2026-09-17): the scheduler's `New topic -- start fresh` row is a cold-start
hint for "what should I review now", not a lapsed review, and counting it
would tell a learner who just ticked every milestone to start fresh.
`plan evaluate` keeps the row; the exclusion is the completion review's.

The engine reads through the preview path — `assess(AssessPlan(phase="end",
record=False))`, one call per fully-checked plan per `build_now_plan` — so
the document, its status and the checkpoint log are untouched. A failed
assessment keeps the plain sentence with `proposal` None and the counts 0,
plus a warning naming the plan: the counts are then unknown, not zero, and no
renderer reads a clean slate or outstanding work into a failure. Measured on
the live 877 MB sessions.db the preview costs ~320 ms per fully-checked plan,
a transient state by design. New JSON keys appear only inside
`completion_actions` entries; the no-plan golden is byte-identical.

`plan close <id>` mirrors `plan repair`: `_inspect`, refuse with the open
count (exit 1) while any milestone is open, leave a `complete` plan alone,
and for a fully-checked plan launch the architect through the one
`study --mode plan-architect` chain with `### Closing review` as the brief's
first section (the four count/proposal lines, then the evidence). The shared
"plan as it stands" section is refactored out of the repair brief
byte-identically. The persona gains "Extend or close" (three projections
re-projected, manifest regenerated with `updated` restored on unmoved
entries, secrets baseline refreshed whole-repo: 72 -> 72 files, exactly the
two manifest hashes). CLI `now` prints the evidence lines beneath the
sentence; the Today card gains `completionEvidence()` for the same lines
(JS 136/136); the recap speaks the composed sentence. Docs, two spec deltas,
and rubric row 4b (re-run primary and both proposals, verdict PENDING for
the owner; row 4 kept as the record of the finding).

Verification: 7 RED -> green (64/64 across the two files); golden sha
ec451ce8 unchanged; `just lint`, `just typecheck`, `mkdocs build --strict`,
`openspec validate` clean. Full suite 30 failed / 7233 passed / 14 errors,
diffed against a clean f1c52ce control run in parallel: item4 - control =
empty, control - item4 = exactly the seven REDs, 44 shared environmental ids
committed by name in receipts/full-suite-control-item4-2026-09-18.md.
…em 4's GREEN-time decisions

T4.2 GREEN is 8229329 (7 RED -> green, matched-control full suite with zero
regressions, 44 environmental ids now committed by name); T4.3's rubric row
4b is in the same commit with the verdict pending the owner. Design section 4
closes its one open decision - `proposal` is nullable, None on a failed
assessment, represented once at the action - and records the two decisions
taken in GREEN: the Today card renders the review's evidence lines, and the
docs' pinned six-boundary list stays at six with the consensual close stated
beside it.
…oth readings

The owner scored rubric row 4b on 2026-09-18: (a) the evidence-backed
'extend' proposal is one they would walk, (b) a clean review's 'close' is
one they would agree to. This closes row 4's 'no as phrased' finding from
2026-09-16; row 4 stays as the record of the original finding.

Recorded in the receipt's status line and verdict cell exactly as given
(no rationale was supplied beyond the two answers, and none is invented),
in tasks.md (T4.3 note; T5.5's archive condition now names 3b as the one
row still outstanding), and on the public study-plans page, which said the
re-run row 'awaits the maintainer's score'. The contract pin keys on
'owner verdicts RECORDED' and stays satisfied.
…e repair/close refusals; a red suite names its nodes (T6.3)

Design §6 of plan-integration-followons asks the verification receipt for
two more registered checks beside the golden: the two architect grants (the
nine plan tools + record_plan_learning, derived from the inventory, in the
spelling each harness honours) and the plan repair / plan close refusal
texts as the CLI seam tests pin them. Five tests pin their names, shapes
and one negative case (a copy of the tree with one grant removed and one
stray grant added must fail naming both).

The fifth test is receipt honesty: the first real verify run on 9d10fee
recorded full-suite-studyloop as ok:false with a 12-line output_tail, so
the receipt could not say WHICH of the sandbox's 30 failures + 14 errors
occurred and could not be reconciled against the 44 named environmental
ids. A failed pytest check now keeps every FAILED/ERROR short-summary line
on its row.

RED: 5 failed / 21 passed, each for the intended missing name, function or
key.
… checks; a red pytest check names its failed nodes (T6.3) — GREEN

Design §6 of plan-integration-followons: two registered checks beside the
golden. `architect-grants` is in-process — it derives the ten names from
studyloop.mcp.inventory (PLAN_TOOL_NAMES + LEARNING_RECORD_TOOL) and judges
agents/kiro/study-plan-architect.json (`@studyloop` visible in `tools`;
allowedTools carries exactly the ten as `@studyloop/<tool>`, no bare server
grant, no inert mcp_studyloop_* spelling) and the Claude architect's
`tools:` frontmatter (exactly the ten as `mcp__studyloop__<tool>`, no other
server), measuring what it found so the receipt reads without the files.
`repair-close-refusals` runs exactly the four CLI seam node ids that pin the
husk refusal (both exits), `plan repair` on a ready plan and an unknown id,
and `plan close` on an unfinished plan.

Receipt honesty: a failed pytest check now keeps every FAILED/ERROR
short-summary node id on its row (`failed_nodes`), so a red full suite can
be reconciled against the named environmental set from the receipt alone —
the first real run on 9d10fee could not be.

Registry 29 -> 31. Verify-script tests 5 RED -> 26 passed; both new checks
pass on this tree; ruff, format, pyright clean.
…ncil review 6, F1)

evaluate_plan turns a history reader that fails into one warning
('<reader> unavailable — evaluation is partial') and an empty default, so a
count read while that reader was down is unread, not zero. CompletionReview
read those zeros as a clean slate: the now engine said 'the closing review is
clean — it proposes closing the plan' and plan close printed 'proposes:
close' with the gap relegated to a separate section. GPT-Astra 🔴 F1, Grok 🔵
(v); reproduced with two RED tests before this change.

The fix lives in the one definition. CompletionReview keys on the evaluator's
own marker (PARTIAL_READ_MARKER, defined once in evaluation.py and used by
_safe), keeps the counts it did read, sets partial=True, proposes None —
neither close nor extend — and names each gap among its evidence lines
('Not read: …'), so both surfaces carry it without a second path.
CompletionAction gains partial; the sentence has a third branch that says the
review is partial and could not propose; plan close's proposal line reads
'unassessed — the review is partial' and its status line no longer says the
review proposes. The separate '### Data gaps' brief section is gone: the gaps
are the review's own lines now.

Persona 'Closing a Plan' tells the architect what an unassessed proposal
means (walk what was read, prefer re-running the review, never infer a clean
slate); three projections re-projected, manifest regenerated (updated
restored on the 20 unmoved entries), baseline refreshed whole-repo with the
pinned detect-secrets 1.5.0 (exactly the two manifest hashes, 72 -> 72
files). Spec delta states the rule and adds the scenario.

66/66 across the two item-4 files (2 RED -> green), 614 across the plan
suites + pins, JS 136/136, golden sha unchanged, ruff/pyright/mkdocs/openspec
clean.
…council review 6)

626ea12 made a planning click await init()'s /api/session/options fetch
before judging whether an agent exists, so a cold server no longer refused
with 'Select an agent'. That put an unbounded network wait on the click's
critical path: a request that never settles (a server that accepts the
connection and stalls) held the learner's click forever — no POST, no
refusal, no message. qwen3-coder 🔴, GPT-Astra F3, Grok 🔵 (i); reproduced by
a JS test that never settles the fetch (timed out at 5 s).

The wait now races the options promise against optionsWaitMs (8 s, on the
timer's state so tests can shorten it); past the bound the launch judges the
agent as it stands and gives the picker's own refusal, and a settlement that
arrives later launches nothing on its own. The two existing wait tests
(resolve-then-launch once; empty picker still refuses) are unchanged.

JS 137/137 (+1), browser journey 11/11 -m e2e; web-ui spec delta states the
wait and its bound.
…ument is named (council review 6, F4)

Three sentences claimed more than the evidence. husk_provenance read a
creation stamp before the gate as 'was never judged by it' — but a plan
created before 2026-09-15 can be saved ready after it and hand-edited into a
husk later; the stamp establishes when the document was created, not what
judged it or when it became incomplete. It now says 'This plan's creation
stamp predates the readiness gate (…); the seam cannot tell when it became
incomplete.' doctor's healthy row promised 'every write the gate judges will
pass' — a future write can remove a required field — and now says 'all ready
as they stand.' And husks() skipped a document it could not load, so doctor
could report all-ready over a directory holding a file no listing can read.
GPT-Astra 🔴 F4; reproduced with a parametrised provenance test and two
doctor tests before this change (the first fixture was not unreadable at
all — the parser is lenient with malformed front matter — the real class is
a non-UTF-8 file or a filename that is not a valid id).

PlanApplication.survey_husks() is the one pass that returns both facts
(HuskSurvey: husks + unreadable ids); husks() is now a view over it, so
plan list/repair are unchanged. doctor emits one warn row per unreadable
document beside the readiness rows, fix_auto=False. Health and cli-surface
deltas state the honest wording and the new row.

312 across the pinning files, guard 30/30, ruff/format/pyright clean,
openspec valid.
… mentor grants went live, and when to re-probe kiro-cli (council review 6, F5)

The Claude architect's frontmatter names ten mcp__studyloop__ tools, and no
diff in the reviewed range showed where the server itself is declared for
Claude Code — so three seats asked whether the grant was inert the way the
mentor's had been (GPT F5 🟡, Grok 🔵 c). It is not: installers._MCP_HARNESSES
includes claude and `studyloop install agents` merges both servers into
~/.claude.json's mcpServers. The install doc now says so, and the pin takes
the path from installers._mcp_config_path('claude') rather than a remembered
string, so the sentence cannot drift from the code.

Two more disclosures the seats asked for: existing Kiro study-mentor installs
start seeing the twelve MCP tools their file always named once re-installed
(f5c2057 fixed the inert spelling; nothing a user read said so — Grok 💡 d,
GPT 🔵), and the grant spelling is evidence pinned to kiro-cli 2.21.4/2.22.0,
so the doc and the probe receipt's header both say to re-run probes A and B
on a newer binary. The receipt also states that the session-db
'visible, prompts' reading is an expectation, not a measurement (Grok 🔵 a).

Persona/install/docs/prompt-contract pins 137 passed; mkdocs --strict clean.
…s plan; the label says review, not complete (council review 6, F6)

The card rendered every completion sentence, then every evidence line
flattened beneath them, so with two finished plans a line lost the plan it
belonged to — the contextual review D-G asks for, weakened at the surface
most learners read. And both the card and CLI now labelled the note 'Plan
complete' over a plan whose status is still active until the learner agrees
with the architect. GPT-Astra F6 🟡; reproduced from the markup.

today-panel.js gains completionReviews(): one block per action — sentence and
its own evidence lines — keyed by plan_id in the engine's order; the flat
completionNotes()/completionEvidence() helpers now derive from it. index.html
renders one .today-plan-review block per plan (data-plan-id) with the lines
nested inside, labelled 'Closing review'; cli/_now.py prints the same label.
A JS test pins the grouping (an extend review, a partial review carrying its
'Not read:' line, and a pre-D-G entry) and a markup test pins the keyed
block, the nesting and the absence of both a flat evidence loop and the word
'Plan complete'. No test had pinned either old label.

JS 139/139 (+2), CLI now/guidance/seam 67/67, e2e Today/plan subset 29
passed; design §4 records the correction.
…ues the way the Web brief does (council review 6, F2)

The Web door one-lines every value the planning brief quotes (review-3 F4);
the CLI plan repair / plan close briefs interpolated the plan's title, id,
topics, created stamp and the review's evidence lines raw. A YAML-quoted
front-matter title holding '\n## Forged section' survives parse_plan, and
_render_plan_as_it_stands printed it as a real heading inside the brief the
architect reads. GPT-Astra F2 🟡; reproduced by rendering such a document
(the '##' line was a heading of the brief).

The containment is now one definition on the seam,
studyloop.planning.one_line (whitespace runs, newlines included, collapse to
one space, so no value can start a line): the CLI briefs quote every
learner-authored value through it, and the Web door's _one_line delegates to
it, so every adapter's brief is contained the same way. The repair brief's
blocker lines are seam-authored sentences and are left as they are. The
architecture guard passes: one_line is a studyloop.planning import (D-6).

RED test pins both briefs against a hostile title, topic and evidence line;
guard + CLI seam + Web brief tests 119 passed; ruff/pyright clean.
…only branches, the evidence cap, and plan close's three no-launch exits (council review 6, F7)

The seven item-4 REDs established due-work extension, clean closure,
new-topic exclusion, the exception fallback, no plan writes and the two
principal CLI paths. GPT-Astra F7 🟡 named what they did not: a struggle
alone or an unverified milestone alone carrying 'extend', the evidence cap's
overflow line, and plan close on an already-complete plan, on a plan with no
milestones, and on an unknown id. Six tests now pin those branches; the
zero-milestone test also proves no assessment is read on that exit.

Discrimination proved by mutation: with the proposal rule changed to read the
due count alone, exactly the struggle-only and unverified-only tests fail and
the rest stay green; the source was restored byte-identical.

73 passed across the two files; ruff/pyright clean.
Brief over items 1-4 of plan-integration-followons, reviewed tree 9d10fee
(range 1565234..9d10fee: 29 commits, 73 files): the owner's decisions and
the four taken during the batch, design §1-§4 verbatim, the agent's own
T-notes verbatim, every diff grouped by item, the seven outside-the-items
commits classified, reference facts each verified on the tree, and ten
numbered deliverables. 6,082 lines, ~95k tokens; the manifest records the
sha256 of the bytes the seats received (7ef91328…) — the pre-commit
whitespace hook then stripped trailing spaces from the committed copy and
from one seat file, so the committed brief's digest differs from the
manifest's by whitespace only (recorded in the arbitration).

Seats (scripts/council/run_council.py, max-tokens 40000, timeout 1700 s, run
09:59:01Z): openai.gpt-6-astra ACCEPT-WITH-CORRECTIONS (2 🔴 / 5 🟡 / 2 🔵),
grok-4.6 ACCEPT (0 / 0 / 5 🔵 / 6 💡), qwen3-coder ACCEPT (1 🔴 / 1 🟡 already
addressed / 1 🔵). All three finish_reason=stop with content; no re-run.
Arbitration and the reproductions follow in their own commits.
…/31 with the 44 sandbox ids named; tick T6.1–T6.3

Arbitration over the three seats: every red and yellow reproduced before
acceptance (a RED test that failed for the seat's reason, or a probe of the
renderer/markup), seven corrections landed one commit each, one claim
refuted by a named test (qwen's empty-agent launch), three questions carried
to the owner as decisions rather than invented (the abandonment contract,
plan close on checked non-active plans, the session-db prompt probe), and the
arbiter's own error recorded (brief §9's reconnect-purpose sentence
contradicted the code; the code and its pin are right).

Verify on the clean head d0251fd: 31 registered checks, 30 ok; the one red
is full-suite-studyloop (30 failed / 5108 passed / 14 errors) whose new
failed_nodes field lists exactly the 44 sandbox-environmental ids named in
the item-4 control receipt — the receipt reconciles itself now, which is what
failed_nodes was added for. plan-suites 518, browser journey 11,
agent-session-tools 2146, both new checks green.

Also: the active-learning delta said "Rule 8" where the code and rubric say
rule 9 (GPT spec review); fixed.
…leted; the architect asks (owner decision 2026-09-18)

Council review 6 left open whether plan close should refuse a fully-checked
abandoned/paused/draft plan. The owner decided: it should be closable or
deletable, and the agent asks the learner which. Four tests pin it: plan
close on each of the three statuses launches once with the closing section's
last line naming the status and both doors (set_study_plan_status to
complete; delete_study_plan with confirmed=True) and 'ask', the status line
naming the status, nothing written and the status unchanged; and the
persona's Closing a Plan section names the three statuses, both tools, and
keeps deletion behind the learner's explicit word.

RED: 4 failed for the intended reasons (no Status line in the brief; the
persona does not name the statuses).
…he learner's word — GREEN (owner decision 2026-09-18)

plan close reviews a fully-checked draft, paused or abandoned plan like an
active one (it never refused them; D-G did not restrict the command). What
changes: the brief's closing section ends with a 'Status: <status> — not
active; ask the learner whether to close it (set_study_plan_status(plan_id,
"complete")) or delete it (delete_study_plan(plan_id, confirmed=True), only
after they say so in this conversation)' line, and the command's status line
names the state the plan is in. The command still writes nothing.

Persona 'Closing a Plan' gains the rule: say what each door means (closing
keeps the document as complete; deleting leaves only the checkpoint log),
ask, delete only after the learner has said in so many words that this plan
goes — the standing deletion rule holds, never because the status was
abandoned — and offer closing as the reversible choice when they are unsure
(verified: the seam accepts complete -> active). Three projections
re-projected; manifest regenerated (updated restored on 20 unmoved entries);
baseline refreshed whole-repo with the pinned detect-secrets 1.5.0 (Claude's
entry replaced; OpenCode's new 16-hex hash is below the entropy threshold and
has no entry; 72 -> 72 files).

The cli-surface delta also catches up with review-6 F1/F2: the closing
section has no separate Data gaps section any more, the proposal line may
read 'unassessed — the review is partial', and learner-authored values are
one-lined. Docs: study-plans.md and cli-reference.md say a paused, abandoned
or never-activated plan can be closed or deleted the same way.

4 RED -> green; 157 across the seam/persona/install/docs/guidance files;
ruff/format/pyright/mkdocs/openspec clean.
…s open items; the cross-harness rule

Open item 2 (plan close on a checked non-active plan) is decided and landed
(dfc75d6 RED, 983aa72 GREEN): close or delete, the architect asks. Open
item 3 is reframed by the owner's standing rule — nothing in StudyLoop may be
kiro-cli specific; every process and steering rule is stated for all six
supported harnesses — so the question is no longer 'probe kiro-cli' but 'what
can the harness-launched architect reach, per harness', and the arbitration
now states that per harness as verified on the tree: Kiro (visible + ten
trusted; session-db trust unmeasured), Claude (ten allow-listed, server in
~/.claude.json, session-db unreachable), OpenCode (both servers global,
permission-block grants), Codex/Grok/pi (no named-agent feature; reached
through StudyLoop's own harness-neutral launch chain — six adapters, one
canonical persona, --agent threaded through plan repair and plan close; pi
CLI-fallback by design). What remains owner-side is a per-harness
reachability receipt on a real install, one row per harness.

The rule is recorded in design.md's preamble so item 5 is written under it.
…ore it lands is a session like any other (owner decision 2026-09-18, review 6 F3b: option b)

Design §2 promised 'navigate away / cancel before the console attaches → no
live slot'. The web-ui spec and the landed test refused that (navigating away
detaches with a grace period by design; End is the abandon path), and council
review 6 asked the owner to settle which was the contract. The owner chose
(b): a requested launch is a session like any other — leaving never destroys,
End abandons, and there is no cancel for a pending launch, because a second
meaning for 'leave' that depends on sub-second timing is the less predictable
rule for the learner this design is for. (c) — recording cancellation as debt
— was rejected: it would keep a wrong promise alive as a TODO.

The barrier test holds the start POST while the learner navigates to Plans,
releases it, and proves: the 201's session exists with purpose planning (and
still does half a second later), no plan was created, returning to the Study
Session view reattaches to the same session with one label and one start
event, and End releases it. Discrimination proved by mutating the client to
end any launch that lands after the learner left (option a): the test fails at
the existence assertion; source restored byte-identical. Design §2 rewritten
with the retraction and the reason; web-ui delta gains the scenario; the
arbitration's open item 1 is closed.

Journey module 12/12 -m e2e in natural order; ruff/format/pyright clean;
openspec valid.
…settings journeys (CI run 35341660469)

The e2e job on 795ba6b failed two pre-existing tests while 567 passed;
both read a loading state as the answer on a runner 25 minutes into the job,
and neither touches a path the branch changed (the two previous heads ran
e2e green on identical product code).

test_second_brain_ui::test_settings_shows_the_destination_command_pattern_when_it_is_missing
counted `.brain-active` cards the instant the section was visible; the active
class arrives with the launch-target response, so it read 0. Its sibling was
given `active.first.wait_for(state="attached")` in 46262d2; this one now has
the same wait.

test_session_recovery_journey::TestStudyPickerRecovery::test_409_from_a_second_tab_offers_reattach_that_adopts_the_session
created tab A's session over HTTP 300 ms after tab B's picker rendered, while
tab B's init() /api/session/state fetch was still in flight; on the loaded
runner that fetch landed after the session existed, tab B adopted it, and the
Start button the test then clicked was hidden ("waiting for element to be
visible" for 30 s). The test now waits for the timer's own settled signal
(topic 'No active session', sessionActive false) before the session exists,
so the 409 path is the one under test.

Both modules 13 passed / 2 skipped locally in natural order; ruff/format clean.
…(CI run 35341660469 attempt 2)

The e2e rerun on 795ba6b failed exactly one test, and a different one from
attempt 1: test_abandoning_a_launch_mid_flight_leaves_no_session_and_no_plan
read `assert [''] == []` — a purpose label still visible with empty text the
instant /api/session/state reported the session gone. The label is cleared by
the console's stop() on the study-session-stop event, an async path separate
from the state poll the test gates on, so a bare DOM read can land one frame
early. The test had passed on this branch's three earlier e2e runs; the
assertion is right, its timing was not. It now waits for every label to be
hidden before asserting, as the two sibling fixes in c8b832f do.

Journey module 12/12 -m e2e locally; ruff/format clean.
…where the journey reads it (CI run 35347509248)

Third e2e run, third head, fourth different test — one failure each time
with 567-568 passing. All four are one class: a DOM read the frame the
console's `connected` flips, while the purpose label's x-show is applied on
the next frame (and its clearing arrives on the study-session-stop event).
This run: test_console_is_labelled_planning_and_label_survives_reconnect read
[] one frame early after the reload.

One helper, _wait_for_planning_label, waits for exactly one rendered label
and returns the labels; the three post-connect reads in the module (the
reconnect test before and after reload, the barrier test's return to the
console) go through it. The abandon test's cleared-label wait from 725b9e0
is the mirror case and stays. The two failures in attempt 1 were in modules
that run before this one and were fixed in c8b832f; nothing leaks from the
new barrier test into its neighbours.

Journey module 12/12 -m e2e locally in natural order; ruff/format/pyright
clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant