Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,12 @@ breaking changes may land in a minor release.

- Escalate an environment fault at the review-budget rescue gate instead of
deferring the story as unconverged (DW-523).
- Explain that unpinned result-artifact scans search only the configured artifact
directories themselves, so a nested story spec no longer produces an opaque
`no-artifact` breadcrumb (#780).
- Surface why a sprint-mode dev session found no result: `session-end` carries
the last resultless verdict, the `no-artifact` crumb names specs found one level
down, and `validate` warns `queue.nested-specs` on a nested spec layout (#780).

## [0.13.1] — 2026-10-01

Expand Down
1 change: 1 addition & 0 deletions docs/FEATURES.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,6 +85,7 @@ See [README.md](../README.md) for the narrative overview and [setup-guide.md](se
- Hook events from a nested coding CLI are attributed, not trusted (#767). A CLI process started from inside a session inherits the relay environment, so its `SessionStart`/`Stop`/`SessionEnd` land in the parent's event stream under the parent's task id and could complete or crash it. The generic adapter applies a second layer after the task-id filter (`SessionAttribution`): the first `SessionStart` is the launched session's — identified or not, so an anonymous start (a payload the relay could not read) still takes the parent's slot — and a later `SessionStart` with a new id marks that id foreign, unless its `source` is `clear` or `compact`, which is the launched session rotating its id and rebinds (`resume` and `startup` stay foreign). `source` alone does not rebind: an id already found foreign stays foreign (a child compacting), and a `clear` start right after a foreign id's `SessionEnd` is that child clearing, so its new id is foreign too. A foreign id's events are dropped before they can complete the session, re-point its transcript, spend nudges or re-arm the stall timer; the first drop per foreign id writes a `foreign-hook-event-ignored` breadcrumb (`hook_event`, `foreign_session_id`, `dropped_so_far`), every drop counts into `foreign_hook_events` on `heartbeat.json` and the `timeout-fired` crumb, and the sweep's failed-session diagnostic replays the same rule (`hook_foreign_ids`, `; foreign ignored: N` in the escalation suffix). The rule fails toward acceptance: id-less events, ids that never announced a `SessionStart` (a rotated id, a copilot `toolu_…` subagent `Stop`) and everything before the first start — including an identified `SessionEnd` from a CLI that exited before its `SessionStart` fired (#727) — are admitted, so attribution only ever drops a known child's events and never adds a completion path. A profile declaring `session_id_flag` pins the launched session's id instead (DW-505/508): the `claude` profile ships `--session-id`, so the adapter mints a UUID4 per launch, passes it on the command line and knows the parent before any event arrives — every id ever bound (the pin and each `clear`/`compact` rebind) is the launched session's own, and an identified `SessionStart` or `SessionEnd` from any other id is foreign, including a child `SessionEnd` whose child never announced a start; other events (`Stop` included) from an unannounced id are still admitted. If the CLI does not honour the pin — its first identified `SessionStart` that is not a `clear`/`compact` rotation reports another id (a future claude ignoring `--session-id`, or an `extra_args`/`launch_args` overlay adding `--resume`, `--continue` or `--fork-session`) — attribution treats the launched session as foreign, so its own `Stop` is dropped and it ends only on window death or its timeout. One `pinned-session-id-mismatch` breadcrumb (`pinned_session_id`, `reported_session_id`, `source`) makes that visible (DW-509). It is observation only and leaves the attribution rules unchanged. A start the pin rejects and the relay tags `mismatch` (a nested CLI, e.g. one a parallel `SessionStart` hook launched, reaching the events directory first) is not checked; the next start is. So a CLI whose own starts all read `mismatch` (a miscalibrated relay, below) and that ignores the pin writes no crumb. The check needs the launched session's own identified start: if that start was anonymous (a payload the relay could not read) or was lost behind a `clear`/`compact` rotation, the next identified non-rotation start is the first one checked, and it may be a nested child's. The crumb then names the child's id while the launched session completes normally, so read it as "the pin was not confirmed" rather than as proof the session is stuck. Accepted limitation, for profiles without `session_id_flag` (codex, gemini, copilot, antigravity): a child `SessionEnd` whose child never announced a `SessionStart` is indistinguishable from the parent's own and is admitted; nested CLIs announce their start. The sweep's failed-session diagnostic always replays the unpinned rule. A profile that maps no `SessionStart` (Stop-only, e.g. `antigravity`) gets no nested-CLI protection. A relay vendored before this change forwards no `source`, so there a `clear`/`compact` start with a new id reads as foreign and the session falls back to window death or its timeout — re-run `bmad-loop init` to re-vendor the relay. Relay-side lineage closes the remaining gap (DW-507). A tmux launch runs the window command behind a POSIX prelude, `/bin/sh -c 'BMAD_LOOP_LAUNCH_PID=$$; export BMAD_LOOP_LAUNCH_PID; exec "${SHELL:-/bin/sh}" -c "$1"' sh <command>`, so the variable holds the launched pid and the command runs under tmux's `default-shell` exactly as before (tmux sets `SHELL` to it in every pane), so whatever that shell sources for a `-c` command (fish's `config.fish`, zsh's `.zshenv`) still applies. Both relays (`bmad-loop relay` and the legacy copied script) tag every event `lineage`: `match` when the launched CLI itself fired the hook, `mismatch` when a nested CLI did, `unknown` when that cannot be read. The relay walks its parent chain from `/proc` and skips only two kinds of process: the launch chain — any process started within 5 s of the launched pid, which covers a shell that forks the CLI (fish, or dash as Debian/Ubuntu `/bin/sh`: the launched pid is the shell, the CLI its child) and a node shim's real binary a level below (a limitation: anything the CLI itself starts at launch, such as an MCP server or a project `SessionStart` hook running another CLI, is inside the window too and reads `match`, so only the rules above apply to it) — and a hook-command wrapper, read structurally from its argv: a shell (`sh`, `bash`, `zsh`, `fish`, …) whose `-c` command string holds the relay invocation (`…/bmad-loop relay` or `…/bmad_loop_hook.py`) followed by the event name as consecutive words, split on whitespace, quotes and shell operators (so `sh -c '…/bmad-loop relay Stop && true'` counts), or a runner whose argv ends in it (`uv run --no-project python …/bmad_loop_hook.py Stop`). Nothing else in a process's argv is read, so a nested CLI whose prompt merely names the relay is still a nested CLI. Any other process on the way means a nested CLI. The relays only tag, never drop. The launched session's own first `SessionStart` calibrates lineage (with a pin, the first identified start the pin admits, so a child's start that reaches the events directory first calibrates nothing, identified or anonymous): tagged `match`, it is trusted, and every later `mismatch` event is foreign — id-less events and a child's `clear`/`compact` rotation without a preceding `SessionEnd` included, which the rules above alone would admit or rebind. Tagged `mismatch` (the CLI's hook architecture defeats the walk) or `unknown`/untagged (psmux on Windows and macOS have no `/proc`; a relay that predates the tag sends none), lineage is ignored for the attempt, the rules above apply unchanged, and one `hook-lineage-untrusted` breadcrumb (`reason`: `miscalibrated` or `unavailable`, `lineage`: the first start's tag) records the degrade. Lineage needs a `SessionStart` to calibrate on: a Stop-only profile (e.g. `antigravity`) never engages it and writes no crumb. An event carrying one of the launched session's own ids is never made foreign by lineage, and events before calibration ignore it. Id-less drops share one `foreign-hook-event-ignored` crumb with `foreign_session_id: null`, and the sweep's failed-session diagnostic, which replays the same rule, counts them as `hook_foreign_idless` (`; foreign id-less events ignored: N` in the escalation suffix).
- Transport faults in the generic (multiplexer-driven) adapter leave crumbs instead of healthy-looking answers (DW-447/449/453/454). Like `session-probe-failed`, a window-liveness probe that raises `MultiplexerError` never reads as death, and every verdict is unchanged; what changes is the record. `liveness-probe-failed` (`site`, `error`) is written by the `tick` site at a wait-loop streak's first failed tick only, never once per tick, and by the one-shot sites (`over-budget`, `stall`, `post-kill`) once per verdict probe that raised. `liveness-probe-recovered` (`failures`) closes a tick streak when a later tick probes cleanly; a session that ends mid-streak (Stop, SessionEnd, abort, over-budget) leaves no recovered crumb. `heartbeat.json` carries the running streak as `probe_failures`, and `timeout-fired` carries the final one. `over-budget-fired` and `kill-escalated` gain `liveness_unknown`, which is true when the probe behind that verdict (for the kill, the last poll before escalation) raised. A stall, budget, stop or contract nudge whose send raised writes `nudge-send-failed` (`nudge`, `error`) and is never reported as sent: `stall_nudges_sent` on the heartbeat counts delivered nudges, `stall_nudges_failed` counts the failed attempts, and `contract-nudge-sent` is written only after a successful send (a failed contract nudge is still not retried). The stall-nudge cap and the #727 activity window count attempts, delivered or not. A post-kill rescue abandoned on unknown liveness or an unreadable artifact writes `post-kill-rescue-abandoned` (`reason` = `liveness-unknown` / `unreadable-artifact`, `status`, plus `error` for the read fault). The opencode HTTP adapter's heartbeat carries neither new key (its `stall_nudges_sent` counts attempts), but its nudges are crumbed the same way (DW-503): its `send_text` raises a `MultiplexerError` when the prompt POST fails, so an undelivered contract nudge writes `nudge-send-failed` instead of `contract-nudge-sent`, and its own budget, stall and stop nudges write `nudge-send-failed` and carry on exactly as before. Its dev sessions do share the post-kill reconcile, so they write `post-kill-rescue-abandoned` too, on the `unreadable-artifact` arm only (their liveness probe never answers "unknown").
- Three generic-adapter observation folds are crumbed the same way, verdicts unchanged (DW-448/450/451). A stall-expiry look at the pane whose capture raises `MultiplexerError`, or whose parked-prompt search blows `PARKED_PROMPT_MATCH_TIMEOUT_S`, still reads as "not parked", and writes `parked-probe-failed` (`reason` = `capture-failed` / `match-timeout`, `pattern` for a timeout, `error`) once per expiry; a profile with no `parked_prompt_patterns` or a backend without `capture_pane` stays silent. A pane-log stat fault other than absence still leaves the #261/#727 proof-of-work signal unknown, and writes `log-evidence-failed` (`error`) once per session however many verdict sites consult it. A present `result.json` the read-back refuses (unreadable, unparseable, not an object) still reads as no result: the Stop read-back's give-up record in `resultless-stops.jsonl` says `malformed-result-json` with the refusal instead of `no-result-json`, once per give-up rather than per poll, and the exit read-back writes `result-json-refused` (`error`). The opencode HTTP adapter shares the `resultless-stops.jsonl` verdict but not the other crumbs.
- A nested sprint spec explains its own timeout (#780). The unpinned sprint-mode dev read-back reads only specs directly under the artifacts dir, never its subdirectories — widening the scan would widen the #261 shared-dir hazard — so a spec kept in, say, `implementation-artifacts/stories/` is never found and the session rides to timeout. The `no-artifact` crumb in `resultless-stops.jsonl` now names any qualifying spec one level down ("qualifying spec(s) found in subdirectories, which are never read back: … — move the spec directly under the artifacts dir, or use stories mode ([stories] source) for a stories/ layout"); a nested hit is named, never read back as a result, and a nested `*.md` the probe cannot read, or a subdirectory it cannot list, is reported as a probe fault rather than read as "nothing nested". A dev session that ends non-completed carries the last such verdict on its `session-end` journal entry as `resultless_verdict` / `resultless_detail` (present-only, diagnostic, never routing), so the reason surfaces without opening the task dir. `bmad-loop validate` warns `queue.nested-specs` in sprint mode on a `*.md` with a non-empty frontmatter `status:` one level down, before any tokens are spent (a nested `*.md` it cannot read, or a subdirectory it cannot list, is warned on too, never silently skipped); stories mode never warns, since `stories/` is its own layout.
- Four more generic-adapter folds are crumbed, verdicts unchanged (DW-452/455/456/457). A budget usage sample whose transcript read raises still reads as "no sample" — with a persistent fault, `token_budget_mode = "enforce"` stays off for the session — and its streak writes `usage-sample-failed` (`error`) at the first failure and `usage-sample-recovered` (`failures`) at the next clean sample, never once per tick (one torn mid-append read is one pair); `heartbeat.json` carries the running count as `usage_sample_failures`. A #727 transcript activity scan that raises still counts as no model-side evidence, with its streak crumbed the same way (`transcript-scan-failed` / `transcript-scan-recovered`). A launch-snapshot fault that turns the #276 M1 refuse gate inert — the identity check's `resolve()` raising (the 3.11 symlink-loop `RuntimeError` included, which used to escape) or the digest read raising — still reads NEUTRAL, and writes `spec-identity-unreadable` (`spec`, `snapshot`, `error`) or `spec-digest-unreadable` (`spec`, `error`) once per read-back, not per grace poll. A spec read-back that gives up on a fault says so: `stat-failed` for a stat fault the stories read-back used to file as `stale-mtime`, `unreadable-spec` for an unreadable or undecodable spec it filed as `not-terminal` (and the scan fallback as `no-artifact`), each with the error, in `resultless-stops.jsonl` on a Stop read-back and as `spec-readback-failed` (`reason`, `spec`, `error`) on a one-shot or dead-window (post-kill reconcile) read, which recorded nothing before.
- CRITICAL resolution: `bmad-loop resolve <run-id>` opens an interactive resolve agent seeded with the escalation + frozen spec; you disambiguate, it re-arms the story (`escalated → pending`, spec reset to `ready-for-dev`) and resumes. `--no-interactive` skips to re-arm if you fixed the spec yourself. A DEFERRED or environment-fault escalated story whose kept work is sound takes [`--reverify`](#resolve-reverify) instead, which replays verification rather than re-driving a dev session. The re-arm advances the story's
baseline in the **code tree** and is honest when it cannot: a failed advance is narrowed to typed git
Expand Down
4 changes: 3 additions & 1 deletion docs/tui-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -326,7 +326,9 @@ One row per story (or sweep bundle/triage task) in the selected run:
read fault, with the error),
`ambiguous-frontmatter`, `unmodified-since-launch` — the spec's bytes were
unchanged since review launch, so it is a prior `done` re-opened, not this
session's output (#276) — or `terminal-frontmatter-pending`).
session's output (#276) — or `terminal-frontmatter-pending`). A dev session
that ends non-completed also shows the last of these on its `session-end`
journal entry, as `resultless_verdict` / `resultless_detail` (#780).
- **Log** — the active agent session's pane output (`logs/<task-id>.log`),
ANSI colors preserved, starting with a dim `— <task-id>.log —` header. The
active task is the last `session-start` without a matching `session-end`
Expand Down
9 changes: 9 additions & 0 deletions src/bmad_loop/adapters/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -308,6 +308,15 @@ class SessionResult:
# into the same prompt. `parked_evidence` names what matched. APPENDED.
parked: bool = False
parked_evidence: str | None = None
# Why the dev read-back last came up empty (#780): the verdict and detail of
# the session's last `resultless-stops.jsonl` crumb, folded in by
# `_DevSynthesisMixin.run` when the session ends non-completed. That file has
# no reader, so a spec nested one directory down rode to timeout with no
# visible reason; the engine journals these on `session-end` instead.
# Diagnostic only: never read by routing, never set on a completed result.
# APPENDED.
resultless_verdict: str | None = None
resultless_detail: str | None = None


class CodingCLIAdapter(ABC):
Expand Down
Loading
Loading