Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ There is intentionally no `.claude/` self-install in this repo: the plugin is en
- **Dogfood the kit**: evolve dobby through its own stages (`/dobby:scope` → … → `/dobby:wrap`) or `/dobby:dispatch` for small fixes. Friction found while doing so is signal — fix the kit, not the workaround.
- **Everything in English.** Three skill categories coexist. (1) The work-session **stage** skills (`scope` → … → `wrap`), the **side-path** skills that plug into the flow on demand (`handoff`, `triage`, `map`, `resolve-conflicts`, `wizard`, `teach`, `upgrade`), the worker agents, and their supporting skills are **methodology** — project-agnostic (no references to any specific codebase) but assuming the single **terminal execution host** (with optional cmux enrichment when `CMUX_WORKSPACE_ID` is set) and reaching the `@kvnwolf/dobby` CLI (`bunx dobby` — `env`/`instructions`/`check`/`up`/`down`/`dev`/`db:*`/`update`) for the mechanics it still executes and the instruction catalogue for what it cannot; the kit reaches the running app via the devUrl `dobby env` reports plus a curl liveness check, and drives the UI per environment — `dobby instructions browser`, which under cmux is its own cmux-browser → claude-in-chrome → curl ladder and under Claude Desktop / t3 code is that host's own MCP browser tools; see the statement above). (2) The kit ALSO carries **convention** skills (`data-processing`, `data-fetching`, `module-conventions`) that encode the user's standard application stack (TanStack Start + Drizzle/Neon + Better Auth, the `@/shared` form/data system) and intentionally reference its module file conventions — deep-path imports and the role-based file taxonomy (`{export}.server.ts` / `functions.ts` / `{descriptor}.browser.ts` / `schema.gen.ts`), no barrels. That coupling is deliberate, not a leak to genericize. (3) **Kit self-improvement tooling** (`mark`, `learn`) couples to the *host* — Claude Code session storage (`~/.claude/projects`, `CLAUDE_CODE_SESSION_ID`) — not to any project, and exists to evolve the kit from how it behaved in real field sessions. That host-coupling is intentional and each such `SKILL.md` must label itself as this category so it isn't mistaken for project-agnostic methodology.
- **Skills carry NO `model:`/`effort:`; each agent's own frontmatter has one owner.** Skills inherit the interactive SESSION's model/effort; the maintainer chooses the main-thread Architect's intelligence manually for the task. Agent PROMPT BODIES in `plugin/agents/*.md` remain the authoritative role instructions, and their frontmatter is the SOLE source of that role's model/effort — there is no external recipe to mirror or keep in sync: researcher Sonnet/medium, test-author Opus/high, implementor Sonnet/high, reviewer Opus/high, qa Sonnet/medium. There is no normal execute reviewer loop; reviewer remains available for explicit dispatch and missing-work-log safety review. Claude Code's operator-level `CLAUDE_CODE_SUBAGENT_MODEL` may still override subagent pins externally; record that when evaluating a run because it is host control, not a Dobby setting.
- **Namespacing is mandatory**: cross-references between kit pieces are always `/dobby:<skill>` and `dobby:<agent>`. Bare names only for things outside the plugin. After any rename/addition, grep for bare references.
- **Namespacing is mandatory**: cross-references between kit pieces are always `/dobby:<skill>` and `dobby:<agent>`. Bare names only for things outside the plugin. The Agent tool's `name` is a separate per-task ADDRESS (`test-author-t<id>` / `implementor-t<id>` / `qa-t<id>` — a `:` is rejected there), never the namespaced id. After any rename/addition, grep for bare references.
- **The Architect never implements.** The interactive main thread owns planning, host mechanics, `AskUserQuestion`, persistence, routing, and worker dispatch. It delegates exploration and code changes to workers; a stage skill that makes main grep around or implement code is a regression.
- **Spec prefers scoped evidence but keeps the original standalone path.** In a work session, `/dobby:spec` writes only the existing `STATE.md`; for an already-understood standalone task with no state, it may run `dobby state init` before persisting and linting the plan. It never invents missing understanding—when inputs are thin, route to interview/research instead.
- **No interruptions mid-flow**: stages run to completion; gates exist only at stage handoffs (Next-step) and plan approval. No unsolicited explanations; teaching is opt-in via `/dobby:teach`.
Expand Down
6 changes: 3 additions & 3 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,9 +31,9 @@ The vocabulary of the dobby kit. Use these terms exactly — in skills, agents,
- **Design tree** — the dependency structure of a task's open decisions in `/dobby:interview`: every question hangs off the answers or research it depends on, and each answer pushes the **frontier** outward.
- **Frontier** — every open decision in the **design tree** whose prerequisites are already settled: everything `/dobby:interview` could ask RIGHT NOW without guessing — what a **round** asks, split across consecutive rounds by the four-question popup cap and by vehicle; the interview closes only when it is empty.
- **Round** — one turn of frontier questioning in `/dobby:interview`, homogeneous by vehicle: an `AskUserQuestion` batch of AT MOST four questions with anticipatable options, or a plain-text batch of open-ended ones.
- **Build loop** — the per-task loop that turns a spec into locally verified code: `test-author` (conditional — only when the repo has a test suite AND the spec marked the task test-first) → `implementor` (implements, then runs its own **Exit gate**) → **QA** (proves real behaviour and returns a verdict). A defect QA finds opens a **build conversation** with the implementor, or with the test-author when the failure traces to the test contract rather than the implementation; that conversation is capped at five rounds, after which QA reports the task `needs-human` instead of retrying further. Normal execute never dispatches the reviewer — holistic static review lives at the **External PR review boundary**. Encoded once in the `dobby:execute` skill's `references/build-protocol.md` — the dispatch protocol the Architect follows directly, for every task — and reused by `/dobby:dispatch` and `/dobby:address-review`, which follow that same protocol for a single ad-hoc task instead of a whole plan.
- **Build loop** — the per-task loop that turns a spec into locally verified code: `test-author` (conditional — only when the repo has a test suite AND the spec marked the task test-first) → `implementor` (implements, then runs its own **Exit gate**) → **QA** (proves real behaviour and returns a verdict). A defect QA finds opens a **build conversation** with the implementor, or with the test-author when the failure traces to the test contract rather than the implementation; that conversation is capped at five rounds, after which QA reports the task `needs-human` instead of retrying further. Each task's three workers are dispatched once and stay addressable under `<role>-t<id>` until the task reaches a terminal status, and the Architect prints a per-step status table (one row per task, one emoji per step) after every worker verdict or fix-round message. Normal execute never dispatches the reviewer — holistic static review lives at the **External PR review boundary**. Encoded once in the `dobby:execute` skill's `references/build-protocol.md` — the dispatch protocol the Architect follows directly, for every task — and reused by `/dobby:dispatch` and `/dobby:address-review`, which follow that same protocol for a single ad-hoc task instead of a whole plan.
- **QA** — the worker (`dobby:qa`) that runs the build loop's last step: proves a task's real BEHAVIOUR ONLY, never lint, types, build, or the test suite — that mechanical layer is already closed by the **Edit hook** and the implementor's own **Exit gate** before QA is ever dispatched. It drives the browser (following `dobby instructions browser`, the environment's own **instruction catalogue** entry for the UI-verification topic — never `env`, which reports no browser guide) where an app exists, or exercises the artefact directly (a CLI, a skill, a library) where none does, and returns `{pass, failureKind, evidence, verificationKind, findings}`. A genuine defect opens a **build conversation** with the implementor or test-author; QA never edits code itself. _Avoid_: verifier.
- **Build conversation** — the direct sibling `SendMessage` exchange that closes a build-loop failure: QA reaches back to whichever worker can fix it — `dobby:implementor` for a code defect, `dobby:test-author` (describing expected behaviour only, never a code fragment) for a test-contract problem — rather than starting a fresh agent that would re-read everything. QA numbers every message in the text itself (round 1, round 2, …), so the count survives a fresh context on either side; the conversation is capped at five rounds, after which QA stops and reports the task `needs-human`. An environment failure (a dead browser, an expired credential, anything QA can't attribute to the code or the tests) never enters this conversation — it goes straight to the Architect and costs no round.
- **Build conversation** — the direct sibling `SendMessage` exchange that closes a build-loop failure: QA reaches back to whichever worker can fix it — the task's `implementor-t<id>` for a code defect, `test-author-t<id>` (describing expected behaviour only, never a code fragment) for a test-contract problem (the `dobby:<role>` id is the `subagent_type`, never the address) — rather than starting a fresh agent that would re-read everything. QA numbers every message in the text itself (round 1, round 2, …), so the count survives a fresh context on either side; the conversation is capped at five rounds, after which QA stops and reports the task `needs-human`. An environment failure (a dead browser, an expired credential, anything QA can't attribute to the code or the tests) never enters this conversation — it goes straight to the Architect and costs no round; and the fixer messages the reporter back by name to close the round.
- **Test-author** — the worker (`dobby:test-author`) that runs the conditional first step of the build loop: it writes a task's tests **from the spec alone, never seeing the implementation**, producing the fixed contract the implementor must satisfy. Blindness to the code is what makes the tests anti-tautological — an independent source of truth. It runs once per task; a **build conversation** round re-implements against the same tests without the test-author touching the implementation itself, the implementor's own **Exit gate** runs them, and the external PR reviewer later judges their static quality in the context of the complete change.
- **Green baseline** — the record (`dobby baseline record`, stored at `.dobby/baseline.json`) of which test suites were already failing before a task's work began. The implementor's **Exit gate** consults it (`dobby check --fix --baseline`) so only suites newly red because of this change are reported — a pre-existing failure is exempt, not the implementor's to fix. It replaces hand-written KNOWN-RED exclusions, and an absent record is stated explicitly ("every failing suite counts") rather than passing silently. Distinct from the **gate cache**, which records only proven-GREEN verdicts, never pre-existing red ones.
- **Exit gate** — the full `dobby check --fix --baseline` the implementor runs on itself before handing a task to QA, serialised to one implementor at a time across the shared tree (a task's other steps — its test-author step, its implementation, its QA proof — keep running in parallel with its siblings; only the gate itself queues, because it judges the whole tree). This INVERTS the earlier rule that forbade implementors from running any check: the mechanical layer is closed before QA ever sees the task, which is what lets QA prove behaviour only and never re-run it.
Expand All @@ -44,7 +44,7 @@ The vocabulary of the dobby kit. Use these terms exactly — in skills, agents,
- **blocked** — a terminal task state alongside done/needs-human: the task was skipped WITHOUT spawning agents because a direct or transitive dependency ended `needs-human`; the result names the blocker (`blockedBy`) and the task gets no work-log entry. Transitivity is structural — a blocked task is itself non-`done`, so its own dependents block in turn — while every task independent of the dead one keeps running exactly as if nothing had happened. _Avoid_: skipped, cancelled.
- **Dispatch** — the lightweight ad-hoc path: a scoped task handed to one worker (or the single-task build loop), no `STATE.md`.
- **Prototype** — throwaway code that answers ONE design question, then dies. Two branches: **logic** (a minimal TUI over a pure, portable module) and **UI** (3-5 radically different variants on one route with a floating switcher).
- **Namespacing** — inside the plugin, every cross-reference is fully qualified: `/dobby:<skill>` for skills, `dobby:<agent>` for `subagent_type`/`agentType`. Bare names are reserved for things OUTSIDE the plugin (`deep-research`, `find-docs`, built-in `Plan`/`Explore`).
- **Namespacing** — inside the plugin, every cross-reference is fully qualified: `/dobby:<skill>` for skills, `dobby:<agent>` for `subagent_type`/`agentType`. Bare names are reserved for things OUTSIDE the plugin (`deep-research`, `find-docs`, built-in `Plan`/`Explore`). The Agent tool's `name` is a separate, per-task ADDRESS (`<role>-t<id>`; `:` is rejected there), so the namespaced id is never used as a name.
- **Session indicator** — the copy-pasteable pointer to a recorded Claude Code session that `/dobby:mark` emits (transcript `.jsonl` path + repo, worktree root, the `STATE.md` path, the `/dobby:*` skills it invoked, and a note). Consumed by `/dobby:learn`, which digests that session to improve a kit skill from how it actually behaved in the field. Together `mark` (capture, in the consumer project) and `learn` (digest + edit, in this repo) are the kit's **self-improvement loop**; they couple to host paths (`~/.claude/projects`, `CLAUDE_CODE_SESSION_ID`) on purpose.
- **Teach** — the kit's on-demand pedagogy skill (`/dobby:teach`): the interactive Architect teaches the user ONE topic in-session — explains it from trusted resources, runs a tight feedback loop to verify understanding, and records the demonstrated understanding as evidence. It is conversational interaction, not planning or code work, so no worker is dispatched. It is NOT a work-session stage. Distinct from the `mark`/`learn` self-improvement loop, which evolves the kit rather than the user.
- **Capability** — a detected fact about a project, derived from signals in its dependencies/marker files (e.g. `vite`, `tanstack-start`, `react`, `neon`, `drizzle`, `react-email`, `vitest`, `expo`). Capabilities DRIVE the mechanical layer — `dobby` infers each task from them (the check pipeline, the `dev` composition, the `db:*` set, the capability-filtered help); see **Task inference**. `dobby env` reports the detected list; a capability is triggered by one or more **detection signals**.
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -177,7 +177,7 @@ The coordinator makes sure the app is up — `/dobby:execute` runs `bunx dobby u

The implementor runs its own **Exit gate** (`bunx dobby check --fix --baseline`) before handing off — the full quality gate, serialised to one implementor at a time across the shared tree — so QA only ever proves real behaviour and never re-runs lint/types/build/the suite. The leading test step is conditional: when the repo has a test suite and the spec marked a task test-first, a `dobby:test-author` writes the failing tests before the implementor touches the code. There is no fixed batch grouping tasks into waves — independent tasks keep running the moment they're ready, and a destructive task (one that mutates shared backend state during its proof) runs alone, with nothing else touching that state at the same time. A genuine defect QA finds opens a **build conversation** with the implementor, or the test-author when the failure traces to the tests rather than the implementation; the conversation is capped at five rounds, after which the task reports `needs-human` instead of retrying further. Every task that depended on a dead one is skipped as `blocked` — no agents spawned, the blocker named in its row — while everything independent of it keeps running. The normal loop deliberately has no per-task reviewer: after commit/push, the repository's external reviewer (currently Greptile) reviews the complete PR, and merge readiness requires a review of the current HEAD rather than a stale summary or silence.

**You'll see:** the Architect narrating the run live as it works — a line when each task starts, one line per task the moment it lands (`✓ verified`, `✗ needs-human`, `⊘ blocked`), a line per build-conversation round, and extra detail for a task in trouble — while `STATE.md` stays current as the run advances, so a compaction or a fresh session can reconstruct exactly where things stand. Once every task in the plan has reached a terminal status, the run closes with one summary table: rounds per task, first-attempt success, what died and why, and wall clock.
**You'll see:** the Architect narrating the run live as it works — a line when each task starts, one line per task the moment it lands (`✓ verified`, `✗ needs-human`, `⊘ blocked`), a line per build-conversation round, and extra detail for a task in trouble — while `STATE.md` stays current as the run advances, so a compaction or a fresh session can reconstruct exactly where things stand. After every worker verdict or fix-round message, the Architect also prints a per-task status table — one row per task, one emoji per step for not-started/in-progress/failed/passed — so you always have a live view of the whole plan. Each task's workers are addressed as `test-author-t<id>` / `implementor-t<id>` / `qa-t<id>` and stay addressable under those same names across fix rounds. Once every task in the plan has reached a terminal status, the run closes with one summary table: rounds per task, first-attempt success, what died and why, and wall clock.

### 6. Wrap

Expand Down
Loading
Loading