Add spec-review plugin: adversarial panel between brainstorming and writing-plans - #124
Merged
Merged
Conversation
added 2 commits
September 7, 2026 08:53
…riting-plans Superpowers reviews its own specs. Both brainstorming's spec self-review and writing-plans' self-review are same-agent, same-context passes that end "fix inline and move on", so the only independent critic of a spec in that flow is the human. This adds the missing pass. A panel of independent reviewers judges the committed spec under one lens each — subtraction, completeness, framing — then the lead verifies every finding against the spec and the repo, applies the unambiguous ones, and escalates only genuine trade-offs, batched into a single choice with a recommendation. Escalating trivia is treated as a failure of the adjudication step. Lenses are ordered by marginal value rather than by blast radius: subtraction first, because brainstorming tells itself to apply YAGNI ruthlessly but only ever self-checks it, and framing last, because the brainstorming dialogue has usually already interrogated it. The reviewer model is chosen at dispatch, never pinned in agent frontmatter: an unresolvable alias silently inherits the parent model, which produces a review that only looks independent. The rule is "not weaker than the authoring session, superior where available" — a reviewer below the author's capability rubber-stamps. For the same reason the receipt states only what is verifiable and never claims model independence.
Every plugin directory must ship a README.md — enforced by tests/lib/test_cross_references.sh::test_every_plugin_has_readme, which was the sole failure on the previous commit.
2 of 4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
When we design something with the superpowers workflow, the written spec is checked only by the
same assistant that wrote it, and then by a human. There is no independent second opinion at the
point where mistakes are cheapest to fix — before any code exists. This adds a small plugin that
fills that gap: a panel of independent reviewers reads the finished spec, and their findings are
filtered before anything is changed.
It sits between the two existing stages: brainstorming produces the spec, this reviews it, then
writing-plans turns it into an implementation plan.
Why this rather than an existing tool
code-review-suitereviews a diff after implementation. Its vocabulary is code-shaped —includes/severity-definitions.mdis about SQL injection and unclosed resources, andincludes/verdict-rubric.mdis explicitly "PR mode only" with a GitHub posting policy. Neithertransfers to a spec, where the real defects are "solves the wrong problem", "readable two ways",
and "over-scoped". So this ships its own small spec-shaped severity vocabulary instead of
importing an ill-fitting one, and lives in its own plugin rather than accreting onto the PR
machinery.
Changes
plugins/spec-review/skills/review-spec/SKILL.md— the workflow. Resolves the spec, dispatchesthe panel, verifies findings, classifies them, batches choices, applies, commits, hands over.
plugins/spec-review/agents/spec-reviewer.md— the reviewer subagent. Read-only tools, oneassigned lens, spec-shaped severity tiers (
BLOCKING/IMPORTANT/SUGGESTION), strictoutput shape with a mandatory
SUBTRACTIONSsection, capped at 60 lines.plugins/spec-review/.claude-plugin/plugin.json— manifest..claude-plugin/marketplace.json— register the plugin.Design decisions worth reviewing
scaled by blast radius. Reviewer count should follow how many genuinely orthogonal questions you
can ask, and a spec can be wrong in exactly three non-overlapping ways: wrong problem,
incomplete answer, over-built answer.
framingis the documented one to drop for a two-lenspanel, since the brainstorming dialogue usually already interrogated it.
model:in the agent frontmatter, deliberately. An alias that fails to resolve silentlyinherits the parent model, so a pinned alias can yield a review that only looks independent. The
caller picks the model at dispatch under the rule "not weaker than the authoring session,
superior where available".
but never asserts model independence, because the skill cannot confirm which model actually ran.
user-visible behaviour, or reviewers disagree and neither spec nor repo settles it, or it trades
off a stated preference. Four choices is the cap; more means the spec needs rework rather than a
four-round interrogation.
description— explicit mention and thepost-brainstorming seam. Deliberately avoids the every-prompt reminder pattern, which costs
context in every session.
Test plan
jq -eparsesmarketplace.jsonandplugin.jsonmodeltest_conventions.shenforces)tests/run.shgreen in CI/review-specagainst a real spec — not yet done; the skill hasnever been executed, only structurally validated
framingis droppedUnrelated issues noticed, not addressed here
tests/run.shcannot run unless the working directory is inside the repo:tests/lib/test_specialist_score.sh:5callsgit rev-parse --show-toplevelat source time underset -euo pipefail, aborting the whole harness. Other tests resolveREPO_ROOTfromBASH_SOURCE, so this one looks like an oversight.plugins/s3-search/exists on disk but is not registered in.claude-plugin/marketplace.json.