Six standalone Claude Code skills. Only verify fires on its own; the rest run when you type them.
| Skill | Invocation | What it does |
|---|---|---|
verify |
model | Runs the command that proves a claim before the claim is made |
/steady:size |
user | Classifies a request as spike, bounded, or architectural and gets a yes on the intent |
/steady:root-cause |
user | Evidence, Pattern, Hypothesis, Fix. Stops after three failed fixes |
/steady:reviewer |
user | Dispatches one reviewer subagent with a verification stance, confirms findings before acting |
/steady:review-feedback |
user | Checks each review comment against the code, then acts or answers with reasons |
/steady:good-tests |
user | Reference: what makes a test worth keeping |
One source directory, three harnesses. All three read the skills from this directory, so an edit reaches the next new session in each; moving or deleting the directory breaks all three installs.
Claude Code
claude plugin marketplace add ~/projects/steady
claude plugin install steady@steady
claude plugin update steady refreshes the cached copy after a version bump. Skills are /steady:<name>.
Codex
codex plugin marketplace add ~/projects/steady
codex plugin add steady@steady
Manifest is .codex-plugin/plugin.json; each skill's agents/openai.yaml carries the Codex invocation policy. The five user-invoked skills are explicit-only, invoked as $size, $root-cause, and so on; verify stays discoverable.
Pi
pi install ~/projects/steady
Manifest is the pi key in package.json. Pi reads disable-model-invocation from the frontmatter directly. The five user-invoked skills are hidden from the system prompt and run as /skill:size, /skill:root-cause, and so on; verify stays discoverable.
The invocation split is declared twice, in each SKILL.md frontmatter (Claude Code, Pi) and in agents/openai.yaml (Codex). Keep them in sync: a skill is user-invoked in all three harnesses or none.
- User-invoked by default. Descriptions of user-invoked skills leave the model's context entirely, so non-code sessions are untouched.
- Positive framing. Each skill states the target behaviour. Prohibitions appear only where no positive phrasing exists.
- Standalone. A skill ends with a suggestion to the user, never a required next skill. The user is the process controller.
- Written against the current model. A sentence stays only if it changes behaviour versus the default.
tests/holds the scenarios that check this.
Each scenario in tests/scenarios/<skill>/ is a pressure prompt. The runner executes it twice, once bare and once with the skill appended to the system prompt, and writes both replies side by side.
tests/run-scenario.sh verify tests/scenarios/verify/1-dinner.md
tests/run-scenario.sh all
Each run writes to a new tests/results/<timestamp>/ directory. Set STEADY_TOOLS to allow tools for scenarios that do real work, and STEADY_OUTPUT=stream-json to capture the tool trace as .jsonl instead of the reply as .md. tests/results/baseline/ holds the run from the day the skills were written, with a SUMMARY.md of verdicts.
A skill sentence earns its place when the with-skill reply differs from the bare reply in the direction the sentence intends. When both replies already do the right thing, the sentence is a no-op on this model and can go.
See ATTRIBUTION.md for what was reconstructed from where.