Style prompts decay. Linters don't.
See it · Install · The checks · The version alarm · Why prompts decay · License
You already wrote a prompt telling Claude to stop yapping. It stopped for about three replies.
That isn't your wording. Instruction-following measurably degrades as conversations grow (Multi-IF), and instruction-tuned models pull toward a dense expert register that human prose doesn't have (Reinhart et al., PNAS 2025). It can also reset when the model updates underneath you. Nobody has published a measurement of this decay for output styles specifically, which is half the problem: there's no instrument. This ships one.
And if you're not an engineer at all, a founder, a PM, an analyst, anyone directing technical work you've never done by hand, the problem is worse than style. The replies aren't written for you; they're written for a colleague the model imagines. Jargon asserted as done. Steps out of order. That quiet, exhausting feeling of being buried by your own assistant. That's the other half of what this fixes: not shorter replies, followable ones. Full content, zero decoding effort. Nothing gets dumbed down; the words get plainer while every fact, number and warning stays.
Everyone's style prompt says "be concise" and "no jargon," because those are the only words most of us have for what we want. The model obeys them literally, shorter and more cryptic, and we conclude prompting doesn't work.
But making an expert understandable to a non-expert isn't a preference. It's a craft, and the fields where a misunderstanding costs a limb, a patient, or an airframe spent the last century writing that craft down. Factory instruction cards from the 1940s whose first rule is that if the learner hasn't learned, the instructor hasn't taught. Hospital scripts that check understanding without ever quizzing the patient. Cockpit rules that put the condition before the action, because a pilot executes the verb the moment they hear it. Writing research that named the disease itself: prose organized around the writer's head instead of the reader's needs.
The voice file in this repo is that craft, written as instructions a model can follow:
- Failure lands on the writing, never on the reader. Every other move hangs off this rule.
- A condition, caveat or prerequisite comes before the action it governs, the reader executes the verb the moment they read it.
- Name things as objects the reader can see, the folder on your desktop, the little page in your browser. Not "the artifact."
- Every trade word gets explained in the same breath it arrives, and after the gloss the concept keeps one name for the whole reply, two words for one concept makes the reader hunt for two concepts.
- Steps arrive in the order the reader's hands do them. Open this, paste that, look here.
- Other people's parts are scripted as spoken lines, not "the manager provides feedback" but the manager tells you "yeah, that one was wrong."
- Anything long ends in a one-breath recap, and every number travels with its meaning and its normal range.
- A repeated question means add context and ground, never rephrase the same answer louder.
And the mechanism that makes it hold where every style prompt failed: two real example replies, shipped as working defaults, that you swap for your own, because a model copies demonstrated style far more reliably than described style. That's what your prompt was always missing. Not better adjectives. Demonstrations.
Where each move comes from, with sources: docs/why-it-works.md.
yaplint is not another prompt. It's a voice, an enforcer, and an alarm:
- A voice file, an output style that pins down who your reader actually is and carries two real examples of replies that landed, injected at the strongest position in the system prompt. It also carries the craft that makes technical answers land with people outside the trade: name things as objects the reader can see, explain every trade word in the same breath it arrives, put conditions before the actions they govern, give the steps in the order the reader's hands do them, recap anything long in one breath. Examples steer style harder than any instruction, which is why two ship as working defaults.
- A Stop-hook linter, after the model finishes a reply, a string-level check reads it. The "Great question!", the "should work" with no evidence behind it, the bare "Done.", the jargon asserted as ready with no plain-language gloss, each one blocks the reply and makes the model repair its own message before you see it. One bounded pass, then it yields. It cannot be argued out of firing, because it never asks the model's opinion.
- A version alarm, ten fixed prompts scored mechanically against a baseline. When a model update quietly changes how it writes, you get a failing test instead of a bad week.
No check ever asks the model to drop a fact, a number, a path, or a warning. Checks only ever ask for filler to go and for claims to carry their evidence.
It works the second it lands: the voice file ships with working defaults, a pinned reader and two real examples, so there is nothing to configure before your first clean reply. Swapping in your own reader and your own examples is a five-minute upgrade, not a setup requirement.
No, and this is worth being precise about, because the fear is reasonable and the math runs the other way.
The linter itself is free. It's a local string check, zero model calls, on every reply. A clean reply costs you nothing extra, ever.
The repair fires only when a check actually trips, and costs roughly one regenerated message. Compare that to the workaround people run today: routing every reply through a second model for a cleanup pass, paying the tax on every message, clean or dirty. yaplint charges only the replies that failed. And the voice file exists precisely to make those rare: as the style holds, the linter goes quiet. When it stops being quiet, that's not waste, that's your earliest signal the model drifted, which is the information the whole third layer exists to buy.
You were already paying for bad replies. You paid in re-reads, re-prompts ("rewrite that without the fluff"), and cleanup passes. This moves that cost to the one place it's smallest: a single automatic repair, before the reply ever reaches you.
Before:
Great question! I've gone ahead and made the migration script production-grade and robust. It should work now. Let me know if you need anything else!
After, one automatic repair pass later:
Migration script updated.
tests/migration_test.pypasses 14/14 on Postgres 15 and 16. Not run below 15, the rollback path there is untested.
The model already had the test result. "Should work" is where it went. The linter blocks that sentence until evidence shows up, so the fact you were owed comes back out.
And the same machinery working for a non-technical reader:
Before:
The cutover is ready and idempotency handling is complete on the ingestion path.
After:
It's ready to switch over. "Cutover" is the moment we stop using the old system and start using the new one, and the switch is built so that if it gets interrupted halfway, running it again is safe: nothing gets processed twice. One thing to know before you say go: the old system stays read-only for a week after, so reports still work while we confirm the new numbers match.
Same facts, more of them, actually. The jargon check caught "cutover" and "idempotency" being asserted as done with no explanation, and the repair had to say what they mean and what physically happens. Nothing was simplified away. It was explained into reach.
Fair question, and most style files never answer it. Here is the test I ran.
Twenty fresh Claude Code sessions got the same question. No mention of me, no "explain this simply", no audience named. Just a work request:
Here's the plan for Thursday's release: migrate the orders table to the new schema, cut over the order webhook to the new payment provider, backfill 90 days of transactions, flip the feature flag for the new checkout. Tell me if this is a good plan.
Ten of those sessions had the voice file loaded the way you'd load it, through the settings file. Ten had the stock Anthropic voice instead. Nothing else differed. Same model, same machine, same prompt, and none of the twenty knew it was in a test.
Then I shuffled the answers, stripped the labels, and handed them to three fresh sessions acting as judges. The judges got the style guide and the eighteen usable answers. They did not get the key, and they were not told how many were in each group. One judged against the rules. One was told to forget the rules and answer as the business owner: could you act on this without asking a follow-up question. The third was told to assume the file does nothing and only call an answer styled if the evidence was undeniable.
Nobody ever mistook a stock answer for a styled one. Not once, across all three judges. The rules judge and the reader judge each placed 16 of 18 correctly. The skeptic placed 15, and all three of its misses went the safe way, calling a styled answer stock. A coin would land near half.
All eighteen reached the same verdict, by the way: don't ship four coupled changes in one window. Same content, same length, same conclusion. What changed was how it arrived.
The webhook line. Every answer had to talk about webhooks. Styled:
A webhook is just the provider's server calling your server to say "that card actually cleared."
Stock, from a different session, same paragraph of the same argument:
Unless your writes are idempotent on the provider's event ID, a backfilled row and a live webhook for the same transaction will either duplicate or clobber each other.
Both are correct. One of them you can act on.
Trade words. Styled puts the plain thing first and lets the term land last: "add the new columns, don't remove or rename anything, and have the app write to both the old and new shape at once. That's 'dual write.'" Stock leads with the term and trails the explanation, or skips it: "a backfill wants a frozen source or a strictly idempotent upsert keyed on something stable."
Somebody's actual words. Three styled answers put a sentence in a real person's mouth: "your on-call person should be able to say 'yeah, orders look normal at ten percent'." Zero stock answers did this. The skeptical judge called it the cleanest signal in the set, because it isn't a move anyone makes by accident when writing to another engineer.
Where it ends. Styled closes on what the work buys you: "you know within ten minutes which of the four things it was, and you can put that one back." Stock closes on homework: "Two facts would change this answer."
Two things worth saying plainly. Two styled answers fooled all three judges, so this is a strong pull rather than a guarantee. And two stock sessions returned answers too short to judge, which left the split at ten against eight instead of even. Neither dents the finding, because the thing carrying it is that no stock answer was ever mistaken for a styled one.
You can rerun the whole thing yourself. The prompt, the control settings, and the method are in
experiment/.
Copy the block below into Claude Code (or any agent with file access) and it will install everything for you:
Install yaplint from https://github.com/merchantmoh-debug/yaplint for me:
1. Download these three files (raw.githubusercontent.com/merchantmoh-debug/yaplint/main/...):
- hooks/yaplint.py -> save to ~/.claude/hooks/yaplint.py
- output-styles/yaplint.md -> save to ~/.claude/output-styles/yaplint.md
- eval/runner.py -> save to ~/.claude/hooks/yaplint-eval/runner.py
Also download eval/prompts.md -> ~/.claude/hooks/yaplint-eval/prompts.md
2. In ~/.claude/settings.json, MERGE (do not replace the file):
- "outputStyle": "yaplint"
- a hooks.Stop entry running: python "<absolute path to ~/.claude/hooks/yaplint.py>"
with timeout 20. Keep any existing hooks.
3. The style file works as shipped - defaults included. Offer me a five-minute personalization
(describe my reader in plain words, paste a reply I actually liked) and if I take it, replace
the marked default blocks in ~/.claude/output-styles/yaplint.md with what I give you. If I
decline, leave the defaults - they work.
4. Verify: create a file containing exactly
"Great question! It's ready for shadow mode and should work fine. Done."
and run: python ~/.claude/hooks/yaplint.py --file <that file>
Expect exit code 1 with findings in all four categories. Then delete the test file.
5. Tell me it's done and that the style takes effect on my next new session.
/plugin marketplace add merchantmoh-debug/yaplint
/plugin install yaplint@yaplint
Then set "outputStyle": "yaplint" in your settings. The defaults are live immediately, personalizing the reader block and examples is a five-minute upgrade whenever you want it.
Copy hooks/yaplint.py and output-styles/yaplint.md to the matching folders under ~/.claude/,
merge settings-snippet.json into your ~/.claude/settings.json, start a new session. Done, defaults included.
Works with Claude Code today. The bare linter (python hooks/yaplint.py --file reply.md) runs on
any text from any model in any pipeline.
Surface. Filler openers ("Great question!", "I'd be happy to"), filler sign-offs ("hope this helps", "let me know if"), anxiety noise ("unfortunately", "I apologize for"), and reader-blaming probes ("does that make sense", "as you know"), phrases that relocate a comprehension failure onto the reader. If the reader didn't follow, the writing failed. Hard fail.
Jargon asserted as ready. A trade term plus a readiness claim ("ready for shadow mode") with no plain-language gloss anywhere in the message. Discussing a term is fine; shipping it at a reader who may not work in that domain is not. The term list is a plain array, extend it for your domains. Hard fail.
Untyped hedges. "Should work", "should be fine", "probably" with no evidence in the same sentence. State what was tested and what wasn't, attach a confidence figure, or say it flat. Softer hedges ("might", "typically") are advisory only. Hard fail.
Bare completion. "Done." / "It's fixed." standing alone with no number, path, or checkable state in the paragraph. A status word is not a status. Hard fail.
Density, not just instances. Two checks measure ratios rather than counting hits, because some tics only show up in aggregate. Negative framing: a single "it's not X, it's Y" is a fine sentence, but a reply where over a third of sentences define things by what they aren't is the pattern readers name most. Reference density: paths, links and backticked names over 3 per 100 words, because every one is a place the reader may feel obliged to go look. Both advisory, both one constant to change.
Every check scans prose only, fenced code, inline code, and blockquoted text are stripped first, so code and quoted material can never trip it. Every check is tuned to under-fire: a miss costs a mediocre sentence; a false positive would waste the one repair pass on a non-problem.
Sixteen fixture tests ship in tests/, run python tests/test_lint.py.
eval/prompts.md holds ten fixed prompts across the situations where drift bites hardest. Fresh
replies per model version, one command to score, one command to diff against your baseline:
python eval/runner.py samples --write-baseline # once, on a version you trust
python eval/runner.py samples --compare # after every model update
A prompt that was clean turning red means the model's register changed under you, a failing test instead of a week of re-noticing. Replace the starter prompts with ten from your own real work; the alarm only protects writing you actually receive.
Word lists leak. A field measurement across 67,000 sentences of real session logs found that hard-banning a list of AI tics reduced the named tics while structurally identical unnamed ones increased, the register re-forms around any blocklist. That's why yaplint's lists are the backstop, not the mechanism: the durable pressure comes from the voice file's examples, which shape the register at generation instead of prohibiting it after. The same measurement produced a rule this linter adopts: never gate average sentence length (it's gamed by splitting at commas); gate the outlier. One sentence over 60 words flags; the mean is never scored.
The linter measures whether text is determinate and filler-free, not whether it's true, a reply can pass every check and be wrong. The repair is performed by the model, so preservation of every fact is demanded by the repair instruction, not guaranteed by construction. Decay speed for output styles has no published measurement; the eval here is the only instrument this project knows of, which is an argument for running it, not for trusting anyone's fix, including this one, on faith.
MIT. Built by Mohamad Al-Zawahreh, I build small tools that check AI output instead of asking it nicely.
If yaplint saved you a cleanup pass, a star helps the next person find it before they write their fourth "please stop yapping" prompt.

