Repository navigation
bench: publish a reproducible fixture benchmark - #4
Merged
Merged
Conversation
benchmark-methodology.md described a fixture benchmark but withheld numbers until the corpus, policy snapshot and per-fixture results were checked in. This adds all three, so the results reproduce from a clean clone: pnpm install pnpm --filter @fencier/core build node benchmark/run.mjs The corpus is 82 fixtures covering all seven policy rules. 58 carry a violation, 14 are near misses that resemble one and must not be flagged (exactly 8 files against a limit of 8, a variable named tokenizer, a comment mentioning an API key), and 13 are clean patches including ignored paths. The near misses and clean patches are the point: a detection rate without them is unfalsifiable, since a checker that failed every input would report 100%. Results on this commit, policy pinned in benchmark/policy.snapshot.yaml: 74/74 expected signals detected 25/25 seeded secrets, at least two per pattern across all nine detectors 1 false positive across 24 clean fixtures 81/82 fixtures matching expectations exactly The one mismatch is a real defect, kept in the corpus rather than removed: `process.env.API_KEY = undefined;` is reported as generic_api_key, because the pattern matches eight or more non-quote characters after `=` and `undefined;` qualifies. Clearing a key is reported as leaking one. Also recorded: the default policy blocks `.env.*`, which matches a committed `.env.example`. The engine behaves as configured; whether the shipped policy should carve that out is a policy question. docs/benchmark.md carries the results and a limitations section stating what the numbers do not establish. 100% detection measures coverage of implemented rules on inputs written to exercise them; it is not evidence that the policy model covers every way an agent can damage a repository, and it is not an independent audit, since the corpus and the detectors share an author. Runs are deterministic: no network, no model calls, no clock-dependent behaviour in the evaluated path. Existing suite unchanged at 54 tests.
CI failed on the previous commit: biome lints benchmark/results/, and the runner writes JSON with JSON.stringify(..., null, 2), which formats arrays across multiple lines where biome wants them inline. Those files are generated artifacts marked "do not edit by hand", so they are excluded from linting rather than formatted to match. I missed this locally because a Windows checkout gives every file CRLF, so `pnpm lint` already failed on all 45 files before my change and the real error was buried. Verified this time in a fresh LF clone, which is what CI sees: lint, typecheck, 54 tests, build and release:check all pass. The published results also recorded commit d9649dc, which predates the benchmark, so that provenance line could not be reproduced. Regenerated at c81d17d, which contains the benchmark code. Numbers are unchanged: 74/74 signals, 25/25 seeded secrets, 1 false positive across 24 clean fixtures.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
docs/benchmark-methodology.mddescribed a fixture benchmark but deliberately withheld numbers until the corpus, policy snapshot and per-fixture results were in the repository. This adds all three.Corpus
82 fixtures covering all seven policy rules:
blocked_path,outside_allowed_paths,sensitive_path,missing_tests,max_files_changed,max_lines_changed,secret_patterntokenizer, a comment mentioning an API key, a protected path changed together with its testignored_pathsthat must be dropped before evaluationThe near misses and clean patches are the point. A detection rate without them is unfalsifiable, because a checker that failed every input would report 100%.
Results
The one mismatch is real, and stays
The pattern matches
API_KEYfollowed by=and eight or more non-quote characters;undefined;qualifies. Clearing a key is reported as leaking one. Left in the corpus and documented rather than deleted, so the headline number stays honest.Also recorded: the default policy blocks
.env.*, which matches a committed.env.example. The engine behaves as configured. Whether the shipped policy should carve that out is a separate question.What the numbers do not establish
docs/benchmark.mdcarries a limitations section. 100% detection measures coverage of the rules the engine implements, on inputs written to exercise them. It is not evidence that the policy model covers every way an agent can damage a repository, and it is not an independent audit, since the corpus and the detectors share an author.Notes
benchmark/policy.snapshot.yaml, so editing the project's ownfencier.yamldoes not move published resultspnpm benchmarkadded