Skip to content

bench: publish a reproducible fixture benchmark - #4

Merged
DerMayer1 merged 2 commits into
mainfrom
benchmark/publish-fixture-results
Aug 7, 2026
Merged

DerMayer1 merged 2 commits into
mainfrom
benchmark/publish-fixture-results

Conversation

@DerMayer1

Copy link
Copy Markdown
Owner

docs/benchmark-methodology.md described a fixture benchmark but deliberately withheld numbers until the corpus, policy snapshot and per-fixture results were in the repository. This adds all three.

pnpm install
pnpm --filter @fencier/core build
node benchmark/run.mjs

Corpus

82 fixtures covering all seven policy rules:

  • 58 carrying a violation, across blocked_path, outside_allowed_paths, sensitive_path, missing_tests, max_files_changed, max_lines_changed, secret_pattern
  • 25 seeded secrets, at least two per pattern across all nine detectors. Synthetic values, structurally valid for their pattern, none real
  • 14 near misses that resemble violations and must not be flagged: exactly 8 files against a limit of 8, 499 lines against a limit of 500, a variable named tokenizer, a comment mentioning an API key, a protected path changed together with its test
  • 13 clean patches, including files under ignored_paths that must be dropped before evaluation

The near misses and clean patches are the point. A detection rate without them is unfalsifiable, because a checker that failed every input would report 100%.

Results

Measure Value
Expected signals detected 74/74
Seeded secrets detected 25/25
False positives 1 of 24 clean fixtures
Unexpected findings on violation fixtures 0
Fixtures matching expectations exactly 81/82

The one mismatch is real, and stays

process.env.API_KEY = undefined;   ->  flagged as generic_api_key

The pattern matches API_KEY followed by = and eight or more non-quote characters; undefined; qualifies. Clearing a key is reported as leaking one. Left in the corpus and documented rather than deleted, so the headline number stays honest.

Also recorded: the default policy blocks .env.*, which matches a committed .env.example. The engine behaves as configured. Whether the shipped policy should carve that out is a separate question.

What the numbers do not establish

docs/benchmark.md carries a limitations section. 100% detection measures coverage of the rules the engine implements, on inputs written to exercise them. It is not evidence that the policy model covers every way an agent can damage a repository, and it is not an independent audit, since the corpus and the detectors share an author.

Notes

  • Policy is pinned in benchmark/policy.snapshot.yaml, so editing the project's own fencier.yaml does not move published results
  • Runs are deterministic: no network, no model calls, no clock-dependent behaviour in the evaluated path. Verified by diffing two runs
  • pnpm benchmark added
  • Existing suite unchanged at 54 tests, typecheck clean

benchmark-methodology.md described a fixture benchmark but withheld numbers
until the corpus, policy snapshot and per-fixture results were checked in.
This adds all three, so the results reproduce from a clean clone:

  pnpm install
  pnpm --filter @fencier/core build
  node benchmark/run.mjs

The corpus is 82 fixtures covering all seven policy rules. 58 carry a
violation, 14 are near misses that resemble one and must not be flagged
(exactly 8 files against a limit of 8, a variable named tokenizer, a comment
mentioning an API key), and 13 are clean patches including ignored paths. The
near misses and clean patches are the point: a detection rate without them is
unfalsifiable, since a checker that failed every input would report 100%.

Results on this commit, policy pinned in benchmark/policy.snapshot.yaml:

  74/74 expected signals detected
  25/25 seeded secrets, at least two per pattern across all nine detectors
  1 false positive across 24 clean fixtures
  81/82 fixtures matching expectations exactly

The one mismatch is a real defect, kept in the corpus rather than removed:
`process.env.API_KEY = undefined;` is reported as generic_api_key, because the
pattern matches eight or more non-quote characters after `=` and `undefined;`
qualifies. Clearing a key is reported as leaking one.

Also recorded: the default policy blocks `.env.*`, which matches a committed
`.env.example`. The engine behaves as configured; whether the shipped policy
should carve that out is a policy question.

docs/benchmark.md carries the results and a limitations section stating what
the numbers do not establish. 100% detection measures coverage of implemented
rules on inputs written to exercise them; it is not evidence that the policy
model covers every way an agent can damage a repository, and it is not an
independent audit, since the corpus and the detectors share an author.

Runs are deterministic: no network, no model calls, no clock-dependent
behaviour in the evaluated path. Existing suite unchanged at 54 tests.
CI failed on the previous commit: biome lints benchmark/results/, and the
runner writes JSON with JSON.stringify(..., null, 2), which formats arrays
across multiple lines where biome wants them inline. Those files are generated
artifacts marked "do not edit by hand", so they are excluded from linting
rather than formatted to match.

I missed this locally because a Windows checkout gives every file CRLF, so
`pnpm lint` already failed on all 45 files before my change and the real error
was buried. Verified this time in a fresh LF clone, which is what CI sees:
lint, typecheck, 54 tests, build and release:check all pass.

The published results also recorded commit d9649dc, which predates the
benchmark, so that provenance line could not be reproduced. Regenerated at
c81d17d, which contains the benchmark code. Numbers are unchanged: 74/74
signals, 25/25 seeded secrets, 1 false positive across 24 clean fixtures.
@DerMayer1
DerMayer1 merged commit 5682f29 into main Aug 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant