Repository navigation
feat: replayable C0.4 golden corpus (registry v2) - #20
Merged
Merged
Conversation
The v1 registry pinned a request hash and an expected-result hash for each case but shipped neither the request nor the result, so nothing could recompute them (GKS-MIG-002). Registry c0-qualification/v2 ships both for every runnable case: - request fixtures (cases/<id>.json); - golden transcripts (expected/<id>.json); - a runner that replays each case against a real gks-server stdio process on a fresh store. A replay passes only when all of these hold: - the fixture still hashes to the registry; - every step's intent annotation holds; - stdout carries JSON-RPC frames only; - no fixture credential reaches output or the store; - the normalized transcript hashes to the registry. Server-clock instants, measured durations and clock-derived hashes are replaced by stable labels before hashing. Equal hashes keep equal labels, so idempotent replays stay visible. --write records a case only when two runs agree. Scope and gating: - The 17 tool cases each call their own tool. Eight scenario cases cover API-010 replay, the tenant wall, auth denial, transport denial, GenesisRAG17 receipts, backend failure and lost-response replay. - The Tier-4 physical readback is split out as its own NOT_RUN case. - Manifest: PASS 24, NOT_RUN 1. - check:c0 verifies fixture and transcript hashes statically and allows a PASS only for a replayable case. - The c0-gate CI slice now also runs check:corpus. The baseline is re-locked for the workflow change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 27, 2026
Freshair129
added a commit
that referenced
this pull request
Sep 27, 2026
This addresses the RKOI review of the golden corpus v2 (#20). GENESISRAG17-RECEIPTS now demonstrates each of its assertions: - Wrong-hash, forged-stage, other-tenant, mismatched-scope and wrong-role graph receipts run before any receipt exists, so each is refused by its own check, and the refusal message is annotated. - After acceptance, an identical graph receipt is idempotent and a different one is a conflict. - Publishing before the quality gate is refused. - After a Stage 15 failure, the worker receipt is refused because the execution is terminal, the gate fails and publication is refused. The public evidence export and the store both show stages 9-14 SUCCEEDED, 15 FAILED and 17 FAILED. Normalization is narrowed so it cannot hide a regression: - Instants are normalized only inside the run's wall-clock window. - Durations and hashes the request itself carried stay literal. Worker durations are distinctive values, and the runner rejects any request duration under 1000 ms. - The runner recomputes decision, graph, worker and publication hashes with the contract functions GKS uses. - --write refuses a labelled value that is identical in both of its runs. Other runner changes: - A graceful stop must exit cleanly, the store path must not leak, and --case requires an id. - The legacy default portfolio differs from every explicit scope. AUTH-DENIAL covers both a scope-less and an explicit-scope envelope. - The manifest records productSha and corpusSha separately. better-sqlite3 is declared as a root devDependency: the runner and 12 test files import it from the root. The baseline lock is re-locked. No product code changes. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR closes GKS-MIG-002, one of the owner-decision rows in the P1 report. The owner chose to ship the C0.4 golden corpus as something that can actually be replayed. Design and evidence:
docs/reports/2026-09-27-c0-corpus-v2.md.Before (v1): each case had a request hash and an expected-result hash, but not the request or the result. Nothing could recompute the hashes, and every PASS pointed at ordinary contract tests.
Now (v2): each runnable case ships three things:
cases/<id>.json;expected/<id>.json;gks-serverstdio process on a fresh SQLite store.What a replay checks
A case passes only when all five hold:
ok, a named tool error, a protocol error message, or a partial match);Normalization. Three kinds of value change from run to run and are replaced by stable labels before hashing:
decisionHashand the receipt hashes chained from it. Equal hashes keep equal labels, so an idempotent replay stays visible in the transcript.A request that needs one of those hashes takes it from an earlier response with
$bind.Re-baselining.
--writerecords a case only when two consecutive runs agree. A contract test rebuilds the fixtures and fails if they differ from the committed files, so a change to a shared helper cannot silently change the corpus.Cases (25; 24 runnable)
initializeandtools/listpin the protocol surface.FAILwithallowPublication: false.gks_backend_unavailableand no partial write.The C0.4 result manifest now records PASS 24 and NOT_RUN 1.
productionReadyanddeploymentAuthorizedstayfalse. The real-MSP runs recorded atbe97c93move toexternalRunswith their original date.Gating
check:c0now also checks, without running anything, that each fixture and transcript hashes to the registry, and it allows a PASS only for a case that can be replayed.npm run check:corpus. Thec0-gateCI slice runs it aftercheck:c0andcheck:baseline.tests/integration/c0-corpus.test.mjsruns the same replay insidenpm test.Behaviour
This PR does not change any product code. It only records how the server behaves today.
One thing I noticed is left unchanged: when a batch has no worker receipt, the gate gives the security reason "retrieval benchmark reported a cross-tenant leak", which is misleading because no benchmark exists at all. It is noted in the report for a later PR.
Test plan
npm run check:corpus: all 24 cases replay, in about 3 s, identical across 5 consecutive runs.--writealso ran each case twice.npm test: vitest 288 passed and 2 skipped (the MSP suites needMSP_REPO_ROOT). Security 12/12, unit 9/9.npm run check:c0gives PASS 24 / NOT_RUN 1, andnpm run check:baselineholds.Mutation checks: each mutation was applied to product code, the corpus was run, and the mutation was reverted.
idempotentflag on replayaliasesdroppednormVersionvalueRKOI architecture review. It was still running when this PR was opened; any findings will be fixed in follow-up commits on this PR.
CI on this PR.
🤖 Generated with Claude Code