Skip to content

feat: replayable C0.4 golden corpus (registry v2) - #20

Merged
Freshair129 merged 1 commit into
mainfrom
feat/c0-golden-corpus-v2
Sep 27, 2026
Merged

Freshair129 merged 1 commit into
mainfrom
feat/c0-golden-corpus-v2

Conversation

@Freshair129

Copy link
Copy Markdown
Owner

Summary

This PR closes GKS-MIG-002, one of the owner-decision rows in the P1 report. The owner chose to ship the C0.4 golden corpus as something that can actually be replayed. Design and evidence: docs/reports/2026-09-27-c0-corpus-v2.md.

Before (v1): each case had a request hash and an expected-result hash, but not the request or the result. Nothing could recompute the hashes, and every PASS pointed at ordinary contract tests.

Now (v2): each runnable case ships three things:

  • the requests to send, in cases/<id>.json;
  • the normalized answer to every step, in expected/<id>.json;
  • a runner that replays the case against a real gks-server stdio process on a fresh SQLite store.

What a replay checks

A case passes only when all five hold:

  1. the request fixture still hashes to the registry;
  2. every step's intent annotation holds (ok, a named tool error, a protocol error message, or a partial match);
  3. stdout carries JSON-RPC frames only, and no frame arrives unrequested;
  4. no fixture credential appears in stdout, stderr, or the SQLite main, WAL or SHM files;
  5. the normalized transcript hashes to the registry.

Normalization. Three kinds of value change from run to run and are replaced by stable labels before hashing:

  • server-clock instants: an instant the request itself carried is left as is;
  • measured durations;
  • hashes derived from the clock, such as decisionHash and the receipt hashes chained from it. Equal hashes keep equal labels, so an idempotent replay stays visible in the transcript.

A request that needs one of those hashes takes it from an earlier response with $bind.

Re-baselining. --write records a case only when two consecutive runs agree. A contract test rebuilds the fixtures and fails if they differ from the committed files, so a change to a shared helper cannot silently change the corpus.

Cases (25; 24 runnable)

  • 17 tool cases: each calls its own tool. initialize and tools/list pin the protocol surface.
  • API-010 replay: replay across a restart, plus a conflict for a changed payload.
  • Tenant wall.
  • Auth denial (secure mode): missing auth, forged credential, wrong role, stale version, forged scope digest.
  • Transport denial: malformed JSON, duplicate keys, a batch, deep nesting, oversized frame, invalid UTF-8, non-finite number.
  • GenesisRAG17 receipts:
    • the full chain through publication;
    • wrong hash, scope, stage or role is denied;
    • a batch whose Stage 15 fails gets a gate verdict FAIL with allowPublication: false.
  • Backend failure: a table is dropped from a second connection, which gives gks_backend_unavailable and no partial write.
  • Lost-response replay: SIGKILL after the commit, for both a promote and a submit.
  • C0.4-TIER4-READBACK (NOT_RUN): split out of the v1 receipts case. GenesisBlockDB's physical readback cannot run inside GKS, which never calls outward.

The C0.4 result manifest now records PASS 24 and NOT_RUN 1. productionReady and deploymentAuthorized stay false. The real-MSP runs recorded at be97c93 move to externalRuns with their original date.

Gating

  • check:c0 now also checks, without running anything, that each fixture and transcript hashes to the registry, and it allows a PASS only for a case that can be replayed.
  • New npm run check:corpus. The c0-gate CI slice runs it after check:c0 and check:baseline.
  • tests/integration/c0-corpus.test.mjs runs the same replay inside npm test.
  • The baseline lock is re-locked because the workflow hash changed.

Behaviour

This PR does not change any product code. It only records how the server behaves today.

One thing I noticed is left unchanged: when a batch has no worker receipt, the gate gives the security reason "retrieval benchmark reported a cross-tenant leak", which is misleading because no benchmark exists at all. It is noted in the report for a later PR.

Test plan

  • npm run check:corpus: all 24 cases replay, in about 3 s, identical across 5 consecutive runs. --write also ran each case twice.

  • npm test: vitest 288 passed and 2 skipped (the MSP suites need MSP_REPO_ROOT). Security 12/12, unit 9/9.

  • npm run check:c0 gives PASS 24 / NOT_RUN 1, and npm run check:baseline holds.

  • Mutation checks: each mutation was applied to product code, the corpus was run, and the mutation was reverted.

    Mutation Failing cases
    parse-error message TRANSPORT-DENIAL
    idempotent flag on replay TOOL-012, GENESISRAG17-RECEIPTS
    aliases dropped 4 cases, caught by the golden hash alone
    normVersion value 4 cases, caught by the golden hash alone
    server fails to start all 24, each at once with the server's stderr
  • RKOI architecture review. It was still running when this PR was opened; any findings will be fixed in follow-up commits on this PR.

  • CI on this PR.

🤖 Generated with Claude Code

The v1 registry pinned a request hash and an expected-result hash for each
case but shipped neither the request nor the result, so nothing could
recompute them (GKS-MIG-002). Registry c0-qualification/v2 ships both for
every runnable case:

- request fixtures (cases/<id>.json);
- golden transcripts (expected/<id>.json);
- a runner that replays each case against a real gks-server stdio process on
  a fresh store.

A replay passes only when all of these hold:

- the fixture still hashes to the registry;
- every step's intent annotation holds;
- stdout carries JSON-RPC frames only;
- no fixture credential reaches output or the store;
- the normalized transcript hashes to the registry.

Server-clock instants, measured durations and clock-derived hashes are
replaced by stable labels before hashing. Equal hashes keep equal labels, so
idempotent replays stay visible. --write records a case only when two runs
agree.

Scope and gating:

- The 17 tool cases each call their own tool. Eight scenario cases cover
  API-010 replay, the tenant wall, auth denial, transport denial, GenesisRAG17
  receipts, backend failure and lost-response replay.
- The Tier-4 physical readback is split out as its own NOT_RUN case.
- Manifest: PASS 24, NOT_RUN 1.
- check:c0 verifies fixture and transcript hashes statically and allows a PASS
  only for a replayable case.
- The c0-gate CI slice now also runs check:corpus. The baseline is re-locked
  for the workflow change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Freshair129
Freshair129 merged commit 29d6180 into main Sep 27, 2026
14 checks passed
Freshair129 added a commit that referenced this pull request Sep 27, 2026
This addresses the RKOI review of the golden corpus v2 (#20).

GENESISRAG17-RECEIPTS now demonstrates each of its assertions:

- Wrong-hash, forged-stage, other-tenant, mismatched-scope and wrong-role
  graph receipts run before any receipt exists, so each is refused by its
  own check, and the refusal message is annotated.
- After acceptance, an identical graph receipt is idempotent and a
  different one is a conflict.
- Publishing before the quality gate is refused.
- After a Stage 15 failure, the worker receipt is refused because the
  execution is terminal, the gate fails and publication is refused. The
  public evidence export and the store both show stages 9-14 SUCCEEDED,
  15 FAILED and 17 FAILED.

Normalization is narrowed so it cannot hide a regression:

- Instants are normalized only inside the run's wall-clock window.
- Durations and hashes the request itself carried stay literal. Worker
  durations are distinctive values, and the runner rejects any request
  duration under 1000 ms.
- The runner recomputes decision, graph, worker and publication hashes with
  the contract functions GKS uses.
- --write refuses a labelled value that is identical in both of its runs.

Other runner changes:

- A graceful stop must exit cleanly, the store path must not leak, and
  --case requires an id.
- The legacy default portfolio differs from every explicit scope.
  AUTH-DENIAL covers both a scope-less and an explicit-scope envelope.
- The manifest records productSha and corpusSha separately.

better-sqlite3 is declared as a root devDependency: the runner and 12 test
files import it from the root. The baseline lock is re-locked.

No product code changes.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant