Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ jobs:
security) npm run test:security ;;
unit) npx vitest run tests/unit ;;
client-pack) npm run pack:client ;;
c0-gate) npm run check:c0 && npm run check:baseline ;;
c0-gate) npm run check:c0 && npm run check:baseline && npm run check:corpus ;;
*) echo "Unknown validation slice: ${{ matrix.slice }}" >&2; exit 1 ;;
esac

Expand Down
20 changes: 18 additions & 2 deletions docs/ADR-GKS-C0-QUALIFICATION.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
version: "0.2.1"
version: "0.3.0"
created_at: "2026-09-22T00:00:00+07:00,RWANG,working-tree"
last_update: "2026-09-27T10:00:00+07:00,Claude"
last_update: "2026-09-27T12:00:00+07:00,Claude"
status: "beta"
approval_owner: "Boss (บอส)"
approval_recorded_at: "2026-09-22T00:00:00+07:00"
Expand Down Expand Up @@ -169,6 +169,21 @@ No production secrets, user data, or external live endpoint is allowed in the
corpus. A qualification run must report `PASS`, `FAIL`, `NOT_RUN`, or
`BLOCKED`; a missing external MSP fixture cannot become a green C0 claim.

**Registry v2 (2026-09-27, owner decision on GKS-MIG-002).** The hashes must
be reproducible, not only recorded. Registry `c0-qualification/v2` therefore
ships:

- the request fixture for every runnable case, under `cases/`;
- its normalized golden transcript, under `expected/`;
- a runner (`npm run check:corpus`) that replays each case against a real
`gks-server` process on a fresh store.

A status of `PASS` is only allowed for a case that has a replayable fixture. The
Tier-4 physical readback has no replayable fixture inside GKS, so it is its
own `NOT_RUN` case. Values the server derives from its clock are replaced by
stable labels before hashing; the rules are listed in
`docs/reports/2026-09-27-c0-corpus-v2.md`.

## Contract and test matrix

| Surface | Required assertion | Test class |
Expand Down Expand Up @@ -217,3 +232,4 @@ review confirms:
| 0.1.0 | 2026-09-22 | beta | User approved D1-D4 for C0 implementation | working-tree | RWANG |
| 0.2.0 | 2026-09-22 | beta | Clarified compatibility versus secure MSP auth mode and recorded real-chain qualification boundary | working-tree | RWANG |
| 0.2.1 | 2026-09-27 | beta | Cross-reference only: ADR-GKS-PIPELINE-VISIBILITY (accepted 2026-09-27) changes what `gks_search`, `gks_entity_get`, `gks_relations_get` and `gks_artifact_link` return for unpublished GenesisRAG17 entities. No request is newly rejected; see that ADR's observable-changes table. | working-tree | Claude |
| 0.3.0 | 2026-09-27 | beta | D4: registry v2. The owner decided to ship replayable request fixtures, golden transcripts and a corpus runner (GKS-MIG-002). A PASS now requires a replayable fixture. The Tier-4 physical readback is split out as its own NOT_RUN case. | working-tree | Claude |
179 changes: 179 additions & 0 deletions docs/reports/2026-09-27-c0-corpus-v2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,179 @@
---
version: "0.1.0b"
created_at: "2026-09-27T12:00:00+07:00,Claude,working-tree"
last_update: "2026-09-27T12:00:00+07:00,Claude"
status: "candidate"
superseded_by: null
attributes:
domain: "genesis-knowledge-system"
doc_type: "baseline-evidence"
scope: "C0.4 golden corpus registry v2: replayable request fixtures, golden transcripts and a corpus runner"
baseline_sha: "f50116abdc3fe606851df4d21697027e252341c7"
---

# C0.4 golden corpus v2

## Status

GKS-MIG-002 was the last P1 row that needed an owner decision about the corpus
([P1 report](2026-09-27-p1-c0-closure.md)). The owner chose to ship it. The
C0.4 corpus is now a replayable artifact rather than a list of hashes:

- registry `c0-qualification/v2` replaces v1;
- the result manifest is `PASS_WITH_LIMITATIONS`: PASS 24, NOT_RUN 1;
- `productionReady` and `deploymentAuthorized` stay `false`.

## What was wrong with v1

The v1 registry held a request hash and an expected-result hash for each case,
but not the request or the result. Nothing could recompute those hashes, so
they were labels, not evidence. Each PASS in the manifest pointed at ordinary
contract tests. Two further problems:

- the registry's own `status` field still said `NOT_RUN` for the 17 tool
cases while the manifest said `PASS`;
- the GenesisRAG17 receipts case was `NOT_RUN` as a whole, although only its
Tier-4 physical readback was out of reach.

## What v2 is

| Part | Path | Role |
|---|---|---|
| Builder | `scripts/c0-corpus/build-cases.mjs` | Defines every case, both its metadata and its steps. It writes the request fixtures and the registry deterministically. |
| Request fixtures | `tests/fixtures/c0-qualification/cases/<id>.json` | The exact frames to send, in order, with server environment and fixture credentials |
| Golden transcripts | `tests/fixtures/c0-qualification/expected/<id>.json` | The normalized answer to every step |
| Runner | `scripts/c0-corpus/runner.mjs`, `scripts/run-c0-corpus.mjs` | Replays each case against a real `gks-server` stdio process on a fresh SQLite store |
| Registry | `tests/fixtures/c0-qualification/registry.json` | Case metadata plus the SHA-256 of each canonical fixture and transcript |

A case is a list of steps:

- `call` sends a JSON-RPC frame.
- `raw` sends raw bytes, for transport denials.
- `restart` restarts the server on the same store.
- `kill-after-commit` sends a frame, waits until a read-only SQL probe sees
the durable write, then SIGKILLs the process without reading the reply.
- `store-exec` runs SQL from a second connection, to inject a fault.
- `store-query` reads rows and records them.

### What a replay checks

A replay passes only if all of the following hold:

1. the request fixture still hashes to the registry;
2. every step's intent annotation holds, for example `ok`, a named tool error
such as `gks_conflict`, a protocol error message, or a partial
`structuredContent` match;
3. every stdout line is a JSON-RPC 2.0 frame and no frame arrives unrequested;
4. no fixture credential appears in stdout, in stderr, or in the SQLite main,
WAL or SHM files;
5. the normalized transcript hashes to the registry.

The annotations exist so that `--write` cannot quietly re-baseline an answer
that contradicts the purpose of the case. The hash exists so that a change no
annotation mentions still fails.

### Normalization

Some values legitimately differ from run to run. The runner replaces three
kinds before hashing:

- **Server-clock instants** become `<server-time>`. An instant that the request
itself carried is left alone, so echoed fixture times still count.
- **Measured durations** (`duration_ms`, `processing_time_ms`) become
`<duration-ms>`.
- **Clock-derived hashes** (`decisionHash`, and the graph, derived, receipt,
verdict, publication and failure hashes chained from it) become numbered
labels such as `<decisionHash:1>`. Equal values keep equal labels, so
"the replay returned the committed decision" stays visible in the
transcript.

A decision's facts carry the transaction time, so its hash depends on the
server clock. A request that needs such a value binds it from an earlier
response with `{"$bind": "<step>.<path>"}`; every other request value is
literal.

### Guarding against unnoticed changes

`--write` records a case only if two consecutive runs produce identical
transcripts, so a volatile field that normalization misses cannot become the
golden result. A contract test rebuilds the fixtures and fails if they differ
from the committed files, so a change to a shared test helper (`makeBatch`,
`promotion`) cannot silently change the corpus.

## Cases

The registry has 25 cases, of which 24 are runnable.

| Case | Replays |
|---|---|
| C0.4-TOOL-001 to 017 | Each of the 17 public tools, preceded by the setup it needs. `initialize` and `tools/list` pin the protocol surface and the registry. The contract test requires each tool case to call its own tool. |
| API010-REPLAY | promote, restart, identical replay (idempotent), `gks_conflict` for a changed payload; one promotion row and one evidence row |
| TENANT-WALL | Tenant and tenantless rows never cross, `portfolio-shared` does not bypass the wall, and cross-tenant reads return `gks_scope_denied` |
| AUTH-DENIAL | Secure mode. Missing auth, a forged credential, the wrong role, a stale version and a forged scope digest are all denied and write nothing. A valid envelope is accepted, and the MSP credential never reaches output or the store. |
| TRANSPORT-DENIAL | Malformed JSON, duplicate keys, a batch, 33-deep nesting, an oversized frame, invalid UTF-8 and a non-finite number are each rejected. The server keeps serving afterwards. |
| GENESISRAG17-RECEIPTS | A full chain through publication. A duplicate graph receipt is idempotent. A wrong decision hash, wrong scope, forged stage identity and wrong role are each denied. A second batch whose Stage 15 fails gets a gate verdict `FAIL` with `allowPublication: false`. |
| TIER4-READBACK | **NOT_RUN.** GenesisBlockDB's physical readback cannot run inside GKS, which never calls outward. This case was split out of the v1 receipts case. |
| BACKEND-FAILURE | The last table a promotion writes is dropped from a second connection. The promotion returns `gks_backend_unavailable`, and no entity or promotion row is left behind. |
| LOST-RESPONSE-REPLAY | SIGKILL after the durable commit, for a legacy promote and for a pipeline submit. The replay after a restart is idempotent, a changed payload is `gks_conflict`, and there is one row each. |

The real-MSP runs recorded at `be97c93` cannot be replayed from this
repository. They move to `externalRuns` in the manifest, keeping their
original date.

## Evidence that it detects regressions

Each mutation below was applied to the product code, the corpus was run, and
the mutation was reverted:

| Mutation | Cases that failed |
|---|---|
| Parse-error message changed | TRANSPORT-DENIAL (golden hash) |
| Replay reports `idempotent: false` | TOOL-012, GENESISRAG17-RECEIPTS (annotation and hash) |
| `aliases` dropped from the entity output | TOOL-003, TOOL-004, TENANT-WALL, AUTH-DENIAL (hash only; no annotation names the field) |
| `normVersion` value changed | the same four (hash only) |
| Server fails to start | all 24, each failing at once with the server's stderr |

Local runs:

- `npm run check:corpus` replays all 24 cases in about 3 seconds, stable over 5
consecutive runs;
- `npm test` gives 288 passed and 2 skipped (MSP suites that need
`MSP_REPO_ROOT`); security 12/12, unit 9/9;
- `check:c0` reports PASS 24 / NOT_RUN 1, and `check:baseline` holds after the
re-lock described below.

## CI and baseline

The `c0-gate` slice now runs `check:c0`, `check:baseline` and `check:corpus`.
`tests/integration/c0-corpus.test.mjs` runs the same replay inside `npm test`.
The workflow hash changed, so the baseline lock was re-locked
(`f50116a+working-tree`).

`check:c0` now also verifies, without running anything, that every fixture and
transcript hashes to the registry. It refuses a PASS for any case that has no
replayable fixture.

## Maintaining the corpus

1. Change `scripts/c0-corpus/build-cases.mjs`, or the behaviour a case records.
2. Run `node scripts/c0-corpus/build-cases.mjs`. It rewrites the fixtures and
clears the expected hash of every case whose request changed.
3. Run `node scripts/run-c0-corpus.mjs --write`. It re-baselines the
transcripts, registry hashes and manifest. It refuses if any annotation
fails or two runs disagree.
4. Review the diff of `expected/*.json`. This diff is the behaviour change
being approved.

## Observation, not changed here

A gate verdict for a batch with no worker receipt includes the security reason
"retrieval benchmark reported a cross-tenant leak". That reason is misleading:
there was no benchmark at all, and the retrieval dimension already reports
"retrieval benchmark is missing". The golden transcript records the current
text. Correcting it is a behaviour change for a later PR.

## Change log

| Version | Date | Status | Summary | Commit Hash | Agent |
|---|---|---|---|---|---|
| 0.1.0b | 2026-09-27 | candidate | C0.4 golden corpus registry v2: replayable fixtures, golden transcripts, runner, CI gate, Tier-4 readback split out | working-tree | Claude |
7 changes: 4 additions & 3 deletions docs/reports/2026-09-27-p1-c0-closure.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
version: "0.1.0b"
version: "0.1.1b"
created_at: "2026-09-27T10:30:00+07:00,Claude,working-tree"
last_update: "2026-09-27T10:30:00+07:00,Claude"
last_update: "2026-09-27T12:00:00+07:00,Claude"
status: "candidate"
superseded_by: null
attributes:
Expand Down Expand Up @@ -144,10 +144,11 @@ versioned rollout, not with this closure.
| GKS-PIP-008 | a legacy export with an out-of-range cursor returns empty | return `gks_invalid_request`, as the pipeline export does |
| GKS-IDN-006 | human repair proof is only an `msp:proof/` prefix | verify it against MSP |
| GKS-NFR-006 | the gate trusts benchmark numbers the worker reports | pin a GKS-side benchmark manifest |
| GKS-MIG-002 | the registry holds request and result hashes without payloads, so it cannot be replayed | ship registry v2 with request fixtures and a corpus runner, then re-baseline the hashes |
| GKS-MIG-002 | ~~the registry holds request and result hashes without payloads, so it cannot be replayed~~ | **Done:** registry v2 ships replayable fixtures and a runner; see [the corpus v2 report](2026-09-27-c0-corpus-v2.md) |

## Change log

| Version | Date | Status | Summary | Commit Hash | Agent |
|---|---|---|---|---|---|
| 0.1.0b | 2026-09-27 | candidate | P1 C0 closure: executable baseline lock and CI gate, lost-response replay PASS, IMMEDIATE write transactions and error-code discipline, C0 acceptance evidence | working-tree | Claude |
| 0.1.1b | 2026-09-27 | candidate | GKS-MIG-002 closed by the C0.4 golden corpus v2 | working-tree | Claude |
1 change: 1 addition & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@
"test:security": "node --test --test-concurrency=1 tests/security/*.security.mjs",
"check:c0": "node scripts/check-c0-qualification.mjs",
"check:baseline": "node scripts/check-baseline-lock.mjs",
"check:corpus": "node scripts/run-c0-corpus.mjs",
"test:all": "npm test",
"pack:client": "npm pack --workspace @freshair129/gks-client-js --dry-run",
"desktop": "node apps/wiki-desktop/server.mjs",
Expand Down
Loading
Loading