From 39198c79957c065db06183e536c0ed0b033d2bd4 Mon Sep 17 00:00:00 2001 From: Freshair129 <94353529+Freshair129@users.noreply.github.com> Date: Sun, 27 Sep 2026 09:51:48 +0700 Subject: [PATCH 1/2] fix: hide unpublished GenesisRAG17 entities from legacy reads Proposed in ADR-GKS-PIPELINE-VISIBILITY (awaiting owner approval). - Migration 0007 adds a GKS-owned entities.origin column. Only the pipeline submit writer sets 'pipeline'. The backfill marks a row pipeline-origin only when a run recorded it and the legacy path never created it, and adds an entity_id-leading index on pipeline_mentions. - search, getEntity (and so relations_get and artifact_link), getRelations, the legacy resolution pool (so resolveTo too) and D9 BIND/MERGE operands hide pipeline entities until a mentioning run is PUBLISHED. Stage 9 reuse opts in with includeUnpublishedPipeline. - A legacy norm key that imitates a hidden typed pipeline key is created under the D2 discriminator instead of surfacing gks_conflict. Co-Authored-By: Claude Opus 5.5 --- docs/ADR-GKS-C0-QUALIFICATION.md | 5 +- docs/ADR-GKS-GENESISRAG17.md | 5 +- docs/ADR-GKS-PIPELINE-VISIBILITY.md | 217 +++++++++++++++++ docs/GKS-DATA-MODEL.md | 17 +- docs/GKS-PORT-CONTRACT.md | 11 +- migrations/0007_pipeline_entity_origin.sql | 24 ++ packages/gks-core/src/index.mjs | 4 +- packages/gks-persistence/src/index.mjs | 118 ++++++--- tests/contract/pipeline-genesisrag17.test.mjs | 226 +++++++++++++++++- tests/contract/stage9-migration.test.mjs | 9 +- tests/security/cross-tenant-deny.security.mjs | 49 ++++ 11 files changed, 631 insertions(+), 54 deletions(-) create mode 100644 docs/ADR-GKS-PIPELINE-VISIBILITY.md create mode 100644 migrations/0007_pipeline_entity_origin.sql diff --git a/docs/ADR-GKS-C0-QUALIFICATION.md b/docs/ADR-GKS-C0-QUALIFICATION.md index 79f1865..371c0bb 100644 --- a/docs/ADR-GKS-C0-QUALIFICATION.md +++ b/docs/ADR-GKS-C0-QUALIFICATION.md @@ -1,7 +1,7 @@ --- -version: "0.2.0" +version: "0.2.1" created_at: "2026-09-22T00:00:00+07:00,RWANG,working-tree" -last_update: "2026-09-22T00:00:00+07:00,RWANG" +last_update: "2026-09-27T10:00:00+07:00,Claude" status: "beta" approval_owner: "Boss (บอส)" approval_recorded_at: "2026-09-22T00:00:00+07:00" @@ -216,3 +216,4 @@ review confirms: | 0.1.0b | 2026-09-22 | candidate | Proposed C0 qualification decisions for G1 review | working-tree | RWANG | | 0.1.0 | 2026-09-22 | beta | User approved D1-D4 for C0 implementation | working-tree | RWANG | | 0.2.0 | 2026-09-22 | beta | Clarified compatibility versus secure MSP auth mode and recorded real-chain qualification boundary | working-tree | RWANG | +| 0.2.1 | 2026-09-27 | beta | Cross-reference only: ADR-GKS-PIPELINE-VISIBILITY (proposed) changes what `gks_search`, `gks_entity_get`, `gks_relations_get` and `gks_artifact_link` return for unpublished GenesisRAG17 entities. No request is newly rejected; see that ADR's observable-changes table. | working-tree | Claude | diff --git a/docs/ADR-GKS-GENESISRAG17.md b/docs/ADR-GKS-GENESISRAG17.md index 0648653..7480ceb 100644 --- a/docs/ADR-GKS-GENESISRAG17.md +++ b/docs/ADR-GKS-GENESISRAG17.md @@ -1,7 +1,7 @@ --- -version: "0.5.0b" +version: "0.5.1b" created_at: "2026-09-07T23:30:00+07:00,RWANG,working-tree" -last_update: "2026-09-11T22:30:00+07:00,Claude Opus 5,working-tree" +last_update: "2026-09-27T10:00:00+07:00,Claude Opus 5.5,working-tree" status: "accepted" approval_owner: "Boss (บอส)" approval_recorded_at: "2026-09-07T23:00:00+07:00" @@ -221,6 +221,7 @@ evidence. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| +| 0.5.1b | 2026-09-27 | accepted | Cross-reference: ADR-GKS-PIPELINE-VISIBILITY (proposed) keeps submit-time canonical entities invisible to legacy reads until a mentioning run is PUBLISHED; pipeline tools and Stage 9 reuse are unchanged. | working-tree | Claude Opus 5.5 | | 0.5.0b | 2026-09-11 | accepted | Implemented structured-record profile contract revision 2 on the GKS side (ADR-075 Phase 2, rollout step 2 of 3): `PIPELINE_ONTOLOGY_VERSION` is `ontology_v2`, `PIPELINE_SUPPORTED_ONTOLOGY_VERSIONS` is `{ontology_v1, ontology_v2}`, the Stage 11 `validEndpoint` ternary is replaced by the per-version predicate -> endpoint table `PIPELINE_ONTOLOGY_ENDPOINTS` (v2 adds `HAS_COMPONENT`, `PRICED_AT`, `IN_CATEGORY` and the `PACKAGE`/`CATEGORY`/`PRICE_TIER` endpoint types), and the Stage 17 knowledge dimension checks set membership and each fact against its own version instead of equality with `ontology_v1`. `rule_v1`, `parseStructuredClaim`, Stage 12 and the C-10 bitemporal count are unchanged. Must merge after the GenesisBlock worker change that accepts both versions. | working-tree | Claude Opus 5 | | 0.4.4b | 2026-09-11 | accepted | Fixed a residual of contract item C-10: the Stage 17 expected bitemporal lane count now includes HELD rows that carry valid time, not only facts, using the same row classification the worker applies (`temporalRows()` covers facts and held together; a row is dated unless validFrom/validTo are undefined/null/`not_applicable` and status is undefined/`not_applicable`). A decision with a dated held row previously disagreed with the worker's count even though the 0.4.3b fix already matched on facts alone. The all-`not_applicable` path (expect 0, lane `not_applicable`) now also considers facts and held together. Held rows already make the knowledge dimension WARN; this only corrects the graph-dimension reason, never whether anything publishes. | working-tree | Claude Opus 5 | | 0.4.3b | 2026-09-11 | accepted | Fixed contract item C-10: the Stage 17 expected bitemporal lane count is the number of facts that carry valid time, not `facts.length`, matching the GenesisBlock worker's mapped-only count. A generation mixing dated and `not_applicable` facts previously failed the graph dimension on the lane-count comparison. The all-`not_applicable` path (expect 0, lane `not_applicable`) and all-dated generations are unchanged. | working-tree | Claude Opus 5 | diff --git a/docs/ADR-GKS-PIPELINE-VISIBILITY.md b/docs/ADR-GKS-PIPELINE-VISIBILITY.md new file mode 100644 index 0000000..1a45c2b --- /dev/null +++ b/docs/ADR-GKS-PIPELINE-VISIBILITY.md @@ -0,0 +1,217 @@ +--- +version: "0.1.0" +created_at: "2026-09-27T10:00:00+07:00,Claude,working-tree" +last_update: "2026-09-27T10:00:00+07:00,Claude" +status: "proposed" +approval_owner: "Boss (บอส)" +approval_recorded_at: null +superseded_by: null +attributes: + domain: "genesis-knowledge-system" + doc_type: "architecture-decision" + scope: "legacy-read visibility of GenesisRAG17 canonical entities before publication" +--- + +# ADR: GenesisRAG17 entities stay invisible to legacy reads until published + +## Decision status + +**Proposed, awaiting owner approval.** The implementation lands in the same +pull request so the decision can be judged against working code and tests, +but it must not merge before approval is recorded here. It changes the +observable read behaviour of four frozen C0 tools, and +[`ADR-GKS-C0-QUALIFICATION.md`](ADR-GKS-C0-QUALIFICATION.md) does not +authorize that on its own. + +## Context + +`transactPipelineSubmit` writes each GenesisRAG17 decision entity into the +shared canonical `entities` table at submit time, before Stage 13 has written +anything and before Stage 17 has gated anything. The legacy read tools +(`gks_search`, `gks_entity_get`, `gks_relations_get`, `gks_artifact_link`) +read that table with no notion of publication. So an entity that exists only +because of a `PENDING`, `REJECTED` or `FAILED_STAGE` run is returned as +canonical knowledge. The legacy resolver's pool reads the same table, so a +legacy promote can also resolve `MATCHED` onto such an entity, or use +`resolveTo` to probe whether an unpublished entity exists. + +A first attempt to fix this (withdrawn from the P0 pull request after +architecture review) keyed on `metadata.pipelineVersion`. That key is caller +data on the legacy promote path, so a caller could hide its own entities, and +the attempt left the resolution pool, D9 BIND/MERGE and relations +inconsistent with the read tools. + +Requirements this addresses from the proposed SRS blueprint: +`GKS-GOV-002` (promotion is not publication) and the legacy-read half of +`GKS-RET-002` (published-only default). It does not implement K2 retrieval. + +## Decision + +### D1 — Visibility rule + +A canonical entity is **visible to legacy reads** when it is: + +- of legacy origin, or +- of pipeline origin and mentioned by at least one GenesisRAG17 run whose + batch status is `PUBLISHED`. + +Every other pipeline-origin entity is **hidden**. Publication is monotonic in +C0: there is no unpublish, so once visible an entity stays visible, and a later +`REJECTED` run that mentions it does not hide it again. + +### D2 — Origin is a GKS-owned column, never caller data + +Migration `0007_pipeline_entity_origin.sql` adds +`entities.origin TEXT NOT NULL DEFAULT 'legacy' CHECK (origin IN ('legacy', 'pipeline'))`. + +- Only `insertPipelineEntity` writes `'pipeline'`. The legacy promote path + never sets the column. No request field, metadata key, candidate string or + norm key can change it. +- **Backfill** marks an existing row `'pipeline'` only when both of these hold: + - a GenesisRAG17 run recorded it as a decision entity (a `pipeline_mentions` + row names it); + - the legacy path never created it (no `entity_mentions` row with outcome + `CREATED` names it). +- Every legacy creation writes a `CREATED` mention, including the 0002 + backfill of pre-Stage-9 entities, so a legacy entity that a pipeline run + later reused stays `'legacy'`. Norm-key shape is deliberately not used: + `norm_v1` keeps U+0000, so a legacy candidate could imitate the typed + pipeline key. +- The migration also adds `idx_pipeline_mentions_entity_ref (entity_id, + scope_key, batch_id)`, so the visibility check is an index lookup, not a + scan of every mention. + +### D3 — Legacy resolution pool excludes hidden entities; the pipeline pool does not + +`lookupResolutionCandidates` excludes hidden entities by default, so neither +the legacy ladder nor `resolveTo` can match or probe them. A hidden +`resolveTo` target answers exactly like a missing one (`REJECTED`). + +The GenesisRAG17 Stage 9 reuse lookup passes +`includeUnpublishedPipeline: true`. Two runs of the same entity must still +converge on one identity whether or not the first run was published. + +Consequence: while a pipeline entity is hidden, a legacy promote of the same +real-world thing creates its own legacy entity. If the pipeline entity is +published later, the duplicate is repaired the sanctioned way, by a D9 +`MERGE` **with the pipeline entity as the survivor**. This trades a repairable +over-split for never serving unpublished knowledge. + +GenesisRAG17 Stage 9 reuse does not follow supersession: a superseded +pipeline entity drops out of both pools, and a later run would reuse the +superseded row by its deterministic id. That gap predates this ADR. Until +reuse follows supersession, a repair must not supersede a pipeline entity. + +**Norm-key collisions.** Usually the typed pipeline key +(`norm_v1(resolutionKey) + U+0000 + TYPE`) and a legacy `norm_v1` key differ. +They can coincide, though: + +- `norm_v1` keeps U+0000; +- a semantic type with no letter case (digits, Thai) is unchanged by + upper-casing. + +So a legacy candidate such as `"acme\u0000123"` produces exactly the key of a +pipeline entity `acme` of type `123`. When a legacy insert loses +`UNIQUE(scope_key, norm_key)` to a hidden pipeline row, the adapter does not +surface `gks_conflict`, which would reveal the hidden row. It creates the +legacy entity under the existing D2 human-distinct discriminator +(`norm_key#mention_id`). Later promotes of the same string reach that entity +through the EXACT rung, which compares candidate strings, so the split does +not repeat. A collision with a *visible* row keeps the decision-5 retry. + +### D4 — D9 BIND/MERGE operate only on visible entities + +`transactHumanResolution` resolves its `canonicalRef`, `survivorRef` and +`supersededRef` with the same visibility rule. A hidden ref answers with the +existing "does not resolve to a canonical entity" error. It is never +`gks_scope_denied`, which would confirm that the entity exists. + +### D5 — Relations never expose a hidden endpoint + +`getRelations` omits any relation whose `from_ref` or `to_ref` names a hidden +entity. After D3 no new legacy relation can point at a hidden entity. This +covers rows written before the upgrade. + +### D6 — What does not change + +- No row, ref, hash, receipt, cursor or ledger entry is deleted or rewritten. +- Response shapes are unchanged: `origin` is not added to any tool output. +- The GenesisRAG17 tools (submit, claim, receipts, gate, publication, + evidence) and their scope rules are unchanged. The worker still receives + every decision entity through `gks_pipeline_claim`. +- The legacy scope mapping that drops `agentId` (`GKS-SCP-004` / `GKS-SCP-005`) + is out of scope. A published entity is visible to the whole legacy scope, as + before. +- **The rule is enforced on reads and resolution, not on every write.** + Promotion's fill of an existing row, supersession following, and pipeline + submit still look entities up by ref without the visibility predicate. That + is safe because core only hands those paths refs taken from the visible pool + or computed from the caller's own scope and candidate string, never a + caller-chosen ref. A future write path that accepts a caller ref must use the + visible lookup. + +## Observable changes after upgrade + +For an entity that exists only through unpublished runs: + +| Tool | Before | After | +|---|---|---| +| `gks_search` | returned | omitted | +| `gks_entity_get` | returned (or `gks_scope_denied` from a foreign scope) | `null` from any scope | +| `gks_relations_get` on it | relations returned (or `gks_scope_denied`) | `[]` | +| `gks_relations_get` on a neighbour | relation to it returned | relation omitted | +| `gks_artifact_link` to it | link written | `gks_invalid_request` ("does not resolve") | +| `gks_knowledge_promote` | may resolve `MATCHED` / `resolveTo` onto it | never resolves onto it; may create a separate entity (D3) | +| D9 BIND/MERGE naming it | accepted | refused as not resolving | +| Earlier promote snapshots and `gks_stage_evidence_export` rows that name it | ref resolved | the ref is kept unchanged (D6), but reads it as `null` until publication | + +No accepted request shape is rejected at validation, so no versioned wire +rollout is needed. The change is recorded here and in the C0 qualification +ADR's revision history instead. + +## Rollback + +- Migration 0007 is additive. +- A binary that includes the schema-ahead guard (`GKS_SCHEMA_AHEAD`) refuses to + open a 0007 store with an older artifact. Rollback is then restoring the + pre-migration backup, as the production runbook requires. +- An older binary without the guard would open the store, ignore `origin`, and + show unpublished entities again. No data is lost either way. + +## Alternatives rejected + +- **Stage pipeline entities in a separate table until publication.** This is + the cleanest end state, but it rewrites Stage 9 reuse, the decision payload + and every receipt check that joins `entities`. It is disproportionate for a + visibility rule. Revisit with K1 revisions. +- **Filter in `gks-core` instead of SQL.** The pool, D9 and relations + predicates live in the adapter's SQL by design (`GKS-PORT-CONTRACT`: caller + filtering is a repairable leak). Filtering in core would split one rule + across two layers. +- **Key on `metadata.pipelineVersion` or on norm-key shape.** Both are + reachable by a legacy caller (D2). +- **Let legacy promote keep matching hidden entities.** The caller then holds a + canonical ref that `gks_entity_get` reports as `null`, and `resolveTo` + becomes an existence oracle for unpublished content. + +## Acceptance criteria + +- Pipeline entities are hidden from all four legacy read tools while their run + is `PENDING`, `REJECTED` or `FAILED_STAGE`, and appear after publication. +- An entity first mentioned by a `REJECTED` run becomes visible when a later + run that reuses it is `PUBLISHED`. +- A legacy entity whose metadata carries `pipelineVersion` stays visible. +- A legacy promote neither matches nor `resolveTo`-probes a hidden entity. A + legacy string whose norm key imitates a hidden typed key is created, not + refused, and a repeat of that string matches the entity it created. +- A `FAILED_STAGE` run stays hidden. +- D9 BIND/MERGE refuse hidden refs with the not-resolving error. +- Relations touching a hidden entity are omitted. +- A foreign-tenant caller learns nothing about a hidden entity. +- The backfill marks pre-existing rows exactly as D2 states. + +## CHANGELOG + +| Version | Date | Status | Summary | Commit Hash | Agent | +|---|---|---|---|---|---| +| 0.1.0 | 2026-09-27 | proposed | Proposed hiding unpublished GenesisRAG17 entities from legacy reads through a GKS-owned `origin` column, with consistent resolution-pool, D9 and relation rules. | working-tree | Claude | diff --git a/docs/GKS-DATA-MODEL.md b/docs/GKS-DATA-MODEL.md index 23ae615..bf6a591 100644 --- a/docs/GKS-DATA-MODEL.md +++ b/docs/GKS-DATA-MODEL.md @@ -1,7 +1,7 @@ --- -version: "0.7.0b" +version: "0.8.0b" created_at: "2026-08-12T10:05:34+07:00,ATHER,working-tree" -last_update: "2026-09-08T04:20:00+07:00,RWANG" +last_update: "2026-09-27T10:00:00+07:00,Claude" status: "beta" approval_owner: "Boss (บอส)" approval_recorded_at: "2026-08-12T10:16:19+07:00" @@ -245,11 +245,21 @@ the mention string is no longer the identity — replaced by over-split can never be matched — or reported `AMBIGUOUS` — against its own ghost; its spellings stay reachable through the aliases the merge copied onto the survivor. +- `origin` (`TEXT NOT NULL DEFAULT 'legacy'`, `'legacy'` or `'pipeline'`, added + by `0007_pipeline_entity_origin.sql`) — which path created the row. Only + the GenesisRAG17 submit writer sets `'pipeline'`; no request field or + metadata key reaches it. A pipeline-origin entity is invisible to legacy + reads, the legacy resolution pool and D9 operands until a run that mentions + it is `PUBLISHED` (`ADR-GKS-PIPELINE-VISIBILITY.md`). The migration + backfills `'pipeline'` only for rows a pipeline run recorded and the legacy + path never created (no `CREATED` mention). Indexes: `idx_entities_search (portfolio_id, type, title, canonical_ref)` (unchanged), `idx_entities_pool (portfolio_id, tenant_id, business_id, workspace_id, project_id)` (new, created by migration 0002's backfill hook), -`idx_entities_superseded (superseded_by)` (new, migration 0004). +`idx_entities_superseded (superseded_by)` (new, migration 0004), and on the +pipeline side `idx_pipeline_mentions_entity_ref (entity_id, scope_key, batch_id)` +(migration 0007) for the visibility check. ### `entity_mentions` — one row per occurrence, not per string @@ -494,6 +504,7 @@ tables and write rules. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| +| 0.8.0b | 2026-09-27 | beta | Proposed `entities.origin` (migration 0007) and the publication-visibility rule for GenesisRAG17 entities in legacy reads (ADR-GKS-PIPELINE-VISIBILITY). | working-tree | Claude | | 0.7.0b | 2026-09-08 | beta | Recorded the GenesisRAG17 typed Stage 9 entity key and semantic type storage in the shared entities table while preserving every pipeline mention occurrence. | working-tree | RWANG | | 0.6.0b | 2026-09-08 | beta | Expanded the GenesisRAG17 data model with exact migration 0006 table shapes, immutable decision facts/occurrences, receipt/gate snapshots, cursor ordering, and the no-Stage-18 extension boundary. | 9279cfe | RWANG | | 0.5.0b | 2026-09-07 | beta | Clarified the graph-receipt-to-enrichment boundary, post-acknowledgement physical projections, gate statistics and Tier4 failure-only terminal rows. | working-tree | RWANG | diff --git a/docs/GKS-PORT-CONTRACT.md b/docs/GKS-PORT-CONTRACT.md index f535be4..9321749 100644 --- a/docs/GKS-PORT-CONTRACT.md +++ b/docs/GKS-PORT-CONTRACT.md @@ -1,7 +1,7 @@ --- -version: "0.10.0b" +version: "0.10.1b" created_at: "2026-08-12T10:05:34+07:00,ATHER,working-tree" -last_update: "2026-09-24T10:06:32+07:00,RWANG" +last_update: "2026-09-27T10:00:00+07:00,Claude" status: "beta" approval_owner: "Boss (บอส)" approval_recorded_at: "2026-08-12T10:16:19+07:00" @@ -279,7 +279,11 @@ interface GksPersistencePortV2 extends GksPersistencePort { // Stage 9 blocking lookup: the candidate rows a resolver may consider. // MUST filter every scope dimension in SQL, never in the caller. // Excludes superseded entities: a D9-merged row is not a live identity. - lookupResolutionCandidates(input: ScopedResolutionQuery): Promise; + // Excludes unpublished GenesisRAG17 entities unless the caller is Stage 9 + // pipeline reuse and passes includeUnpublishedPipeline: true + // (ADR-GKS-PIPELINE-VISIBILITY D3). getEntity, search, getRelations and the + // D9 operands apply the same visibility rule in SQL. + lookupResolutionCandidates(input: ScopedResolutionQuery & { includeUnpublishedPipeline?: boolean }): Promise; // D9 read: unresolved mentions (REVIEW_REQUIRED / AMBIGUOUS, canonical // ref NULL) within scope. Same SQL scope predicate as the lookup. listUnresolvedMentions(input: ScopedReviewQuery): Promise; @@ -537,6 +541,7 @@ implementation package name appears in the client. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| +| 0.10.1b | 2026-09-27 | beta | Proposed (ADR-GKS-PIPELINE-VISIBILITY): legacy reads, the legacy resolution pool and D9 operands exclude unpublished GenesisRAG17 entities; `lookupResolutionCandidates` gains `includeUnpublishedPipeline` for Stage 9 reuse. No tool request or result shape changes. | working-tree | Claude | | 0.10.0b | 2026-09-24 | beta | Implements optional per-client hash-backed direct HTTP read grants while preserving the MSP profile; grants are not enabled by default and production rollout remains separate. | working-tree | RWANG | | 0.9.0b | 2026-09-24 | beta | Adds the approved direct-client read-only grant profile while preserving MSP auth for governed writes; concrete identity verification and activation remain unimplemented. | working-tree | RWANG | | 0.8.1b | 2026-09-22 | beta | Removed the stale port-version-1 statement that contradicted the selected GKS-owned SQLite production profile; deployment evidence remains separately gated. | working-tree | RWANG | diff --git a/migrations/0007_pipeline_entity_origin.sql b/migrations/0007_pipeline_entity_origin.sql new file mode 100644 index 0000000..4f4139c --- /dev/null +++ b/migrations/0007_pipeline_entity_origin.sql @@ -0,0 +1,24 @@ +-- @req GKS-GOV-002, GKS-RET-002 (legacy-read half) — unpublished GenesisRAG17 +-- entities are invisible to legacy reads. +-- @spec docs/ADR-GKS-PIPELINE-VISIBILITY.md (D1, D2) +-- @tested tests/contract/pipeline-genesisrag17.test.mjs, tests/security/cross-tenant-deny.security.mjs +-- +-- entities.origin records which path created the row. Only the pipeline +-- submit writer sets 'pipeline'; no request field, metadata key or norm key +-- can reach it. Legacy rows keep the default. +ALTER TABLE entities ADD COLUMN origin TEXT NOT NULL DEFAULT 'legacy' CHECK (origin IN ('legacy', 'pipeline')); + +-- Backfill (D2): a row is pipeline-origin when a GenesisRAG17 run recorded it +-- as a decision entity AND the legacy path never created it. Every legacy +-- creation, including 0002's backfill of pre-Stage-9 rows, wrote a CREATED +-- mention, so a legacy entity that a pipeline run later reused stays legacy. +UPDATE entities +SET origin = 'pipeline' +WHERE EXISTS (SELECT 1 FROM pipeline_mentions m WHERE m.entity_id = entities.canonical_ref) + AND NOT EXISTS ( + SELECT 1 FROM entity_mentions em + WHERE em.canonical_ref = entities.canonical_ref AND em.outcome = 'CREATED' + ); + +-- The visibility check looks mentions up by entity, across scopes. +CREATE INDEX idx_pipeline_mentions_entity_ref ON pipeline_mentions (entity_id, scope_key, batch_id); diff --git a/packages/gks-core/src/index.mjs b/packages/gks-core/src/index.mjs index 2b8b24f..5e45e21 100644 --- a/packages/gks-core/src/index.mjs +++ b/packages/gks-core/src/index.mjs @@ -95,7 +95,9 @@ export function createGksService({ persistence, defaultPortfolioId, automergeFlo async function existingPipelineCanonicalRefs(scope, mentions) { requirePipelinePersistence("lookupResolutionCandidates"); - const candidates = await persistence.lookupResolutionCandidates({ scope: pipelineLegacyScope(scope) }); + // ADR-GKS-PIPELINE-VISIBILITY D3: Stage 9 reuse sees unpublished pipeline + // entities so repeated runs of one entity converge on one identity. + const candidates = await persistence.lookupResolutionCandidates({ scope: pipelineLegacyScope(scope), includeUnpublishedPipeline: true }); const wanted = new Set(mentions.map((mention) => pipelineEntityNormKey(mention.resolutionKey, mention.semanticType))); const refs = new Map(); for (const candidate of candidates) { diff --git a/packages/gks-persistence/src/index.mjs b/packages/gks-persistence/src/index.mjs index c4add70..003712c 100644 --- a/packages/gks-persistence/src/index.mjs +++ b/packages/gks-persistence/src/index.mjs @@ -281,6 +281,17 @@ function rowScope(row) { }; } +// ADR-GKS-PIPELINE-VISIBILITY D1: a legacy-origin entity is always visible to +// legacy reads; a pipeline-origin entity only once a run that mentions it is +// PUBLISHED. `origin` is written by GKS alone (D2), never taken from a request. +function visibleEntity(alias) { + return `(${alias}.origin = 'legacy' OR EXISTS ( + SELECT 1 FROM pipeline_mentions vm + JOIN pipeline_batches vb ON vb.batch_id = vm.batch_id AND vb.scope_key = vm.scope_key + WHERE vm.entity_id = ${alias}.canonical_ref AND vb.status = 'PUBLISHED' + ))`; +} + function entityFromRow(row) { if (!row) return null; return { @@ -572,44 +583,54 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO const diffs = fillExistingEntity(stored, entity, now, graphVersion); if (diffs.length) fieldDiffs = JSON.stringify(diffs); } else if (entity.canonicalRef) { + const row = { + canonical_ref: entity.canonicalRef, + scope_key: input.scopeKey, + candidate_ref: entity.candidateRef, + type: entity.type, + title: entity.title, + summary: entity.summary, + source_ref: entity.sourceRef, + confidence: entity.confidence, + portfolio_id: scope.portfolioId, + tenant_id: scope.tenantId, + business_id: scope.businessId, + workspace_id: scope.workspaceId, + project_id: scope.projectId, + sharing: scope.sharing, + metadata_json: JSON.stringify(entity.metadata), + aliases_json: JSON.stringify([...(entity.aliases ?? [])].sort()), + external_refs_json: JSON.stringify([...(entity.externalRefs ?? [])].sort()), + norm_key: entity.normKey, + norm_version: entity.normVersion, + created_at: now, + updated_at: now, + graph_version: graphVersion, + }; try { - insertEntity.run({ - canonical_ref: entity.canonicalRef, - scope_key: input.scopeKey, - candidate_ref: entity.candidateRef, - type: entity.type, - title: entity.title, - summary: entity.summary, - source_ref: entity.sourceRef, - confidence: entity.confidence, - portfolio_id: scope.portfolioId, - tenant_id: scope.tenantId, - business_id: scope.businessId, - workspace_id: scope.workspaceId, - project_id: scope.projectId, - sharing: scope.sharing, - metadata_json: JSON.stringify(entity.metadata), - aliases_json: JSON.stringify([...(entity.aliases ?? [])].sort()), - external_refs_json: JSON.stringify([...(entity.externalRefs ?? [])].sort()), - norm_key: entity.normKey, - norm_version: entity.normVersion, - created_at: now, - updated_at: now, - graph_version: graphVersion, - }); + insertEntity.run(row); } catch (error) { - if (isNormKeyUniqueViolation(error)) { + if (!isNormKeyUniqueViolation(error)) throw error; + const winner = selectEntityByNormKey.get(input.scopeKey, entity.normKey); + if (winner && !selectVisibleEntityByRef.get(winner.canonical_ref)) { + // ADR-GKS-PIPELINE-VISIBILITY D3: the key is held by an unpublished + // pipeline entity (norm_v1 keeps U+0000, so a legacy string can + // imitate a typed pipeline key whose type has no letter case). A + // hidden row must answer like an absent one, so the legacy entity + // is created under D2's human-distinct discriminator instead of + // surfacing a conflict. Later promotes of the same string reach it + // through the EXACT rung, which compares candidate strings. + insertEntity.run({ ...row, norm_key: `${entity.normKey}#${mentionId(input.scopeKey, input.idempotencyKey, entity.candidateRef)}` }); + } else { // Decision 5: surface the loss of the UNIQUE(scope_key, norm_key) // race with the winning row attached. The whole envelope rolls // back; the domain layer retries and returns MATCHED against the // winner rather than over-splitting or silently merging. - const winner = selectEntityByNormKey.get(input.scopeKey, entity.normKey); throw new GksNormKeyConflictError( `An entity with norm_key "${entity.normKey}" already exists in this scope.`, { candidateRef: entity.candidateRef, normKey: entity.normKey, winner: entityFromRow(winner) }, ); } - throw error; } } // D1: one mention row per occurrence — the audit trail promotion @@ -737,7 +758,7 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO // the merge a repair -- a repaired over-split can never be MATCHED, or // reported AMBIGUOUS, against its own ghost. Its spellings stay reachable // through the aliases the merge copied onto the survivor. - const selectResolutionPool = db.prepare(` + const resolutionPoolSql = (visibilityPredicate) => ` SELECT * FROM entities WHERE portfolio_id = @portfolioId AND tenant_id = @tenantId @@ -745,8 +766,17 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO AND (workspace_id = '' OR workspace_id = @workspaceId) AND (project_id = '' OR project_id = @projectId) AND superseded_by IS NULL + AND ${visibilityPredicate} ORDER BY created_at, canonical_ref - `); + `; + // ADR-GKS-PIPELINE-VISIBILITY D3: the legacy ladder (and resolveTo) never + // sees an unpublished pipeline entity; GenesisRAG17 Stage 9 reuse does, so + // repeated runs of one entity converge whether or not the first published. + const selectResolutionPool = db.prepare(resolutionPoolSql(visibleEntity("entities"))); + const selectPipelineResolutionPool = db.prepare(resolutionPoolSql("1 = 1")); + // D4: D9 BIND/MERGE operands resolve under the same visibility rule; a + // hidden ref answers exactly like a missing one. + const selectVisibleEntityByRef = db.prepare(`SELECT * FROM entities WHERE canonical_ref = ? AND ${visibleEntity("entities")}`); // ------------------------------------------------------------------------- // D9: the unresolved-mention consumer (ADR-GKS-ENTITY-RESOLUTION D9, D10.2, @@ -908,7 +938,7 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO if (mention.canonical_ref !== null || !UNRESOLVED_OUTCOMES.includes(mention.outcome)) { throw new GksInvalidRequestError("mentionId does not name an unresolved mention."); } - const target = selectEntityByRef.get(input.canonicalRef); + const target = selectVisibleEntityByRef.get(input.canonicalRef); if (!target) throw new GksInvalidRequestError("canonicalRef does not resolve to a canonical entity."); if (!inMentionPool(target, mention)) { throw new GksScopeDeniedError("bind target is outside the mention's resolution pool."); @@ -966,9 +996,9 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO }; } - const survivor = selectEntityByRef.get(input.survivorRef); + const survivor = selectVisibleEntityByRef.get(input.survivorRef); if (!survivor) throw new GksInvalidRequestError("survivorRef does not resolve to a canonical entity."); - const loser = selectEntityByRef.get(input.supersededRef); + const loser = selectVisibleEntityByRef.get(input.supersededRef); if (!loser) throw new GksInvalidRequestError("supersededRef does not resolve to a canonical entity."); if (!withinRequestScope(survivor, input.scope) || !withinRequestScope(loser, input.scope)) { throw new GksScopeDeniedError("merge operands must both be within the request scope."); @@ -1166,8 +1196,8 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO `); const updatePipelineBatchStatus = db.prepare("UPDATE pipeline_batches SET status = @status, updated_at = @updated_at WHERE scope_key = @scope_key AND decision_id = @decision_id"); const insertPipelineEntity = db.prepare(` - INSERT INTO entities (canonical_ref, scope_key, candidate_ref, type, title, summary, source_ref, confidence, portfolio_id, tenant_id, business_id, workspace_id, project_id, sharing, metadata_json, aliases_json, external_refs_json, norm_key, norm_version, created_at, updated_at, graph_version) - VALUES (@canonical_ref, @scope_key, @candidate_ref, @type, @title, '', @source_ref, NULL, @portfolio_id, @tenant_id, @business_id, @workspace_id, '', 'private', @metadata_json, '[]', '[]', @norm_key, @norm_version, @created_at, @updated_at, @graph_version) + INSERT INTO entities (canonical_ref, scope_key, candidate_ref, type, title, summary, source_ref, confidence, portfolio_id, tenant_id, business_id, workspace_id, project_id, sharing, metadata_json, aliases_json, external_refs_json, norm_key, norm_version, created_at, updated_at, graph_version, origin) + VALUES (@canonical_ref, @scope_key, @candidate_ref, @type, @title, '', @source_ref, NULL, @portfolio_id, @tenant_id, @business_id, @workspace_id, '', 'private', @metadata_json, '[]', '[]', @norm_key, @norm_version, @created_at, @updated_at, @graph_version, 'pipeline') `); function pipelineLegacyScope(scope) { @@ -1602,19 +1632,29 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO transactPromotion, search({ query, portfolioId }) { const pattern = `%${query.toLowerCase()}%`; - return db.prepare(`SELECT * FROM entities WHERE portfolio_id = ? AND (lower(canonical_ref) LIKE ? OR lower(title) LIKE ? OR lower(summary) LIKE ? OR lower(type) LIKE ?) ORDER BY canonical_ref`).all(portfolioId, pattern, pattern, pattern, pattern).map(entityFromRow); + return db.prepare(`SELECT * FROM entities WHERE portfolio_id = ? AND (lower(canonical_ref) LIKE ? OR lower(title) LIKE ? OR lower(summary) LIKE ? OR lower(type) LIKE ?) AND ${visibleEntity("entities")} ORDER BY canonical_ref`).all(portfolioId, pattern, pattern, pattern, pattern).map(entityFromRow); }, getEntity(ref) { - return entityFromRow(db.prepare("SELECT * FROM entities WHERE canonical_ref = ?").get(ref)); + return entityFromRow(selectVisibleEntityByRef.get(ref)); }, + // D5: a relation whose endpoint is a hidden entity is omitted, so no + // legacy read can surface an unpublished ref through a neighbour. getRelations(ref) { - return db.prepare("SELECT * FROM relations WHERE from_ref = ? OR to_ref = ? ORDER BY canonical_ref").all(ref, ref).map(relationFromRow); + return db.prepare(` + SELECT * FROM relations + WHERE (from_ref = ? OR to_ref = ?) + AND NOT EXISTS ( + SELECT 1 FROM entities e + WHERE e.canonical_ref IN (relations.from_ref, relations.to_ref) AND NOT ${visibleEntity("e")} + ) + ORDER BY canonical_ref + `).all(ref, ref).map(relationFromRow); }, - lookupResolutionCandidates({ scope } = {}) { + lookupResolutionCandidates({ scope, includeUnpublishedPipeline = false } = {}) { if (!scope || typeof scope !== "object" || typeof scope.portfolioId !== "string" || !scope.portfolioId) { throw new GksInvalidRequestError("lookupResolutionCandidates requires a scope with a portfolioId."); } - return selectResolutionPool.all({ + return (includeUnpublishedPipeline ? selectPipelineResolutionPool : selectResolutionPool).all({ portfolioId: scope.portfolioId, tenantId: scope.tenantId ?? "", businessId: scope.businessId ?? "", diff --git a/tests/contract/pipeline-genesisrag17.test.mjs b/tests/contract/pipeline-genesisrag17.test.mjs index a1207f0..8d8badf 100644 --- a/tests/contract/pipeline-genesisrag17.test.mjs +++ b/tests/contract/pipeline-genesisrag17.test.mjs @@ -1,4 +1,5 @@ import { afterEach, describe, expect, it } from "vitest"; +import Database from "better-sqlite3"; import { mkdtempSync, rmSync } from "node:fs"; import { tmpdir } from "node:os"; import path from "node:path"; @@ -15,6 +16,7 @@ import { validatePipelineBatch, } from "@freshair129/gks-contracts"; import { buildPipelineDecision, pipelineReadbackExpectations } from "@freshair129/gks-core"; +import { promotion } from "../fixtures/candidates.mjs"; const cleanups = []; afterEach(() => { @@ -240,7 +242,8 @@ describe("GenesisRAG17 pipeline contract", () => { expect(decision.facts).toHaveLength(1); expect(decision.facts[0]).toMatchObject({ predicate: "PURCHASED" }); expect(decision.facts[0].subjectId).not.toBe(decision.facts[0].objectId); - const rows = persistence.lookupResolutionCandidates({ scope: { ...batch.scope, projectId: "" } }); + // Stage 9 reuse reads the pipeline pool, which includes unpublished runs. + const rows = persistence.lookupResolutionCandidates({ scope: { ...batch.scope, projectId: "" }, includeUnpublishedPipeline: true }); expect(rows.map((row) => row.type).sort()).toEqual(["Person", "Product"]); expect(new Set(rows.map((row) => row.normKey)).size).toBe(2); @@ -601,3 +604,224 @@ describe("GenesisRAG17 pipeline contract", () => { expect(gate.verdict.dimensions.knowledge).toEqual({ result: "PASS", critical: false, reasons: [] }); }); }); + +// ADR-GKS-PIPELINE-VISIBILITY: a GenesisRAG17 entity written at submit is not +// canonical knowledge for the legacy tools until a run that mentions it is +// PUBLISHED. The legacy view of a pipeline scope drops agentId and projectId. +describe("GenesisRAG17 entities before publication (ADR-GKS-PIPELINE-VISIBILITY)", () => { + const legacyView = (pipelineScope) => ({ portfolioId: pipelineScope.portfolioId, tenantId: pipelineScope.tenantId, businessId: pipelineScope.businessId, workspaceId: pipelineScope.workspaceId, projectId: "", sharing: pipelineScope.visibility }); + + async function publishRun(service, batch) { + const { decision } = await submitAndClaim(service, batch); + const worker = auth(batch.scope, "worker"); + const graphReceipt = graphReceiptFor(decision); + const graphResult = await service.pipelineGraphReceipt({ receipt: graphReceipt, ...worker }); + const receipt = receiptFor(decision, graphResult, graphReceipt); + const written = await service.pipelineWriteReceipt({ receipt, ...worker }); + const gate = await service.pipelineGate({ schemaVersion: PIPELINE_SCHEMA_VERSION, scope: batch.scope, decisionId: decision.decisionId, decisionHash: decision.decisionHash, ...worker }); + expect(gate.verdict.verdict).toBe("PASS"); + await service.pipelinePublicationReceipt({ receipt: { schemaVersion: PIPELINE_SCHEMA_VERSION, scope: batch.scope, runId: batch.runId, decisionId: decision.decisionId, decisionHash: decision.decisionHash, snapshotId: receipt.snapshotId, generation: receipt.generation, receiptHash: written.receiptHash, publishedAt: "2026-09-07T15:00:03.000Z", pointerHash: "c".repeat(64), modelRevision: receipt.model.revision, transactionFrontier: receipt.transaction.frontier, readback: { ok: true } }, ...worker }); + return decision; + } + + // No Tier-4 receipt: the gate FAILs and the batch is REJECTED. + async function rejectRun(service, batch) { + const { decision } = await submitAndClaim(service, batch); + const gate = await service.pipelineGate({ schemaVersion: PIPELINE_SCHEMA_VERSION, scope: batch.scope, decisionId: decision.decisionId, decisionHash: decision.decisionHash, ...auth(batch.scope, "worker") }); + expect(gate.verdict.verdict).toBe("FAIL"); + return decision; + } + + async function legacyRead(service, pipelineScope, entity) { + const view = legacyView(pipelineScope); + return { + entity: await service.getEntity({ ref: entity.id, scope: view }), + searchHit: (await service.search({ query: entity.name, scope: view })).some((row) => row.canonicalRef === entity.id), + relations: await service.getRelations({ ref: entity.id, scope: view }), + }; + } + + const hidden = { entity: null, searchHit: false, relations: [] }; + + let legacyCounter = 0; + function legacyPromotion(pipelineScope, entities, overrides = {}) { + legacyCounter += 1; + return promotion({ + idempotency_key: `visibility-${legacyCounter}`, + provenance_ref: `msp:proof/visibility-${legacyCounter}`, + source_snapshot_hash: legacyCounter.toString(16).padStart(64, "0"), + scope: legacyView(pipelineScope), + candidate: { entities, relations: [] }, + ...overrides, + }); + } + + it("hides entities of a pending or rejected run from every legacy read and reveals them on publication", async () => { + const { service } = harness(); + const pending = makeBatch({ id: "batch-visibility-pending" }); + const { decision } = await submitAndClaim(service, pending); + for (const entity of decision.entities) expect(await legacyRead(service, pending.scope, entity)).toEqual(hidden); + await expect(service.linkArtifact({ knowledgeRef: decision.entities[0].id, artifactRef: "project:PRJ-1", relationType: "RELATED_TO", evidenceRef: "msp:proof/link-hidden", scope: legacyView(pending.scope) })) + .rejects.toMatchObject({ code: "gks_invalid_request" }); + + const rejected = makeBatch({ id: "batch-visibility-rejected", scope: scope({ agentId: "agent-rejected" }), entries: [{ text: "Dana works for Delta Ltd.", mentions: [["Dana", "dana", "Person"], ["Delta Ltd.", "delta", "Organization"]] }] }); + const rejectedDecision = await rejectRun(service, rejected); + for (const entity of rejectedDecision.entities) expect(await legacyRead(service, rejected.scope, entity)).toEqual(hidden); + + const published = makeBatch({ id: "batch-visibility-published", scope: scope({ agentId: "agent-published" }), entries: [{ text: "Erin works for Echo Ltd.", mentions: [["Erin", "erin", "Person"], ["Echo Ltd.", "echo", "Organization"]] }] }); + const publishedDecision = await publishRun(service, published); + for (const entity of publishedDecision.entities) { + const read = await legacyRead(service, published.scope, entity); + expect(read.entity).toMatchObject({ canonicalRef: entity.id }); + expect(read.searchHit).toBe(true); + } + }); + + it("reveals an entity first seen by a rejected run once a later run that reuses it is published", async () => { + const { service } = harness(); + const entries = [{ text: "Frank works for Foxtrot Ltd.", mentions: [["Frank", "frank", "Person"], ["Foxtrot Ltd.", "foxtrot", "Organization"]] }]; + const first = await rejectRun(service, makeBatch({ id: "batch-reuse-rejected", scope: scope({ agentId: "agent-reuse-1" }), entries })); + const frank = first.entities.find((entity) => entity.metadata.resolutionKey === "frank"); + expect(await legacyRead(service, scope(), frank)).toEqual(hidden); + + const second = await publishRun(service, makeBatch({ id: "batch-reuse-published", scope: scope({ agentId: "agent-reuse-2" }), entries })); + expect(second.entities.find((entity) => entity.metadata.resolutionKey === "frank").id).toBe(frank.id); + expect((await legacyRead(service, scope(), frank)).entity).toMatchObject({ canonicalRef: frank.id }); + }); + + it("never lets a legacy promote match or resolveTo-probe a hidden entity", async () => { + const { service } = harness(); + const { decision } = await submitAndClaim(service, makeBatch({ id: "batch-visibility-pool" })); + const alice = decision.entities.find((entity) => entity.metadata.resolutionKey === "alice"); + + const same = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "alice", type: "ENTITY", title: "Alice" }])); + const sameMapping = same.canonical_mappings.find((item) => item.candidateRef === "alice"); + expect(sameMapping.resolution.outcome).toBe("CREATED"); + expect(sameMapping.canonicalRef).not.toBe(alice.id); + + const probe = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "alice probe", type: "ENTITY", title: "Alice probe", resolveTo: alice.id }])); + expect(probe.canonical_mappings.find((item) => item.candidateRef === "alice probe").resolution.outcome).toBe("REJECTED"); + }); + + it("refuses D9 bind and merge on a hidden entity exactly as on a missing one", async () => { + const { service } = harness(); + const { decision } = await submitAndClaim(service, makeBatch({ id: "batch-visibility-d9" })); + const hiddenRef = decision.entities[0].id; + const view = legacyView(scope()); + + const created = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "Beta Co", type: "ENTITY", title: "Beta Industrial Holdings" }])); + const betaRef = created.canonical_mappings[0].canonicalRef; + await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "Beta Co", type: "ENTITY", title: "Totally Different Name" }])); + const [reviewRow] = await service.listUnresolvedMentions({ scope: view }); + expect(reviewRow).toMatchObject({ candidateRef: "Beta Co", outcome: "REVIEW_REQUIRED" }); + + const notResolving = { code: "gks_invalid_request", message: expect.stringContaining("does not resolve") }; + await expect(service.applyHumanResolution({ action: "BIND", mentionId: reviewRow.mentionId, canonicalRef: hiddenRef, provenanceRef: "msp:proof/bind-hidden", scope: view })).rejects.toMatchObject(notResolving); + await expect(service.applyHumanResolution({ action: "MERGE", survivorRef: betaRef, supersededRef: hiddenRef, provenanceRef: "msp:proof/merge-hidden-loser", scope: view })).rejects.toMatchObject(notResolving); + await expect(service.applyHumanResolution({ action: "MERGE", survivorRef: hiddenRef, supersededRef: betaRef, provenanceRef: "msp:proof/merge-hidden-survivor", scope: view })).rejects.toMatchObject(notResolving); + }); + + it("omits a relation whose endpoint is hidden and keeps a legacy entity that carries metadata.pipelineVersion", async () => { + const { service, directory } = harness(); + const { decision } = await submitAndClaim(service, makeBatch({ id: "batch-visibility-relations" })); + const hiddenRef = decision.entities[0].id; + const view = legacyView(scope()); + + const promoted = await service.promoteCandidate(legacyPromotion(scope(), [ + { candidateRef: "Gamma Co", type: "ENTITY", title: "Gamma Co", metadata: { pipelineVersion: PIPELINE_SCHEMA_VERSION } }, + { candidateRef: "Delta Co", type: "ENTITY", title: "Delta Co" }, + ], {})); + const gammaRef = promoted.canonical_mappings.find((item) => item.candidateRef === "Gamma Co").canonicalRef; + const deltaRef = promoted.canonical_mappings.find((item) => item.candidateRef === "Delta Co").canonicalRef; + // A caller-supplied metadata key cannot make a legacy entity pipeline-origin. + expect(await service.getEntity({ ref: gammaRef, scope: view })).toMatchObject({ canonicalRef: gammaRef }); + + // Relations written before the upgrade can still name a hidden entity. + const raw = new Database(path.join(directory, "gks.sqlite")); + try { + const insertRelation = raw.prepare(`INSERT INTO relations (canonical_ref, scope_key, from_ref, relation_type, to_ref, confidence, evidence_ref, portfolio_id, tenant_id, business_id, workspace_id, project_id, sharing, metadata_json, created_at, graph_version) + VALUES (?, ?, ?, 'RELATED_TO', ?, NULL, 'msp:proof/pre-upgrade', ?, ?, ?, ?, '', 'private', '{}', '2026-09-01T00:00:00.000Z', 'gks:graph/1')`); + const legacyKey = [view.portfolioId, view.tenantId, view.businessId, view.workspaceId, "", "private"].join("\u0000"); + insertRelation.run(`gks:relation/${"1".repeat(32)}`, legacyKey, gammaRef, hiddenRef, view.portfolioId, view.tenantId, view.businessId, view.workspaceId); + insertRelation.run(`gks:relation/${"2".repeat(32)}`, legacyKey, gammaRef, deltaRef, view.portfolioId, view.tenantId, view.businessId, view.workspaceId); + } finally { + raw.close(); + } + expect((await service.getRelations({ ref: gammaRef, scope: view })).map((relation) => relation.toRef)).toEqual([deltaRef]); + expect(await service.getRelations({ ref: hiddenRef, scope: view })).toEqual([]); + }); + + it("creates a legacy entity instead of conflicting when its norm key imitates a hidden typed pipeline key", async () => { + const { service } = harness(); + // A caseless semantic type keeps the typed key reachable by a norm_v1 + // string: normKey("acme\u0000123") === pipelineEntityNormKey("acme", "123"). + const { decision } = await submitAndClaim(service, makeBatch({ id: "batch-visibility-normkey", entries: [{ text: "Acme and Atlas", mentions: [["Acme", "acme", "123"], ["Atlas", "atlas", "Product"]] }] })); + const hiddenAcme = decision.entities.find((entity) => entity.metadata.resolutionKey === "acme"); + + const first = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "acme\u0000123", type: "ENTITY", title: "Acme imitation" }])); + const created = first.canonical_mappings[0]; + expect(created.resolution.outcome).toBe("CREATED"); + expect(created.canonicalRef).not.toBe(hiddenAcme.id); + expect(await legacyRead(service, scope(), hiddenAcme)).toEqual(hidden); + + // The same string later reaches the legacy entity, not a new split. + const again = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "acme\u0000123", type: "ENTITY", title: "Acme imitation" }])); + expect(again.canonical_mappings[0]).toMatchObject({ canonicalRef: created.canonicalRef, resolution: { outcome: "MATCHED" } }); + }); + + it("keeps a FAILED_STAGE run hidden", async () => { + const { service } = harness(); + const batch = makeBatch({ id: "batch-visibility-failed-stage" }); + const { decision } = await submitAndClaim(service, batch); + const worker = auth(batch.scope, "worker"); + await service.pipelineGraphReceipt({ receipt: graphReceiptFor(decision), ...worker }); + await service.pipelineStageFailure({ + schemaVersion: PIPELINE_SCHEMA_VERSION, scope: batch.scope, runId: batch.runId, decisionId: decision.decisionId, decisionHash: decision.decisionHash, + stage: decision.stages.find((stage) => stage.stageNumber === 15), startedAt: "2026-09-07T15:00:01.000Z", finishedAt: "2026-09-07T15:00:01.100Z", + metrics: metric({ records_in: batch.chunks.length, records_out: 0, error_count: 1 }), error: { code: "INDEX_WRITE_FAILED", message: "pinned Tier4 index rejected the candidate generation" }, ...worker, + }); + for (const entity of decision.entities) expect(await legacyRead(service, batch.scope, entity)).toEqual(hidden); + }); + + it("backfills origin from GKS's own records, never from caller metadata", async () => { + const { persistence, service, directory } = harness(); + // A legacy entity that GenesisRAG17 Stage 9 reuses (its metadata carries the + // typed identity) is named by pipeline_mentions but was legacy-created. + const reused = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "alice", type: "ENTITY", title: "Alice", metadata: { semanticType: "Person", resolutionKey: "alice" } }])); + const reusedRef = reused.canonical_mappings[0].canonicalRef; + const { decision } = await submitAndClaim(service, makeBatch({ id: "batch-visibility-backfill" })); + expect(decision.entities.find((entity) => entity.metadata.resolutionKey === "alice").id).toBe(reusedRef); + const pipelineCreated = decision.entities.filter((entity) => entity.id !== reusedRef); + const promoted = await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "Kilo Co", type: "ENTITY", title: "Kilo Co", metadata: { pipelineVersion: PIPELINE_SCHEMA_VERSION } }])); + const kiloRef = promoted.canonical_mappings[0].canonicalRef; + persistence.close(); + + // Rewind the store to before 0007: no origin column, no 0007 record. + const dbPath = path.join(directory, "gks.sqlite"); + const raw = new Database(dbPath); + try { + raw.exec("DROP INDEX idx_pipeline_mentions_entity_ref; ALTER TABLE entities DROP COLUMN origin; DELETE FROM schema_migrations WHERE name = '0007_pipeline_entity_origin.sql';"); + } finally { + raw.close(); + } + + const reopened = openSqlitePersistence({ dbPath }); + try { + const check = new Database(dbPath, { readonly: true }); + try { + const origins = Object.fromEntries(check.prepare("SELECT canonical_ref, origin FROM entities").all().map((row) => [row.canonical_ref, row.origin])); + for (const entity of pipelineCreated) expect(origins[entity.id]).toBe("pipeline"); + expect(origins[reusedRef]).toBe("legacy"); + expect(origins[kiloRef]).toBe("legacy"); + } finally { + check.close(); + } + expect(reopened.getEntity(pipelineCreated[0].id)).toBeNull(); + expect(reopened.getEntity(reusedRef)).toMatchObject({ canonicalRef: reusedRef }); + expect(reopened.getEntity(kiloRef)).toMatchObject({ canonicalRef: kiloRef }); + } finally { + reopened.close(); + } + // harness cleanup closes the original handle again; better-sqlite3 tolerates it. + }); +}); \ No newline at end of file diff --git a/tests/contract/stage9-migration.test.mjs b/tests/contract/stage9-migration.test.mjs index 7a0aa8d..6ab322b 100644 --- a/tests/contract/stage9-migration.test.mjs +++ b/tests/contract/stage9-migration.test.mjs @@ -12,7 +12,7 @@ // shape whose norm keys must collide during backfill. import { afterEach, describe, expect, it } from "vitest"; import Database from "better-sqlite3"; -import { mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { mkdtempSync, readFileSync, readdirSync, rmSync } from "node:fs"; import { tmpdir } from "node:os"; import path from "node:path"; import { scopeKey } from "@freshair129/gks-contracts"; @@ -20,6 +20,9 @@ import { createGksService } from "@freshair129/gks-core"; import { openSqlitePersistence } from "@freshair129/gks-persistence"; import { HASH_A, scope } from "../fixtures/candidates.mjs"; +// Every shipped migration is applied on upgrade; count them, not a constant. +const SHIPPED_MIGRATIONS = readdirSync("migrations").filter((name) => name.endsWith(".sql")).length; + const SCOPE_A = scope(); const SCOPE_B = scope({ tenantId: "tenant-b" }); const KEY_A = scopeKey(SCOPE_A); @@ -115,7 +118,7 @@ describe("migration 0002 on a populated pre-Stage-9 store", () => { // pipeline state). The 0005 hook writes one Stage 9 row per pre-existing // promotion, run_id NULL, so the seeded store's promotions are exportable // evidence the moment it is upgraded. - expect(raw.prepare("SELECT COUNT(*) AS n FROM schema_migrations").get().n).toBe(6); + expect(raw.prepare("SELECT COUNT(*) AS n FROM schema_migrations").get().n).toBe(SHIPPED_MIGRATIONS); expect(raw.prepare("SELECT COUNT(*) AS n FROM stage_evidence").get().n).toBe(raw.prepare("SELECT COUNT(*) AS n FROM promotions").get().n); expect(raw.prepare("SELECT COUNT(*) AS n FROM stage_evidence WHERE run_id IS NOT NULL").get().n).toBe(0); expect(raw.prepare("SELECT COUNT(*) AS n FROM entities").get().n).toBe(SEEDED.length); @@ -228,6 +231,6 @@ describe("migration 0002 on a populated pre-Stage-9 store", () => { cleanups.pop(); const raw = openRaw(dbPath); expect(raw.prepare("SELECT COUNT(*) AS n FROM entity_mentions").get().n).toBe(SEEDED.length); - expect(raw.prepare("SELECT COUNT(*) AS n FROM schema_migrations").get().n).toBe(6); + expect(raw.prepare("SELECT COUNT(*) AS n FROM schema_migrations").get().n).toBe(SHIPPED_MIGRATIONS); }); }); diff --git a/tests/security/cross-tenant-deny.security.mjs b/tests/security/cross-tenant-deny.security.mjs index 6250fbe..368ea2b 100644 --- a/tests/security/cross-tenant-deny.security.mjs +++ b/tests/security/cross-tenant-deny.security.mjs @@ -1,5 +1,6 @@ import assert from "node:assert/strict"; import test from "node:test"; +import Database from "better-sqlite3"; import { mkdtempSync, rmSync } from "node:fs"; import { tmpdir } from "node:os"; import path from "node:path"; @@ -478,3 +479,51 @@ test("stageEvidenceExport_foreignScopePagesToNothing_includingTheTenantlessCase" rmSync(dir, { recursive: true, force: true }); } }); + +// ADR-GKS-PIPELINE-VISIBILITY: an unpublished pipeline-origin entity answers +// every caller, own tenant or foreign, exactly as an absent ref does -- so a +// foreign tenant cannot tell "hidden" from "never existed", and the legacy +// resolver pool never offers it. +test("unpublishedPipelineEntity_isIndistinguishableFromAbsent_forEveryTenant", async () => { + const dir = mkdtempSync(path.join(tmpdir(), "gks-security-hidden-")); + const dbPath = path.join(dir, "gks.sqlite"); + const persistence = openSqlitePersistence({ dbPath }); + try { + const service = createGksService({ persistence }); + const tenantA = scope({ tenantId: "tenant-a", projectId: "" }); + const tenantB = scope({ tenantId: "tenant-b", projectId: "" }); + const hiddenRef = `gks:entity/hidden-${"a".repeat(32)}`; + const absentRef = `gks:entity/absent-${"b".repeat(32)}`; + const raw = new Database(dbPath); + try { + raw.prepare(`INSERT INTO entities (canonical_ref, scope_key, candidate_ref, type, title, summary, source_ref, confidence, portfolio_id, tenant_id, business_id, workspace_id, project_id, sharing, metadata_json, aliases_json, external_refs_json, norm_key, norm_version, created_at, updated_at, graph_version, origin) + VALUES (?, ?, 'hidden', 'Person', 'Hidden Person', '', 'source-1', NULL, ?, 'tenant-a', ?, ?, '', 'private', '{}', '[]', '[]', 'hidden', 'norm_v1', '2026-09-01T00:00:00.000Z', '2026-09-01T00:00:00.000Z', 'gks:graph/1', 'pipeline')`) + .run(hiddenRef, [tenantA.portfolioId, "tenant-a", tenantA.businessId, tenantA.workspaceId, "", "private"].join("\u0000"), tenantA.portfolioId, tenantA.businessId, tenantA.workspaceId); + } finally { + raw.close(); + } + + for (const caller of [tenantA, tenantB]) { + assert.equal(await service.getEntity({ ref: hiddenRef, scope: caller }), await service.getEntity({ ref: absentRef, scope: caller })); + assert.deepEqual(await service.search({ query: "Hidden Person", scope: caller }), []); + assert.deepEqual(await service.getRelations({ ref: hiddenRef, scope: caller }), await service.getRelations({ ref: absentRef, scope: caller })); + const link = (knowledgeRef) => service.linkArtifact({ knowledgeRef, artifactRef: "project:PRJ-HIDDEN", relationType: "RELATED_TO", evidenceRef: "msp:proof/hidden-link", scope: caller }); + const [hiddenError, absentError] = await Promise.all([link(hiddenRef).catch((error) => error), link(absentRef).catch((error) => error)]); + assert.equal(hiddenError.code, absentError.code); + assert.equal(hiddenError.message, absentError.message); + // resolveTo cannot probe it either: REJECTED, exactly like an absent ref. + const probe = await service.promoteCandidate(promotion({ + idempotency_key: `hidden-probe-${caller.tenantId}`, + scope: caller, + candidate: { entities: [{ candidateRef: "probe", type: "ENTITY", title: "Probe", resolveTo: hiddenRef }], relations: [] }, + })); + assert.equal(probe.canonical_mappings[0].resolution.outcome, "REJECTED"); + } + assert.deepEqual(persistence.lookupResolutionCandidates({ scope: tenantA }).map((row) => row.canonicalRef), []); + assert.deepEqual(persistence.lookupResolutionCandidates({ scope: tenantA, includeUnpublishedPipeline: true }).map((row) => row.canonicalRef), [hiddenRef]); + assert.deepEqual(persistence.lookupResolutionCandidates({ scope: tenantB, includeUnpublishedPipeline: true }), []); + } finally { + persistence.close(); + rmSync(dir, { recursive: true, force: true }); + } +}); \ No newline at end of file From 2b501868c7e977abcae48ff76fab27f8446e652d Mon Sep 17 00:00:00 2001 From: Freshair129 <94353529+Freshair129@users.noreply.github.com> Date: Sun, 27 Sep 2026 09:55:49 +0700 Subject: [PATCH 2/2] fix: refuse to supersede a pipeline entity; accept visibility ADR - The owner approved ADR-GKS-PIPELINE-VISIBILITY (0.2.0, accepted 2026-09-27). - D4 is now enforced in code: a D9 MERGE whose supersededRef is a pipeline-origin entity is refused with gks_conflict, because Stage 9 reuse finds pipeline entities by deterministic id and does not follow supersession. A pipeline entity can still be the survivor. Covered by a test and mutation-checked. - D3 notes that after publication an imitation string can resolve AMBIGUOUS. - Revision rows in the data model, port contract, C0 qualification and GenesisRAG17 ADRs now reference the accepted decision. Co-Authored-By: Claude Opus 5.5 --- docs/ADR-GKS-C0-QUALIFICATION.md | 2 +- docs/ADR-GKS-GENESISRAG17.md | 2 +- docs/ADR-GKS-PIPELINE-VISIBILITY.md | 41 ++++++++++++------- docs/GKS-DATA-MODEL.md | 2 +- docs/GKS-PORT-CONTRACT.md | 2 +- packages/gks-persistence/src/index.mjs | 6 +++ tests/contract/pipeline-genesisrag17.test.mjs | 18 ++++++++ 7 files changed, 55 insertions(+), 18 deletions(-) diff --git a/docs/ADR-GKS-C0-QUALIFICATION.md b/docs/ADR-GKS-C0-QUALIFICATION.md index 371c0bb..9ef3b24 100644 --- a/docs/ADR-GKS-C0-QUALIFICATION.md +++ b/docs/ADR-GKS-C0-QUALIFICATION.md @@ -216,4 +216,4 @@ review confirms: | 0.1.0b | 2026-09-22 | candidate | Proposed C0 qualification decisions for G1 review | working-tree | RWANG | | 0.1.0 | 2026-09-22 | beta | User approved D1-D4 for C0 implementation | working-tree | RWANG | | 0.2.0 | 2026-09-22 | beta | Clarified compatibility versus secure MSP auth mode and recorded real-chain qualification boundary | working-tree | RWANG | -| 0.2.1 | 2026-09-27 | beta | Cross-reference only: ADR-GKS-PIPELINE-VISIBILITY (proposed) changes what `gks_search`, `gks_entity_get`, `gks_relations_get` and `gks_artifact_link` return for unpublished GenesisRAG17 entities. No request is newly rejected; see that ADR's observable-changes table. | working-tree | Claude | +| 0.2.1 | 2026-09-27 | beta | Cross-reference only: ADR-GKS-PIPELINE-VISIBILITY (accepted 2026-09-27) changes what `gks_search`, `gks_entity_get`, `gks_relations_get` and `gks_artifact_link` return for unpublished GenesisRAG17 entities. No request is newly rejected; see that ADR's observable-changes table. | working-tree | Claude | diff --git a/docs/ADR-GKS-GENESISRAG17.md b/docs/ADR-GKS-GENESISRAG17.md index 7480ceb..3acac54 100644 --- a/docs/ADR-GKS-GENESISRAG17.md +++ b/docs/ADR-GKS-GENESISRAG17.md @@ -221,7 +221,7 @@ evidence. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| -| 0.5.1b | 2026-09-27 | accepted | Cross-reference: ADR-GKS-PIPELINE-VISIBILITY (proposed) keeps submit-time canonical entities invisible to legacy reads until a mentioning run is PUBLISHED; pipeline tools and Stage 9 reuse are unchanged. | working-tree | Claude Opus 5.5 | +| 0.5.1b | 2026-09-27 | accepted | Cross-reference: ADR-GKS-PIPELINE-VISIBILITY (accepted 2026-09-27) keeps submit-time canonical entities invisible to legacy reads until a mentioning run is PUBLISHED; pipeline tools and Stage 9 reuse are unchanged. | working-tree | Claude Opus 5.5 | | 0.5.0b | 2026-09-11 | accepted | Implemented structured-record profile contract revision 2 on the GKS side (ADR-075 Phase 2, rollout step 2 of 3): `PIPELINE_ONTOLOGY_VERSION` is `ontology_v2`, `PIPELINE_SUPPORTED_ONTOLOGY_VERSIONS` is `{ontology_v1, ontology_v2}`, the Stage 11 `validEndpoint` ternary is replaced by the per-version predicate -> endpoint table `PIPELINE_ONTOLOGY_ENDPOINTS` (v2 adds `HAS_COMPONENT`, `PRICED_AT`, `IN_CATEGORY` and the `PACKAGE`/`CATEGORY`/`PRICE_TIER` endpoint types), and the Stage 17 knowledge dimension checks set membership and each fact against its own version instead of equality with `ontology_v1`. `rule_v1`, `parseStructuredClaim`, Stage 12 and the C-10 bitemporal count are unchanged. Must merge after the GenesisBlock worker change that accepts both versions. | working-tree | Claude Opus 5 | | 0.4.4b | 2026-09-11 | accepted | Fixed a residual of contract item C-10: the Stage 17 expected bitemporal lane count now includes HELD rows that carry valid time, not only facts, using the same row classification the worker applies (`temporalRows()` covers facts and held together; a row is dated unless validFrom/validTo are undefined/null/`not_applicable` and status is undefined/`not_applicable`). A decision with a dated held row previously disagreed with the worker's count even though the 0.4.3b fix already matched on facts alone. The all-`not_applicable` path (expect 0, lane `not_applicable`) now also considers facts and held together. Held rows already make the knowledge dimension WARN; this only corrects the graph-dimension reason, never whether anything publishes. | working-tree | Claude Opus 5 | | 0.4.3b | 2026-09-11 | accepted | Fixed contract item C-10: the Stage 17 expected bitemporal lane count is the number of facts that carry valid time, not `facts.length`, matching the GenesisBlock worker's mapped-only count. A generation mixing dated and `not_applicable` facts previously failed the graph dimension on the lane-count comparison. The all-`not_applicable` path (expect 0, lane `not_applicable`) and all-dated generations are unchanged. | working-tree | Claude Opus 5 | diff --git a/docs/ADR-GKS-PIPELINE-VISIBILITY.md b/docs/ADR-GKS-PIPELINE-VISIBILITY.md index 1a45c2b..76a3f2f 100644 --- a/docs/ADR-GKS-PIPELINE-VISIBILITY.md +++ b/docs/ADR-GKS-PIPELINE-VISIBILITY.md @@ -1,10 +1,10 @@ --- -version: "0.1.0" +version: "0.2.0" created_at: "2026-09-27T10:00:00+07:00,Claude,working-tree" -last_update: "2026-09-27T10:00:00+07:00,Claude" -status: "proposed" +last_update: "2026-09-27T09:55:01+07:00,Claude" +status: "accepted" approval_owner: "Boss (บอส)" -approval_recorded_at: null +approval_recorded_at: "2026-09-27T09:55:01+07:00" superseded_by: null attributes: domain: "genesis-knowledge-system" @@ -16,12 +16,11 @@ attributes: ## Decision status -**Proposed, awaiting owner approval.** The implementation lands in the same -pull request so the decision can be judged against working code and tests, -but it must not merge before approval is recorded here. It changes the -observable read behaviour of four frozen C0 tools, and -[`ADR-GKS-C0-QUALIFICATION.md`](ADR-GKS-C0-QUALIFICATION.md) does not -authorize that on its own. +**Accepted.** The owner approved this decision on 2026-09-27, including enforcing +the merge-survivor rule (D4) in code. It changes the observable read behaviour +of four frozen C0 tools, which +[`ADR-GKS-C0-QUALIFICATION.md`](ADR-GKS-C0-QUALIFICATION.md) does not authorize +on its own; this ADR is that authorization. ## Context @@ -100,7 +99,8 @@ over-split for never serving unpublished knowledge. GenesisRAG17 Stage 9 reuse does not follow supersession: a superseded pipeline entity drops out of both pools, and a later run would reuse the superseded row by its deterministic id. That gap predates this ADR. Until -reuse follows supersession, a repair must not supersede a pipeline entity. +reuse follows supersession, D4 refuses any merge that would supersede a +pipeline entity. **Norm-key collisions.** Usually the typed pipeline key (`norm_v1(resolutionKey) + U+0000 + TYPE`) and a legacy `norm_v1` key differ. @@ -117,7 +117,10 @@ surface `gks_conflict`, which would reveal the hidden row. It creates the legacy entity under the existing D2 human-distinct discriminator (`norm_key#mention_id`). Later promotes of the same string reach that entity through the EXACT rung, which compares candidate strings, so the split does -not repeat. A collision with a *visible* row keeps the decision-5 retry. +not repeat. A collision with a *visible* row keeps the decision-5 retry. Once +the pipeline entity is published, a later promote of the imitation string can +match both rows and resolve `AMBIGUOUS`. That only happens for a string that +deliberately copies a caseless pipeline type, and an ambiguous result is safe. ### D4 — D9 BIND/MERGE operate only on visible entities @@ -126,6 +129,11 @@ not repeat. A collision with a *visible* row keeps the decision-5 retry. existing "does not resolve to a canonical entity" error. It is never `gks_scope_denied`, which would confirm that the entity exists. +A `MERGE` whose `supersededRef` is a pipeline-origin entity is refused with +`gks_conflict`, even after publication. The pipeline entity may be the survivor. +It may not be superseded, because Stage 9 reuse would pick the superseded row up +again (D3). This is the one newly rejected request in this ADR. + ### D5 — Relations never expose a hidden endpoint `getRelations` omits any relation whose `from_ref` or `to_ref` names a hidden @@ -163,10 +171,12 @@ For an entity that exists only through unpublished runs: | `gks_artifact_link` to it | link written | `gks_invalid_request` ("does not resolve") | | `gks_knowledge_promote` | may resolve `MATCHED` / `resolveTo` onto it | never resolves onto it; may create a separate entity (D3) | | D9 BIND/MERGE naming it | accepted | refused as not resolving | +| D9 MERGE superseding a *published* pipeline entity | accepted | `gks_conflict`; merge the other entity into it instead | | Earlier promote snapshots and `gks_stage_evidence_export` rows that name it | ref resolved | the ref is kept unchanged (D6), but reads it as `null` until publication | -No accepted request shape is rejected at validation, so no versioned wire -rollout is needed. The change is recorded here and in the C0 qualification +No request shape is rejected at validation. The only newly refused request is +the pipeline-loser `MERGE` above, a D9 write that already has conflict outcomes, +so no versioned wire rollout is needed. The change is recorded here and in the C0 qualification ADR's revision history instead. ## Rollback @@ -206,6 +216,8 @@ ADR's revision history instead. refused, and a repeat of that string matches the entity it created. - A `FAILED_STAGE` run stays hidden. - D9 BIND/MERGE refuse hidden refs with the not-resolving error. +- D9 MERGE refuses to supersede a pipeline-origin entity and accepts it as the + survivor. - Relations touching a hidden entity are omitted. - A foreign-tenant caller learns nothing about a hidden entity. - The backfill marks pre-existing rows exactly as D2 states. @@ -214,4 +226,5 @@ ADR's revision history instead. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| +| 0.2.0 | 2026-09-27 | accepted | Owner approved. D4 now refuses a MERGE that would supersede a pipeline-origin entity (enforced in code and tested). D3 notes the post-publication AMBIGUOUS case for imitation strings. | working-tree | Claude | | 0.1.0 | 2026-09-27 | proposed | Proposed hiding unpublished GenesisRAG17 entities from legacy reads through a GKS-owned `origin` column, with consistent resolution-pool, D9 and relation rules. | working-tree | Claude | diff --git a/docs/GKS-DATA-MODEL.md b/docs/GKS-DATA-MODEL.md index bf6a591..3a650d0 100644 --- a/docs/GKS-DATA-MODEL.md +++ b/docs/GKS-DATA-MODEL.md @@ -504,7 +504,7 @@ tables and write rules. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| -| 0.8.0b | 2026-09-27 | beta | Proposed `entities.origin` (migration 0007) and the publication-visibility rule for GenesisRAG17 entities in legacy reads (ADR-GKS-PIPELINE-VISIBILITY). | working-tree | Claude | +| 0.8.0b | 2026-09-27 | beta | Added `entities.origin` (migration 0007) and the publication-visibility rule for GenesisRAG17 entities in legacy reads (ADR-GKS-PIPELINE-VISIBILITY, accepted 2026-09-27). | working-tree | Claude | | 0.7.0b | 2026-09-08 | beta | Recorded the GenesisRAG17 typed Stage 9 entity key and semantic type storage in the shared entities table while preserving every pipeline mention occurrence. | working-tree | RWANG | | 0.6.0b | 2026-09-08 | beta | Expanded the GenesisRAG17 data model with exact migration 0006 table shapes, immutable decision facts/occurrences, receipt/gate snapshots, cursor ordering, and the no-Stage-18 extension boundary. | 9279cfe | RWANG | | 0.5.0b | 2026-09-07 | beta | Clarified the graph-receipt-to-enrichment boundary, post-acknowledgement physical projections, gate statistics and Tier4 failure-only terminal rows. | working-tree | RWANG | diff --git a/docs/GKS-PORT-CONTRACT.md b/docs/GKS-PORT-CONTRACT.md index 9321749..986ab61 100644 --- a/docs/GKS-PORT-CONTRACT.md +++ b/docs/GKS-PORT-CONTRACT.md @@ -541,7 +541,7 @@ implementation package name appears in the client. | Version | Date | Status | Summary | Commit Hash | Agent | |---|---|---|---|---|---| -| 0.10.1b | 2026-09-27 | beta | Proposed (ADR-GKS-PIPELINE-VISIBILITY): legacy reads, the legacy resolution pool and D9 operands exclude unpublished GenesisRAG17 entities; `lookupResolutionCandidates` gains `includeUnpublishedPipeline` for Stage 9 reuse. No tool request or result shape changes. | working-tree | Claude | +| 0.10.1b | 2026-09-27 | beta | Per ADR-GKS-PIPELINE-VISIBILITY (accepted 2026-09-27): legacy reads, the legacy resolution pool and D9 operands exclude unpublished GenesisRAG17 entities; `lookupResolutionCandidates` gains `includeUnpublishedPipeline` for Stage 9 reuse; D9 MERGE refuses to supersede a pipeline-origin entity (`gks_conflict`). No tool request or result shape changes. | working-tree | Claude | | 0.10.0b | 2026-09-24 | beta | Implements optional per-client hash-backed direct HTTP read grants while preserving the MSP profile; grants are not enabled by default and production rollout remains separate. | working-tree | RWANG | | 0.9.0b | 2026-09-24 | beta | Adds the approved direct-client read-only grant profile while preserving MSP auth for governed writes; concrete identity verification and activation remain unimplemented. | working-tree | RWANG | | 0.8.1b | 2026-09-22 | beta | Removed the stale port-version-1 statement that contradicted the selected GKS-owned SQLite production profile; deployment evidence remains separately gated. | working-tree | RWANG | diff --git a/packages/gks-persistence/src/index.mjs b/packages/gks-persistence/src/index.mjs index 003712c..a5f1693 100644 --- a/packages/gks-persistence/src/index.mjs +++ b/packages/gks-persistence/src/index.mjs @@ -1015,6 +1015,12 @@ export function openSqlitePersistence({ dbPath, migrationsDir = DEFAULT_MIGRATIO if (loser.superseded_by !== null) { throw new GksConflictError("supersededRef names an entity that is already superseded."); } + // ADR-GKS-PIPELINE-VISIBILITY D4: GenesisRAG17 Stage 9 reuse finds its + // entities by deterministic id and does not follow supersession, so a + // pipeline-origin entity may only survive a merge, never be superseded. + if (loser.origin === "pipeline") { + throw new GksConflictError("supersededRef names a GenesisRAG17 entity; merge the other entity into it instead."); + } const graphVersion = `gks:graph/${nextVersion.get().version}`; markSuperseded.run({ canonical_ref: loser.canonical_ref, superseded_by: survivor.canonical_ref, updated_at: now, graph_version: graphVersion }); // The survivor absorbs the loser's identity evidence: its base norm key diff --git a/tests/contract/pipeline-genesisrag17.test.mjs b/tests/contract/pipeline-genesisrag17.test.mjs index 8d8badf..2154f3a 100644 --- a/tests/contract/pipeline-genesisrag17.test.mjs +++ b/tests/contract/pipeline-genesisrag17.test.mjs @@ -769,6 +769,24 @@ describe("GenesisRAG17 entities before publication (ADR-GKS-PIPELINE-VISIBILITY) expect(again.canonical_mappings[0]).toMatchObject({ canonicalRef: created.canonicalRef, resolution: { outcome: "MATCHED" } }); }); + it("lets a published pipeline entity survive a repair merge but never be superseded", async () => { + const { service } = harness(); + const view = legacyView(scope()); + // The legacy spelling is promoted first, under a type the pipeline reuse + // cannot claim, so publication leaves two identities to repair (D3). + const duplicateRef = (await service.promoteCandidate(legacyPromotion(scope(), [{ candidateRef: "alice legacy", type: "ENTITY", title: "Alice" }]))).canonical_mappings[0].canonicalRef; + const decision = await publishRun(service, makeBatch({ id: "batch-visibility-merge" })); + const aliceRef = decision.entities.find((entity) => entity.metadata.resolutionKey === "alice").id; + expect((await service.getEntity({ ref: aliceRef, scope: view })).canonicalRef).toBe(aliceRef); + + await expect(service.applyHumanResolution({ action: "MERGE", survivorRef: duplicateRef, supersededRef: aliceRef, provenanceRef: "msp:proof/merge-pipeline-away", scope: view })) + .rejects.toMatchObject({ code: "gks_conflict", message: expect.stringContaining("GenesisRAG17") }); + expect((await service.getEntity({ ref: aliceRef, scope: view })).supersededBy ?? null).toBeNull(); + + await service.applyHumanResolution({ action: "MERGE", survivorRef: aliceRef, supersededRef: duplicateRef, provenanceRef: "msp:proof/merge-into-pipeline", scope: view }); + expect((await service.getEntity({ ref: duplicateRef, scope: view })).supersededBy).toBe(aliceRef); + }); + it("keeps a FAILED_STAGE run hidden", async () => { const { service } = harness(); const batch = makeBatch({ id: "batch-visibility-failed-stage" });