Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 13 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,16 @@ jobs:
- name: Lint all targets
run: cargo clippy --all-targets --all-features --locked -- -D warnings

- name: Test all targets
run: cargo test --all-targets --locked
- name: Test all targets and verify acceptance coverage
run: python3 tools/run_acceptance_scenarios.py --all-targets

- name: Retain normalized acceptance evidence
uses: actions/upload-artifact@v4
with:
name: acceptance-scenarios-${{ matrix.toolchain }}
path: target/acceptance-scenarios.v1.report.json
if-no-files-found: error
retention-days: 7

- name: Test documentation
run: cargo test --doc --locked
Expand Down Expand Up @@ -82,6 +90,9 @@ jobs:
- name: Check deterministic scenario sources
run: python3 tools/generate_scenario_sources.py --check

- name: Reject incomplete acceptance reports
run: python3 tools/test_acceptance_report.py

- name: Validate architecture specifications
run: python3 .agents/specs/validate-specs.py

Expand Down
5 changes: 5 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,11 @@ nix = { version = "0.30.1", default-features = false, features = ["process", "si

[dev-dependencies]

# Conformance tests copy and hash the exact executable repeatedly. Debug symbols
# are not part of the contract; omit them to keep the PR fixture budget bounded.
[profile.test]
debug = 0

[lints.rust]
unsafe_code = "forbid"

Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,7 @@ or a sandbox.
- [Process authority and isolation](docs/integrations/authority-isolation.md)
- [Artifact binding contract](docs/integrations/artifact-bindings.md)
- [Scenario fixture contract](docs/integrations/scenario-fixtures.md)
- [Executable acceptance matrix](docs/integrations/acceptance-scenarios.md)
- [Hermetic provider kit](docs/integrations/hermetic-provider-kit.md)
- [Versioned contracts](contracts/README.md)
- [Roadmap](ROADMAP.md)
Expand Down
8 changes: 8 additions & 0 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,14 @@ PR #60 merged on 2026-09-24 as
completing the hermetic provider-kit closeout. The exact next Flow checkpoint is
[#30](https://github.com/egohygiene/flow/issues/30).

The #30 review candidate adds the
[executable acceptance matrix](docs/integrations/acceptance-scenarios.md):
versioned deterministic recipes, exact typed outcomes, two fresh-root receipts
per case, PR budgets, and machine-readable coverage/gaps. Its scope ends at
single-execution acceptance and refusal. After this candidate merges and the
default-branch gate passes, #49 is next. FLO-Q03 remains active until the later
real released-provider adapters satisfy its two-adapter exit criterion.

### Central Flow chain

1. [#30](https://github.com/egohygiene/flow/issues/30) — prove compatibility,
Expand Down
93 changes: 93 additions & 0 deletions docs/integrations/acceptance-scenarios.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# Acceptance scenario matrix

## Design and scope

Flow #30 proves one provider execution boundary using the closed scenario
contract and the finalized hermetic provider kit. Test recipes own fault
injection; production code continues to own resolution, process preflight,
transcript validation, filesystem observation, and final artifact acceptance.
No new scheduler, provider algorithm, or recovery state is introduced.

The versioned, test-owned catalog in
`tests/fixtures/acceptance-scenarios.v1.json` assigns stable scenario IDs,
deterministic recipes, expected typed outcomes, PR-tier budgets, and explicit
coverage gaps. Its executor lives under `tests/scenario_matrix/` and reuses the
provider kit's exact package finalization and execution helpers. Catalog rows
are conformance recipes, not public runtime plans or authorization.

Each row must produce exactly one normalized receipt from an actual public API
outcome. An unexpected error, missing row, duplicate row, or expected-outcome
mismatch fails the suite. Raw error strings and provider-authored free text
are excluded from receipts. Recovery instructions name the failed boundary.
The harness compares receipts across fresh roots and checks canaries before
printing portable JSON evidence.

## Boundaries under test

- Resolution: compatibility, missing or unavailable observations, stale and
future versions, exact pins, deterministic conflict, and pre-invocation
fallback policy.
- Contract and preflight: schema rejection, identity, configuration,
authorization, executable integrity, and malformed provider evidence.
- Artifact observation and acceptance: complete, partial, missing, extra,
changed, contradictory, corrupt, escaped, linked, and incorrectly typed
outputs; exit zero cannot authorize promotion.
- Evidence: distinct empty, unavailable, incomplete, unsupported, invalid,
and failed states; bounded operational evidence and portable privacy.

## Running and interpreting the matrix

Run after populating Cargo's locked dependency cache:

```console
python3 tools/run_acceptance_scenarios.py
python3 tools/run_acceptance_scenarios.py --all-targets
```

The driver uses `cargo test --locked --offline` and writes
`target/acceptance-scenarios.v1.report.json`. It removes an older report before
starting, fails on any failed Rust test or missing/duplicate/unexpected receipt,
and records the catalog digest, tested source-file digests, Cargo version,
coverage counts, residual gaps, and actual normalized receipts. Failure logs
stay local; CI uploads only the successful portable report. The source digest
is checked before and after execution so an edited source tree cannot be
mistaken for the tested one.

The PR catalog budgets 30 seconds per execution, two executions per case,
16 KiB per receipt, four immediate output entries, and 1 MiB of generated
regular-file output. The fixed recipes create only files and empty directories;
this output-count budget is not a general recursive filesystem quota. The
driver has a 600-second Cargo wait timeout and a 4 MiB post-capture test-log
acceptance limit; it is not a general descendant-process supervisor. Provider
stdout/stderr and process deadlines remain enforced by the existing locked kit
limits. Budget measurements exclude compilation from the per-case time, but
include compilation in the driver timeout. Memory, whole-filesystem limits,
and kernel network isolation remain explicit gaps.

The test profile omits debug symbols because the exact provider executable is
copied and hashed repeatedly. This keeps generated package size and PR time
bounded without bypassing any package or executable digest check. Each receipt
still names the actual build-specific bytes. The `argv-environment` recipe also
uses a tiny first-party shell probe; its exact observed subjects are retained
in that receipt's `probe_subjects` field.

`observed-empty` means Flow actually observed an empty directory; its receipt
does not claim an accepted provider execution. Resolution success similarly
does not claim execution or artifact acceptance. Only receipts with
`accepted: true` crossed `accept_artifacts` successfully. Transcript mutations
start from actual kit execution and re-enter the public host-neutral transcript
validator; the receipt identifies the validator outcome, not a second launch.

## Claim limits

The kit is synthetic. Real released providers, native format validators,
cryptographic publisher authentication, OS sandboxing, atomic filesystem
snapshots, descriptor-bound launch, and descendant containment are not proven.
Flow #49 owns durable state; #31 owns retry and resume. The scenario catalog
does not expand those checkpoints or claim that two synthetic fixtures are
two real provider adapters.

Operational error objects and rejected provider evidence remain host-local
diagnostics and may contain private values. Only the explicitly allowlisted
normalized receipts are portable. Privacy canaries do not establish a generic
redactor for arbitrary provider-authored messages.
4 changes: 4 additions & 0 deletions docs/integrations/hermetic-provider-kit.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,9 @@
# Hermetic orchestration provider kit

The issue #30 [acceptance matrix](acceptance-scenarios.md) consumes this kit
through test-owned recipes. It adds executable coverage receipts and explicit
gaps while retaining the package, process, and artifact boundaries below.

## Purpose and checkpoint boundary

The hermetic provider kit is Flow-owned conformance infrastructure for issue
Expand Down
8 changes: 8 additions & 0 deletions docs/integrations/scenario-fixtures.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,6 +159,14 @@ contract rather than an executor.

## Checked-in corpus

Issue #30 adds an [executable acceptance matrix](acceptance-scenarios.md) on
top of this contract. Its test-owned catalog reuses the terminal/evidence
vocabulary and public provider-kit boundaries, with two fresh-root executions
per recipe and a completeness-checked JSON report. It does not interpret an
arbitrary scenario manifest as an executable plan. Illustrative digests in the
original intent fixtures are never presented as the executed package identity;
runtime receipts retain the actual finalized kit digests.

The corpus includes a single-provider success example plus multi-provider,
interrupted, observed-empty, and expected-unavailable fixtures. The unavailable
fixture is a valid negative scenario; it is distinct from malformed manifest
Expand Down
Loading
Loading