Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -43,14 +43,16 @@ jobs:
- name: Lint all targets
run: cargo clippy --all-targets --all-features --locked -- -D warnings

- name: Test all targets and verify acceptance coverage
- name: Test all targets and verify acceptance and lifecycle coverage
run: python3 tools/run_acceptance_scenarios.py --all-targets

- name: Retain normalized acceptance evidence
- name: Retain normalized acceptance and lifecycle evidence
uses: actions/upload-artifact@v4
with:
name: acceptance-scenarios-${{ matrix.toolchain }}
path: target/acceptance-scenarios.v1.report.json
path: |
target/acceptance-scenarios.v1.report.json
target/lifecycle-scenarios.v1.report.json
if-no-files-found: error
retention-days: 7

Expand Down Expand Up @@ -93,6 +95,9 @@ jobs:
- name: Reject incomplete acceptance reports
run: python3 tools/test_acceptance_report.py

- name: Reject incomplete lifecycle reports and verify test resource limits
run: python3 tools/test_lifecycle_report.py

- name: Check durable contract references and closed shapes
run: python3 tools/test_durable_contracts.py

Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,7 @@ or a sandbox.
- [Artifact binding contract](docs/integrations/artifact-bindings.md)
- [Scenario fixture contract](docs/integrations/scenario-fixtures.md)
- [Executable acceptance matrix](docs/integrations/acceptance-scenarios.md)
- [Durable lifecycle scenarios and recovery evidence](docs/integrations/lifecycle-scenarios.md)
- [Hermetic provider kit](docs/integrations/hermetic-provider-kit.md)
- [Versioned contracts](contracts/README.md)
- [Roadmap](ROADMAP.md)
Expand Down
16 changes: 11 additions & 5 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ schema: aether.architecture-document/v1
id: flow-roadmap
title: Flow Roadmap
kind: architecture-document
version: 1.3.3
version: 1.3.4
status: draft
owners:
- egohygiene
Expand Down Expand Up @@ -43,9 +43,14 @@ It adds [durable prepared execution](docs/integrations/durable-state.md):
versioned plans and state, atomic immutable snapshots, workspace locking,
accepted checkpoints, fresh resume eligibility, and explicit recovery decisions.
#31 is active through three dependency-ordered review checkpoints:
[#64](https://github.com/egohygiene/flow/issues/64) adds fresh graph assessment and
safe dependent execution; [#65](https://github.com/egohygiene/flow/issues/65) adds
the deterministic durable lifecycle corpus; [#66](https://github.com/egohygiene/flow/issues/66)
[#64](https://github.com/egohygiene/flow/issues/64) merged through
[PR #67](https://github.com/egohygiene/flow/pull/67) as
`83f1ea161aa5aba03da5de6b287005252000cd4b`, with green
[default-branch CI](https://github.com/egohygiene/flow/actions/runs/36226457853).
It adds fresh graph assessment and safe dependent execution.
[#65](https://github.com/egohygiene/flow/issues/65) adds the
[deterministic durable lifecycle corpus](docs/integrations/lifecycle-scenarios.md)
with 34 recipes executed twice, fresh-process restarts, and bounded receipts; [#66](https://github.com/egohygiene/flow/issues/66)
proves authority, duplicate-effect prevention, and recovery residuals. Land one
review PR and verify default-branch CI before starting the next checkpoint.
FLO-Q03 remains active until real released-provider adapters satisfy its
Expand Down Expand Up @@ -104,7 +109,8 @@ Renderflow #421 → #422–#430 → Flow #10 exhaustive comic workflow
```

The suite therefore has four useful parallel ready fronts:
Flow #30, Renderflow #415, Optiflow #88, and Aniflow #8. Keep provider domain
Flow #65 (then #66 after merge and green default-branch CI), Renderflow #415,
Optiflow #88, and Aniflow #8. Keep provider domain
logic in the provider repositories and consume only immutable public contracts
from Flow.

Expand Down
13 changes: 11 additions & 2 deletions docs/architecture/foundation/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ schema: aether.architecture-document/v1
id: flow-architecture
title: Flow Architecture
kind: architecture-document
version: 0.13.0
version: 0.13.1
status: draft
owners:
- egohygiene
Expand Down Expand Up @@ -133,6 +133,13 @@ the assessment, and single-step entry points refuse dependent steps. Automatic
scheduling, retry, cross-plan migration, and provider-native checkpoints remain
later orchestration work.

The test-owned lifecycle corpus exercises these same public durable APIs with
synthetic providers, immutable transition projections, fresh-process reopening,
and explicit recovery decisions. Its versioned receipts are conformance evidence,
not executable plans, authorization tokens, or a second runtime state model.
Portable reports contain allowlisted identities and typed outcomes; raw operational
evidence stays local. The lifecycle guide owns recipe coverage and resource limits.

### External adapters

Adapters isolate process invocation, version/capability probing, structured
Expand Down Expand Up @@ -281,7 +288,9 @@ production graph scheduler. Issue #49 adds the versioned durable coordinator,
immutable prepared plans, atomic snapshots, accepted checkpoints, fresh reuse
assessment, and explicit recovery decisions described above. Issue #64 requires
fresh prerequisite evidence for graph assessment and dependent execution without
adding a scheduler or rewriting accepted history. Flow does not yet supply
adding a scheduler or rewriting accepted history. Issue #65 adds the bounded
[durable lifecycle corpus](../../integrations/lifecycle-scenarios.md), including
fresh-process restart and deterministic recovery receipts. Flow does not yet supply
the public CLI, real holon adapters, signature or transparency verification,
provider-native artifact validation, an atomic filesystem snapshot, an
operating-system sandbox or authenticated enforcement evidence,
Expand Down
18 changes: 12 additions & 6 deletions docs/integrations/acceptance-scenarios.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,8 @@ python3 tools/run_acceptance_scenarios.py
python3 tools/run_acceptance_scenarios.py --all-targets
```

The driver uses `cargo test --locked --offline` and writes
The Linux driver uses `cargo test --locked --offline` with one Rust test worker
to isolate fork-inherited test locks (explicit contention tests still run), and writes
`target/acceptance-scenarios.v1.report.json`. It removes an older report before
starting, fails on any failed Rust test or missing/duplicate/unexpected receipt,
and records the catalog digest, tested source-file digests, Cargo version,
Expand All @@ -57,12 +58,15 @@ The PR catalog budgets 30 seconds per execution, two executions per case,
16 KiB per receipt, four immediate output entries, and 1 MiB of generated
regular-file output. The fixed recipes create only files and empty directories;
this output-count budget is not a general recursive filesystem quota. The
driver has a 600-second Cargo wait timeout and a 4 MiB post-capture test-log
acceptance limit; it is not a general descendant-process supervisor. Provider
Linux driver enforces a 600-second command deadline and a 4 MiB output limit
while capturing. A Cargo target runner limits each test/provider process to
4 GiB address space and each written file to 16 MiB; compilation is excluded
from these kernel limits. Failed or timed-out test commands have their process
group killed and reaped. This is test supervision, not production containment. Provider
stdout/stderr and process deadlines remain enforced by the existing locked kit
limits. Budget measurements exclude compilation from the per-case time, but
include compilation in the driver timeout. Memory, whole-filesystem limits,
and kernel network isolation remain explicit gaps.
include compilation in the driver timeout. Aggregate process-tree memory,
whole-filesystem quotas, and kernel network isolation remain explicit gaps.

The test profile omits debug symbols because the exact provider executable is
copied and hashed repeatedly. This keeps generated package size and PR time
Expand All @@ -83,7 +87,9 @@ validator; the receipt identifies the validator outcome, not a second launch.
The kit is synthetic. Real released providers, native format validators,
cryptographic publisher authentication, OS sandboxing, atomic filesystem
snapshots, descriptor-bound launch, and descendant containment are not proven.
Flow #49 owns durable state; #31 owns retry and resume. The scenario catalog
The [durable lifecycle corpus](lifecycle-scenarios.md) separately qualifies
#49/#64 state and recovery APIs for #31. `--all-targets` verifies both corpora
and retains both reports. This acceptance scenario catalog
does not expand those checkpoints or claim that two synthetic fixtures are
two real provider adapters.

Expand Down
6 changes: 4 additions & 2 deletions docs/integrations/durable-state.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Flow #49 adds a library API for persistent execution of fully prepared process
steps. It builds on the exact subject, authority, transcript, and artifact gates.
Flow #64 adds fresh graph eligibility and safe caller-selected dependent
execution. There is no product CLI or graph scheduler. #31 continues through
#65 (the lifecycle corpus) and #66 (authority/effect and residual-state proofs);
#65 (the [lifecycle corpus](lifecycle-scenarios.md)) and #66 (authority/effect and residual-state proofs);
#53 owns the supported CLI.

## Public entry points
Expand Down Expand Up @@ -205,5 +205,7 @@ cover deterministic fresh-root/reopen reports, transitive and branch invalidatio
all local identity boundaries, complete inventories, stale-report non-authority,
the single-step bypass, blocked future inputs, and preserved history. The
acceptance driver includes these test sources in its evidence identity and runs
them with `--all-targets`; the broader lifecycle receipt catalog belongs to #65.
them with `--all-targets`, along with the 34-row [lifecycle receipt corpus](lifecycle-scenarios.md)
from #65. The latter runs every recipe twice, includes fresh-process recovery,
and checks bounded portable reports.
macOS/Windows CI runs the portable store suite; Linux runs the full provider matrix.
131 changes: 131 additions & 0 deletions docs/integrations/lifecycle-scenarios.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,131 @@
# Durable lifecycle scenarios

Flow #65 qualifies durable recovery through the public APIs introduced by #49
and #64. The test-owned corpus lives in `tests/lifecycle_matrix/`, with 34 fixed
recipes and expected outcomes in `tests/fixtures/lifecycle-scenarios.v1.json`.
It supplements the [acceptance matrix](acceptance-scenarios.md). It does not add
a scheduler, a runtime state model, or a new public contract.

## What the evidence proves

Every recipe runs twice in independent workspaces, with the exact compiled
synthetic provider package. Both executions must produce equal normalized
receipts. Expectations are checked in as reviewed assertions; tests never
regenerate expected outcomes from observed results.

| Family | Scenarios | Required evidence |
| --- | --- | --- |
| Cancellation | Before launch, after intent, during provider execution, between steps | Ordered snapshots, attempt numbers, absent checkpoints, current graph eligibility |
| Provider failures | Timeout, nonzero exit, actual Unix signal, stdout/stderr overflow, partial output | Exact typed process or acceptance error; failed evidence retained without promotion |
| Fresh-process restart | Completed run, partial graph, host exit after intent, host exit after provider reaping | A new host process reopens the same workspace, reconstructs current public contexts, and reuses or continues only eligible work |
| Recovery | Retry cancelled or interrupted work, abandon invalid partial output, refuse unacknowledged or completed retries | Explicit recovery decisions, ordered attempt transitions, preserved uncertain output before retry |
| Staleness | Input, exact plan, configuration values, provider identity, executable bytes, validator implementation, output bytes | Local stale boundary, downstream invalidation, independent branch reuse, dependent launch refusal without state writes |
| Store refusal | Changed validation profile, corrupt checkpoint context, checksum mismatch, partial write, future schema, malformed JSON | Exact reopen refusal; no fallback to an older accepted snapshot |

The graph fixture is A → B plus independent C, with separate source bindings.
It proves ordering dependencies and stale-ancestor propagation. The existing
provider-kit compositions separately prove artifact handoff; #64's graph tests
cover transitive chains, fan-in, and complete context inventories.

## Trace and receipt contract

The test-owned `flow.lifecycle-scenario-catalog/v1` names each recipe's provider
mode, graph shape, recovery classification, expected ordered state frames, and
expected ordered assessments. Frames project the existing `RunState` snapshots:
sequence, status, attempt, checkpoint presence, typed failure, and recorded
recovery actions. Assessments reuse `RunStepAssessment` fields directly.

`flow.lifecycle-scenario-receipt/v1` adds the exact recipe digest, observed package,
executable, manifest, binding and source identities, immutable plan digest,
retained history digest, and candidate artifact digests. On corrupt-history
recipes, frames describe the valid history before fault injection; the final
history digest identifies the rejected snapshot bytes. No successful reopen or
new assessment is implied by those frames.

The fixture preserves originals before deliberate byte corruption and retains
failed candidates. Interrupted output is moved to a separate retained file
before an explicitly approved retry. An ignored Rust helper is a subprocess
entry point, explicitly executed by the matrix; it is not a skipped scenario.
Host exit code 73 bypasses destructors and proves lock release and fresh reopening.
The two exit points are before provider launch and after the runner reaps its
direct child but before durable acceptance. They do not simulate host death
while an unconfined provider or its descendants remain alive.

Recovery classifications describe each **test recipe's operator decision**.
Two provider-failure recipes revalidate the exact retained process transcript
through Flow: a provider-classified retryable failure and a terminal validation
failure. Both remain unaccepted; only an explicit operator decision moves them
to pending or abandoned. Provider retryability hints do not trigger execution.
Cancellation or interruption may be retried after acknowledgement and fresh
context checks; invalid partial output is abandoned in its terminal recipe.
Flow still exposes typed `RunFailure` and `ResumeEligibility`, not an automatic
retryability policy. Reopening, assessing, or deserializing a receipt grants no
authority and launches nothing. The [durable-state guide](durable-state.md) owns
production storage and recovery semantics.

## Running and checking coverage

Populate Cargo's locked cache, then run on Linux:

```console
python3 tools/run_acceptance_scenarios.py --lifecycle-only
python3 tools/run_acceptance_scenarios.py --all-targets
python3 tools/test_lifecycle_report.py
```

The first command writes `target/lifecycle-scenarios.v1.report.json`. The second
runs all Rust targets and verifies both this corpus and all 81 acceptance rows.
CI runs it on Rust 1.85 and stable and uploads both normalized JSON reports.
The driver removes older reports before running, captures output with a bound,
and refuses any failed test, missing/duplicate/unknown receipt, changed expected
trace, unsupported shape/version, recipe digest mismatch, inconsistent provider
identity, or fixture source mismatch. It hashes the tested sources before and
after execution; changes during validation prevent a successful report.

Reports carry source and catalog digests, the actual toolchain, checked budgets,
coverage gaps, and receipts. Rust compares both fresh-workspace executions;
Python independently checks completeness and each exact expected outcome.
Failure logs and raw process output remain local, outside CI uploads. The
portable allowlist excludes absolute roots, PIDs, timing samples, configuration
values, source bytes, provider messages, and raw streams. Canary assertions check
fixture paths and values. This is not a general redactor for arbitrary content.

## Resource bounds and remaining work

| Resource | Qualification bound |
| --- | --- |
| Rust test workers | One; explicit child-process contention tests still overlap their opens |
| Per recipe | Two executions; each completed execution must take at most 60 seconds |
| Entire Cargo command | 600 seconds, including compilation; timeout kills and reaps the test command's process group |
| Captured command output | 4 MiB enforced while reading, before any unbounded capture |
| Test/provider process address space | Linux `RLIMIT_AS`: 4 GiB per process, inherited by host children and providers |
| Single file written by a test/provider | Linux `RLIMIT_FSIZE`: 16 MiB; core dumps disabled |
| Receipt | 32 KiB |
| Durable history | At most 16 snapshots and 2 MiB per recipe workspace |
| Fixture artifacts | At most 16 recursive files and 1 MiB per node artifact root |
| All fixture files | 128 MiB across the recipe's roots, including copied provider binaries and retained evidence |
| Provider execution | Existing pinned kit deadline (5 seconds), 64 KiB stdout and stderr limits, cancellation grace |

The driver serializes Rust test workers because concurrent fork/exec can briefly
inherit another worker's Unix `flock` descriptors and make an unrelated immediate
reopen report `Busy`. This is test isolation, matching the portable store suite;
no lock assertions are weakened and the deliberate contention tests still run.
Concurrent unrelated host spawning is not qualified by this serial tier.

The Cargo target runner applies kernel limits to test executables, not compiler
or linker processes. Direct `cargo test` is useful for debugging but does not
establish the memory/file qualification. Per-recipe elapsed time and aggregate
filesystem totals are checked after execution; an otherwise hung helper is
bounded by the driver's command deadline. The address-space limit is not an
aggregate process-tree RSS quota. Killing the test process group on failure is
test supervision, not a production descendant-containment claim.

This tier qualifies Linux synthetic trusted-unconfined providers. Existing
macOS/Windows durable-store jobs provide narrower portability evidence. Real
provider releases, sandbox enforcement, hardware power loss, every crash window,
provider-native checkpoints, migration, automatic scheduling/retry, and
exactly-once external effects remain outside this proof. **Flow #66** owns the
remaining authority transitions, effect counters, and residual cleanup
qualification. Parent **#31 stays open** until that checkpoint passes. Future
scenario visualization in #62 can consume these checked receipts while retaining
these coverage limits.
Loading
Loading