Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,9 @@ S13_LIVE_SEMANTIC_CHUNKING=1

# Set this to an absolute path before using local file skills.
S13_SANDBOX_ROOT=/absolute/path/to/S13Code/sandbox

# Retry policy for transiently failed graph nodes. Attempts include the first
# try, so 3 means one call plus at most two further attempts. Permanent
# failures are never retried regardless of these values.
S13_RETRY_MAX_ATTEMPTS=3
S13_RETRY_BASE_DELAY=0.05
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -14,4 +14,5 @@ htmlcov/
benchmark.json
benchmark.md
a2a-proof.json
retry-proof.json
.DS_Store
267 changes: 267 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,273 @@ Add one subsection to this README in the same pull request. It must contain:

Do not commit `.env`, credentials, personal memory, generated databases, unrestricted local paths, benchmark output containing private data, or provider responses containing secrets. Use synthetic identities in every proof.

## Extension: a live graph that tells a timeout apart from a 404

### 1. The user-visible capability

A single flaky network call no longer costs the user their answer. Before this
change a failed node was terminal, and because `apply_patch` refuses to re-add
an existing task id the planner had no way to say "that timeout deserves
another attempt, that 404 does not". On the shipped non-browser benchmark that
gap was not theoretical: **4 of the 14 cases returned a completely empty
answer**, because the planner attached the next stage to a node that had
failed, and `GraphStore.ready()` only releases a node once *every* parent has
succeeded. The child sat in `pending` forever, the executor ran out of work,
and the run returned silently with nothing. The live graph now classifies
every failure before it plans around it: a transient failure is re-attempted
as a **new node** with bounded exponential backoff, a permanent one is never
re-attempted, and either way the user gets either a grounded answer or an
explicit explanation of what failed. Across the same benchmark the result goes
from **10/14 completed with 4 empty answers and 4 stranded nodes** to
**14/14 completed, 0 empty answers, 0 stranded nodes**.

Classification is type-first by design, and that ordering is the
security-relevant part: a known permanent exception type short-circuits before
any message inspection, so an attacker-controlled path or URL cannot smuggle
the word "timeout" into a `PermissionError` and buy a retry of a refused
sandbox escape. A retry never edits a node in place — it creates a successor
that inherits the failed node's parents and adopts its children — so every
attempt keeps its own state, its own recorded error and its own place in the
journal.

### 2. The exact API request

```bash
curl -s http://127.0.0.1:8113/v1/agent/runs \
-H 'Content-Type: application/json' \
-d '{
"tenant_id": "synthetic-course",
"project_id": "retry-proof",
"user_id": "student-synthetic",
"agent_id": "assistant",
"prompt": "Fetch http://127.0.0.1:8231/report and tell me what it says."
}'
```

`http://127.0.0.1:8231/report` is a local origin started by
`scripts/repro_retry.py` that answers `503 Service Unavailable` to its first
two requests and serves a real page on the third. The permanent counterpart
uses the same request against `http://127.0.0.1:8231/always-missing`, which
answers `404` forever.

### 3. The graph and the ordered event trace

Transient origin, run `run-910f4ce055b0`:

```
544 run_started -
545 graph_patched - first frontier selected for fetch
546 task_started fetch_1
547 task_failed fetch_1 HTTPStatusError: Server error '503 Service Unavailable'
548 task_retry_scheduled fetch_1__retry2 attempt 2/3 after 0.05s — HTTP 503 is retryable
549 graph_patched - fetch_1 failed transiently (HTTP 503 is retryable); attempt 2 of 3
550 task_started fetch_1__retry2
551 task_failed fetch_1__retry2 HTTPStatusError: Server error '503 Service Unavailable'
552 task_retry_scheduled fetch_1__retry3 attempt 3/3 after 0.1s — HTTP 503 is retryable
553 graph_patched - fetch_1__retry2 failed transiently (HTTP 503 is retryable); attempt 3 of 3
554 task_started fetch_1__retry3
555 task_succeeded fetch_1__retry3
556 graph_patched - research evidence landed; specialist synthesis can begin
557 task_started distill
558 task_succeeded distill
559 graph_patched - specialist synthesis completed
560 task_started answer
561 task_succeeded answer
562 graph_patched - grounded answer produced
```

Final edges: `[["distill","answer"],["fetch_1__retry3","answer"],["fetch_1__retry3","distill"]]`

Two invariants are readable directly from that trace. **No future node exists
before its inputs**: `fetch_1__retry2` appears for the first time at sequence
548, *after* the failure at 547, and `distill` is created at 556 only once a
fetch attempt has actually succeeded — it is wired to `fetch_1__retry3`, never
to the two attempts that failed. And **the failed attempts are still there**:
`fetch_1` and `fetch_1__retry2` remain `failed` with their recorded errors;
nothing was rewritten to make the run look clean.

Permanent origin, run `run-c103337da291` — the whole graph, with no retry
event anywhere in it:

```
563 run_started -
564 graph_patched - first frontier selected for fetch
565 task_started fetch_1
566 task_failed fetch_1 HTTPStatusError: Client error '404 Not Found'
567 graph_patched - every research attempt failed; report the failure rather than stalling
568 task_started answer
569 task_succeeded answer
570 graph_patched - grounded answer produced
```

### 4. The actual final result

Transient origin — the origin's own counter confirms it received exactly three
requests for `/report`:

```
Based on the authorized evidence:
- Quarterly reliability report. [source: graph://run-910f4ce055b0/distill]
- The retry budget absorbed three upstream incidents this quarter. [source: graph://run-910f4ce055b0/distill]
- Quarterly reliability report. [source: http://127.0.0.1:8231/report]
- The retry budget absorbed three upstream incidents this quarter. [source: http://127.0.0.1:8231/report]

Sources consulted: graph://run-910f4ce055b0/distill, http://127.0.0.1:8231/report.
```

Permanent origin — one attempt, no retry, and the failure is reported rather
than swallowed:

```
The authorized evidence did not contain a passage matching this request.

One step did not succeed: The fetch_url step failed: HTTPStatusError: Client error
'404 Not Found' for url 'http://127.0.0.1:8231/always-missing'
[source: graph://run-c103337da291/fetch_1]

No usable source was available.
```

### 5. Evidence and provider/agent assignments

| node | agent | skill | state | attempt | provider/model |
|---|---|---|---|---|---|
| `fetch_1` | fetch_url | `fetch_url` | `failed` | 1 | — (local skill) |
| `fetch_1__retry2` | fetch_url | `fetch_url` | `failed` | 2 | — (local skill) |
| `fetch_1__retry3` | fetch_url | `fetch_url` | `succeeded` | 3 | — (local skill) |
| `distill` | distiller | `distiller` | `succeeded` | 1 | `ollama` / `s13-offline-extractive:1.0` |
| `answer` | answer_with_evidence | `answer_with_evidence` | `succeeded` | 1 | `ollama` / `s13-offline-extractive:1.0` |

Evidence in the final answer is `kind: web_page` from
`http://127.0.0.1:8231/report` (the successful third attempt) plus
`kind: role_output` from `graph://run-910f4ce055b0/distill`. Both model calls
crossed real HTTP to `glc_v3` on `127.0.0.1:8111`, which selected the `ollama`
provider; `S13Code` never saw a credential. Fetch attempts are local skills and
carry no provider.

The provider above is an **offline deterministic model**, not a frontier one.
`scripts/offline_gateway_model.py` serves Ollama's `/api/chat` and `/api/embed`
with an extractive engine so that this section reproduces byte-for-byte with no
GPU and no provider key — every layer above the model tier (`glc_v3` routing,
policy and accounting; the graph, journal, memory and A2A in `S13Code`) is the
real one. Point `OLLAMA_URL` at a real Ollama, or give `glc_v3` a provider key,
and the identical commands produce the identical graph with better prose.

Retry is policy in the planner and mechanism in the store. `GraphStore`
independently enforces the attempt ceiling and the one-attempt-per-failure
rule, so a planner — including an LLM planner emitting `retry` in a
`GraphPatch` — cannot loop the budget away.

### 6. The adversarial failure and its fix

**The attack.** A permanent failure wearing transient words. The user controls
part of a path, therefore part of the error message, and every retry-worthy
keyword can be smuggled into a failure that is in fact a *refused sandbox
escape*:

```
PermissionError: path escapes S13_SANDBOX_ROOT: /var/log/connection-reset/504-timeout-service-unavailable.txt
```

**The failure before the fix.** The first cut of this feature classified by
searching the message for retry-ish words. That implementation is kept in
`tests/test_retry_policy_adversarial.py` as `naive_classify`, so the regression
is executable rather than described:

```python
def test_before_the_fix_a_sandbox_escape_reads_as_retryable():
assert naive_classify(HOSTILE_ERROR) is FailureClass.TRANSIENT
```

It returns `TRANSIENT`, so the graph would re-attempt a path the sandbox had
already refused — up to the ceiling, three times, spending budget on an access
the system exists to deny.

**The fix.** `classify_failure` decides on the **exception type first**. A type
in `PERMANENT_TYPES` short-circuits before any message inspection, so no
message can promote it. Message hints are consulted only for types neither
table recognises, a permanent hint outranks a transient one, and an
unrecognised failure defaults to permanent — a graph that retries what it does
not understand is a graph that burns budget on real bugs.

```python
def test_after_the_fix_the_exception_type_decides_and_the_message_cannot_override_it():
verdict = classify_failure(HOSTILE_ERROR)
assert verdict.failure_class is FailureClass.PERMANENT
assert verdict.error_type == "PermissionError"
```

`test_the_graph_does_not_re_attempt_a_refused_sandbox_escape` runs the same
attack end to end through the executor and asserts the hostile path is
attempted **exactly once**. A parametrised family of disguises is checked the
same way: each must genuinely fool `naive_classify` *and* be refused by
`classify_failure`. A related case falls out of the same rule — `ProxyError` is
a transport type, but `ProxyError: 403 Forbidden` is permanent, because a proxy
that answers 403 will answer 403 again.

Two further attacks are covered in the same file. A **duplicate or replayed
planner decision** must not fan one failure into several concurrent attempts:
replaying the same `trigger_event` is a no-op via the `patches` table, a
duplicate arriving under a *new* event id is rejected with
`task ... has already been retried`, and a crash between a failure and its
patch replays into exactly one attempt on resume. A **result arriving after
cancellation** must not enter the graph: when a sibling finishes the run while
an attempt is still backing off, the attempt is cancelled during its sleep, its
worker never executes, no `task_succeeded` is journalled for it, and
`record_outcome` refuses outright for a node that is no longer `running`.

### 7. Reproducing this from a fresh checkout

Unzip or clone `glc_v3`, `S13Code` and `S13Proof` beside one another. No
provider key and no GPU are required for any step below.

```bash
# 1. tests and lint (44 upstream + 32 added)
cd S13Code
uv sync
uv run ruff check .
uv run pytest -q # 76 passed

cd ../S13Proof && uv sync && uv run pytest -q
uv run python run_a2a_proof.py --output a2a-proof.json

# 2. the offline model tier, so glc_v3's real OllamaProvider has something to call
cd ../S13Code
uv run python scripts/offline_gateway_model.py --port 11434 &

# 3. glc_v3, unmodified, routing to that local model
cd ../glc_v3 && uv sync
OLLAMA_MODEL=s13-offline-extractive:1.0 OLLAMA_URL=http://127.0.0.1:11434 \
LLM_ORDER=ollama uv run uvicorn glc.main:app --host 127.0.0.1 --port 8111 &

# 4. S13Code
cd ../S13Code
export GLC_BASE_URL=http://127.0.0.1:8111
export S13_GATEWAY_PROVIDER=ollama
export S13_SANDBOX_ROOT="$PWD/sandbox"
export S13_CHUNK_MODEL=s13-offline-extractive:1.0
uv run uvicorn s13code.main:app --host 127.0.0.1 --port 8113 &

curl http://127.0.0.1:8113/healthz && curl http://127.0.0.1:8113/readyz

# 5. the retry proof: both scenarios, graph, ordered trace, agents, answers
uv run python scripts/repro_retry.py --output retry-proof.json

# 6. the whole non-browser benchmark
cd ../S13Proof
uv run python run_benchmark.py --base-url http://127.0.0.1:8113 --output benchmark.md
```

To see the *before* state, run step 6 against a checkout of the commit
preceding this branch: `shannon`, `tokyo_weather`, `populations` and
`structured_growth` come back `failed` with an empty answer and one node
stranded in `pending`. On this branch all fourteen cases complete with a
non-empty answer.

Steps 5 and 6 write `retry-proof.json`, `benchmark.json` and `benchmark.md`;
all three are already covered by `.gitignore` and are not committed. Every
identity used above is synthetic.

## License

MIT. See `LICENSE`.
16 changes: 13 additions & 3 deletions s13code/core/live_graph/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,21 @@
GraphSnapshot,
LiveGraphExecutor,
NodeState,
RunReport,
TaskSpec,
)
from .store import GraphStore
from .retry import (
FailureClass,
FailureVerdict,
RetryPolicy,
RetryRequest,
classify_failure,
retry_node_id,
)
from .store import GraphMutationError, GraphStore

__all__ = [
"Event", "GraphPatch", "GraphSnapshot", "GraphStore",
"LiveGraphExecutor", "NodeState", "TaskSpec",
"Event", "FailureClass", "FailureVerdict", "GraphMutationError", "GraphPatch",
"GraphSnapshot", "GraphStore", "LiveGraphExecutor", "NodeState", "RetryPolicy",
"RetryRequest", "RunReport", "TaskSpec", "classify_failure", "retry_node_id",
]
25 changes: 24 additions & 1 deletion s13code/core/live_graph/core.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,8 @@
from enum import StrEnum
from typing import TYPE_CHECKING, Any, Awaitable, Callable, Protocol

from .retry import RetryRequest

if TYPE_CHECKING:
from .store import GraphStore

Expand Down Expand Up @@ -42,13 +44,19 @@ class GraphPatch:

``connect`` uses ``(parent, child)`` pairs. A waited node can be made
runnable later with ``resume``; it is never silently retried.

``retry`` is the explicit alternative to that silence. It re-attempts a
*failed* node as a brand-new node, so the failed attempt keeps its state,
its result and its place in the journal. Nothing here re-runs a node in
place: an attempt is a fact, and facts are not edited.
"""

add: tuple[TaskSpec, ...] = ()
connect: tuple[tuple[str, str], ...] = ()
cancel: tuple[str, ...] = ()
wait: tuple[str, ...] = ()
resume: tuple[str, ...] = ()
retry: tuple[RetryRequest, ...] = ()
finish: bool = False
reason: str = ""

Expand Down Expand Up @@ -82,6 +90,9 @@ class RunReport:
finished: bool
executed: tuple[str, ...]
waiting: tuple[str, ...]
# Nodes created as another attempt at a failed node, oldest first. Empty on
# a clean run, which is what makes a non-empty value worth reading.
retried: tuple[str, ...] = ()


class LiveGraphExecutor:
Expand Down Expand Up @@ -172,8 +183,18 @@ async def _execute(self, task: TaskSpec) -> tuple[TaskSpec, bool, dict[str, Any]
worker = self.skills.get(task.skill)
if worker is None:
return task, False, {"error": f"unknown skill: {task.skill}"}
# Backoff is durable node state, not a scheduler-side timer, so a
# process that restarts mid-wait still honours it. Sleeping inside the
# worker slot also keeps a backing-off retry cancellable: the graph can
# still cancel it, and `asyncio.CancelledError` lands here as it would
# for any other in-flight task.
delay = float(task.metadata.get("retry_delay_seconds") or 0.0)
try:
if delay > 0:
await asyncio.sleep(delay)
return task, True, await worker(task)
except asyncio.CancelledError:
raise
except Exception as exc: # worker failures become planner-visible events
return task, False, {"error": f"{type(exc).__name__}: {exc}"}

Expand Down Expand Up @@ -210,4 +231,6 @@ async def _replay_pending_planner_events(self, run_id: str) -> None:
def _report(self, run_id: str, executed: list[str]) -> RunReport:
snapshot = self.store.snapshot(run_id)
waiting = tuple(nid for nid, node in snapshot.nodes.items() if node["state"] == NodeState.WAITING)
return RunReport(run_id, snapshot.finished, tuple(executed), waiting)
retried = tuple(nid for nid, node in sorted(snapshot.nodes.items())
if node.get("metadata", {}).get("retry_of"))
return RunReport(run_id, snapshot.finished, tuple(executed), waiting, retried)
Loading