Skip to content

Live graph: tell a timeout apart from a 404, and retry only the timeout - #38

Open
AaryanG21 wants to merge 3 commits into
theschoolofai:mainfrom
AaryanG21:s13-live-graph-retry-policy
Open

AaryanG21 wants to merge 3 commits into
theschoolofai:mainfrom
AaryanG21:s13-live-graph-retry-policy

Conversation

@AaryanG21

Copy link
Copy Markdown

What this adds

The live planner could not express the difference between a failure worth
retrying and one that never will be. A failed node was terminal, and
apply_patch refuses to re-add an existing task id, so there was no vocabulary
for "that timeout deserves another attempt, that 404 does not".

That gap was not theoretical. On the shipped non-browser benchmark, 4 of 14
cases returned a completely empty answer
: the planner attached the next stage
to a node that had failed, GraphStore.ready() only releases a node once every
parent has succeeded, so the child sat in pending forever, the executor ran
out of work and the run returned silently with nothing.

  • core/live_graph/retry.py — failure classification and a bounded backoff
    policy. Classification is type-first: a known permanent exception type
    short-circuits before any message inspection, so an attacker-controlled path
    cannot smuggle the word "timeout" into a PermissionError and buy a retry of
    a refused sandbox escape.
  • GraphPatch.retry — a retry creates a new node that inherits the failed
    node's parents and adopts its children. Nothing is edited in place: every
    attempt keeps its own state, its own error and its own place in the journal.
  • GraphStore enforces the attempt ceiling and one-attempt-per-failure
    independently of the planner, so an LLM planner emitting retry cannot loop
    the budget away.
  • The planner classifies before it plans around a failure, resolves its
    branches by retry lineage, and no longer parents the distiller on a failed
    node.

Benchmark, same 14 cases, before → after:
10/14 completed, 4 empty answers, 4 stranded nodes → 14/14 completed, 0
empty answers, 0 stranded nodes
.

Proof

README.md carries the required section: capability, exact API request, graph
and ordered event trace, actual final result, evidence and provider/agent
assignments, the adversarial failure and its fix, and the commands to reproduce
from a fresh checkout with no provider key and no GPU.

scripts/repro_retry.py runs both scenarios end to end against a local origin
that answers 503 twice then 200, and a second that answers 404 forever.

Tests

76 passing (44 upstream, unchanged, plus 32 added), fully offline — no local
Ollama or provider key needed to run the suite.
tests/test_retry_policy_adversarial.py keeps the pre-fix message-matching
classifier as naive_classify, so the regression is executable: each attack
must genuinely fool the old implementation and be refused by the new one. It
also covers duplicate and replayed planner decisions, crash-recovery replay,
and a retry result arriving after cancellation.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ

AaryanG21 and others added 3 commits August 18, 2026 17:30
… a new node

A failed node was terminal, and apply_patch refuses to re-add its id, so the
planner had no way to express 'that timeout deserves another attempt, that 404
does not'. On the shipped benchmark this stranded 4 of 14 cases: a transient
fetch failure left the downstream node pending forever and the run returned an
empty answer.

- core/live_graph/retry.py: type-first failure classification and a bounded
  backoff policy. A permanent exception type short-circuits before any message
  inspection, so a hostile path cannot buy a retry of a refused sandbox escape.
- GraphPatch.retry: a retry creates a NEW node that inherits the failed node's
  parents and adopts its children, so the failed attempt keeps its state and
  its place in the journal and nothing is edited in place.
- GraphStore enforces the attempt ceiling and one-attempt-per-failure
  independently of the planner.
- runtime.py: the planner classifies before planning around a failure, resolves
  branches by retry lineage, and no longer parents the distiller on a failed
  node.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
- answer_with_evidence checks node failure before the per-skill evidence arms.
  A failed researcher previously matched the researcher arm, produced no
  evidence at all, and the run answered 'no authorized memory matched' instead
  of saying the research had failed.
- scripts/repro_retry.py: end-to-end proof against a local origin that is 503
  twice then 200, and a second that is 404 forever.
- scripts/offline_gateway_model.py: deterministic Ollama-compatible model tier
  so the proof reproduces with no GPU and no provider key. glc_v3 and S13Code
  are unmodified above it.
- README: the seven required items for the extension.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
test_an_exhausted_research_branch_answers_instead_of_returning_nothing uses
app_client, which boots a real S13Runtime defaulting to OllamaNomicEmbedder.
Every memory write in that test's run - the inbound prompt, the answer episode
- tries to embed over real HTTP to localhost:11434. It passed in the original
sandbox only because an unrelated local server happened to be running on that
port; in a clean environment with nothing on 11434 it fails with
urllib.error.URLError: Connection refused.

Fix: swap in DeterministicEmbedder before the run, exactly as every other
app_client test in this suite already does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant