Conversation
… a new node A failed node was terminal, and apply_patch refuses to re-add its id, so the planner had no way to express 'that timeout deserves another attempt, that 404 does not'. On the shipped benchmark this stranded 4 of 14 cases: a transient fetch failure left the downstream node pending forever and the run returned an empty answer. - core/live_graph/retry.py: type-first failure classification and a bounded backoff policy. A permanent exception type short-circuits before any message inspection, so a hostile path cannot buy a retry of a refused sandbox escape. - GraphPatch.retry: a retry creates a NEW node that inherits the failed node's parents and adopts its children, so the failed attempt keeps its state and its place in the journal and nothing is edited in place. - GraphStore enforces the attempt ceiling and one-attempt-per-failure independently of the planner. - runtime.py: the planner classifies before planning around a failure, resolves branches by retry lineage, and no longer parents the distiller on a failed node. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
- answer_with_evidence checks node failure before the per-skill evidence arms. A failed researcher previously matched the researcher arm, produced no evidence at all, and the run answered 'no authorized memory matched' instead of saying the research had failed. - scripts/repro_retry.py: end-to-end proof against a local origin that is 503 twice then 200, and a second that is 404 forever. - scripts/offline_gateway_model.py: deterministic Ollama-compatible model tier so the proof reproduces with no GPU and no provider key. glc_v3 and S13Code are unmodified above it. - README: the seven required items for the extension. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
test_an_exhausted_research_branch_answers_instead_of_returning_nothing uses app_client, which boots a real S13Runtime defaulting to OllamaNomicEmbedder. Every memory write in that test's run - the inbound prompt, the answer episode - tries to embed over real HTTP to localhost:11434. It passed in the original sandbox only because an unrelated local server happened to be running on that port; in a clean environment with nothing on 11434 it fails with urllib.error.URLError: Connection refused. Fix: swap in DeterministicEmbedder before the run, exactly as every other app_client test in this suite already does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this adds
The live planner could not express the difference between a failure worth
retrying and one that never will be. A failed node was terminal, and
apply_patchrefuses to re-add an existing task id, so there was no vocabularyfor "that timeout deserves another attempt, that 404 does not".
That gap was not theoretical. On the shipped non-browser benchmark, 4 of 14
cases returned a completely empty answer: the planner attached the next stage
to a node that had failed,
GraphStore.ready()only releases a node once everyparent has succeeded, so the child sat in
pendingforever, the executor ranout of work and the run returned silently with nothing.
core/live_graph/retry.py— failure classification and a bounded backoffpolicy. Classification is type-first: a known permanent exception type
short-circuits before any message inspection, so an attacker-controlled path
cannot smuggle the word "timeout" into a
PermissionErrorand buy a retry ofa refused sandbox escape.
GraphPatch.retry— a retry creates a new node that inherits the failednode's parents and adopts its children. Nothing is edited in place: every
attempt keeps its own state, its own error and its own place in the journal.
GraphStoreenforces the attempt ceiling and one-attempt-per-failureindependently of the planner, so an LLM planner emitting
retrycannot loopthe budget away.
branches by retry lineage, and no longer parents the distiller on a failed
node.
Benchmark, same 14 cases, before → after:
10/14 completed, 4 empty answers, 4 stranded nodes → 14/14 completed, 0
empty answers, 0 stranded nodes.
Proof
README.mdcarries the required section: capability, exact API request, graphand ordered event trace, actual final result, evidence and provider/agent
assignments, the adversarial failure and its fix, and the commands to reproduce
from a fresh checkout with no provider key and no GPU.
scripts/repro_retry.pyruns both scenarios end to end against a local originthat answers 503 twice then 200, and a second that answers 404 forever.
Tests
76 passing (44 upstream, unchanged, plus 32 added), fully offline — no local
Ollama or provider key needed to run the suite.
tests/test_retry_policy_adversarial.pykeeps the pre-fix message-matchingclassifier as
naive_classify, so the regression is executable: each attackmust genuinely fool the old implementation and be refused by the new one. It
also covers duplicate and replayed planner decisions, crash-recovery replay,
and a retry result arriving after cancellation.
🤖 Generated with Claude Code
https://claude.ai/code/session_01FvXm9XSmn5E9fdXGypJKAZ