Skip to content

Add retry policy with transient vs permanent failure classification - #37

Open
AvinashAnad wants to merge 1 commit into
theschoolofai:mainfrom
AvinashAnad:retry-policy
Open

AvinashAnad wants to merge 1 commit into
theschoolofai:mainfrom
AvinashAnad:retry-policy

Conversation

@AvinashAnad

Copy link
Copy Markdown

Summary

  • Live graph track extension: Added a RetryPolicy that classifies task failures as transient (timeouts, rate limits, connection drops) or permanent (bad input, auth errors, missing resources), and a RetryAwarePlanner that wraps any existing planner to automatically retry transient failures up to a configurable bound
  • Permanent failures fail fast without retry; transient retries are tracked with provenance metadata (attempt number, error class, original node ID) in the graph journal
  • Adversarial test proves that a result arriving after cancellation is silently discarded — cancelled work does not leak late results into the graph

Files changed

File Change
s13code/core/live_graph/core.py RetryPolicy dataclass + transient/permanent pattern sets
s13code/runtime.py RetryAwarePlanner wrapper + wiring into the runtime
s13code/core/live_graph/__init__.py Export RetryPolicy
tests/test_retry_policy.py 15 tests: unit + integration + adversarial
part1_proof.py Part 1 benchmark script (4 cases with full traces)
README.md Extension documentation section
.gitignore Exclude part1_traces.json

Test plan

  • All 59 tests pass (uv run pytest -q) — 44 original + 15 new
  • Ruff clean (uv run ruff check .)
  • Part 1 proof runs end-to-end (uv run python part1_proof.py)
  • Transient failure retries and succeeds on second attempt
  • Permanent failure (ValueError) does not retry
  • Max retries exhausted falls through to inner planner
  • Future nodes do not exist before their inputs complete
  • Cancelled retry does not leak result into graph
  • Adversarial: result after cancellation is discarded

🤖 Generated with Claude Code

The live graph planner now distinguishes transient failures (timeouts,
rate limits, connection drops) from permanent failures (bad input, auth
errors, missing resources) and automatically retries transient failures
up to a configurable bound. Permanent failures fail fast without retry.

Changes:
- RetryPolicy dataclass in core/live_graph/core.py classifies errors
  using pattern matching against known transient/permanent signatures
- RetryAwarePlanner in runtime.py wraps any planner to intercept
  task_failed events and re-enqueue transient failures as retry nodes
- 15 new tests covering: transient retry, permanent fail-fast, max
  retries exhaustion, future-node ordering, cancellation leak guard,
  and adversarial result-after-cancellation attack
- Part 1 benchmark script (part1_proof.py) for the four required cases
- README section documenting the extension with full traces

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 7, 2026 16:31

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@theschoolofai theschoolofai left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved for a strong retry-policy implementation and adversarial test suite; contribution credit is compared against earlier retry PRs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants