Skip to content

About

EAG V3 Session 13 live graph, memory, semantic chunking and A2A runtime

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

S13Code

S13Code is the standalone Session 13 agent runtime. It implements a live task graph, scoped and provenance-bearing memory, Rohan's semantic chunking V2, and Agent2Agent interoperability. It asks glc_v3 for model completions over HTTP and never owns provider credentials.

What runs where

Service Default address Responsibility
glc_v3 http://127.0.0.1:8111 Models, keys, routing and channels
S13Code HTTP http://127.0.0.1:8113 Graph, memory, documents and JSON-RPC A2A
S13Code gRPC 127.0.0.1:8114 Official A2A gRPC service
Ollama http://127.0.0.1:11434 Phi-4 segmentation and Nomic embeddings

Requirements

  • Python 3.11 or newer
  • uv
  • A running glc_v3
  • A running Ollama with phi4 and nomic-embed-text
ollama pull phi4
ollama pull nomic-embed-text
ollama serve

Install and run

Unzip glc_v3, S13Code, and S13Proof beside one another. Start glc_v3 first. Then, from this directory:

uv sync

export GLC_BASE_URL=http://127.0.0.1:8111
export S13_GATEWAY_PROVIDER=gemini
export S13_SANDBOX_ROOT="$PWD/sandbox"
export S13_CHUNK_MODEL=phi4:latest
export S13_LIVE_SEMANTIC_CHUNKING=1

uv run s13code serve

State is written under ~/.s13code by default. Set S13_DATA_DIR to use another directory.

Check both services:

curl http://127.0.0.1:8111/healthz
curl http://127.0.0.1:8113/healthz
curl http://127.0.0.1:8113/readyz
curl http://127.0.0.1:8113/.well-known/agent-card.json

Run a prompt

curl -s http://127.0.0.1:8113/v1/agent/runs \
  -H 'Content-Type: application/json' \
  -d '{
    "tenant_id": "course",
    "project_id": "s13",
    "user_id": "student-01",
    "agent_id": "assistant",
    "prompt": "Say hello."
  }'

The response contains the final answer, graph nodes and edges, ordered graph events, and provider/agent assignments. Inspect a persisted run with:

curl http://127.0.0.1:8113/v1/agent/runs/<run-id>

Index the sample corpus

The five files under sandbox/papers/ are fixed .txt fixtures for semantic chunking and retrieval proofs.

curl -s http://127.0.0.1:8113/v1/agent/runs \
  -H 'Content-Type: application/json' \
  -d '{
    "tenant_id": "course",
    "project_id": "papers",
    "user_id": "student-01",
    "prompt": "Index every .txt file under papers/. Confirm how many chunks were indexed in total."
  }'

Document ingestion is versioned and atomic: source preparation, semantic boundaries, exact spans, Nomic embeddings, and visibility succeed together or roll back together.

Architecture

  • s13code/core/live_graph/: durable graph state, patches, event replay and bounded parallel execution
  • s13code/core/memory/: scope checks, provenance, contradiction history, semantic chunking and FAISS retrieval
  • s13code/core/a2a_adapter/: Agent Cards, JSON-RPC, SSE/push, official gRPC and trust checks
  • s13code/gateway.py: the only S13Code → glc_v3 seam
  • s13code/runtime.py: joins graph, memory, tools and model calls into an inspectable run
  • tests/: executable invariants and regression cases

Test before opening a pull request

uv run ruff check .
uv run pytest -q

cd ../S13Proof
uv sync
uv run pytest -q

Student contribution

This section is the required Session 13 deliverable — the assignment rules, and the completed race_fetch write-up (capability, exact request, graph/event trace, final result, evidence and provider assignments, adversarial failure and fix, and reproduction commands) live below.

Fork the official theschoolofai/S13Code repository linked from Axiom, create a branch, implement one meaningful extension, and open one pull request against that repository. Do not open the Session 13 pull request against theschoolofai/glc_v3.

Add one subsection to this README in the same pull request. It must contain:

  1. the user-visible capability,
  2. the exact prompt or API request,
  3. the graph and ordered event trace,
  4. the actual final result,
  5. evidence and provider/agent assignments,
  6. the adversarial failure and its fix,
  7. commands that reproduce the result from a fresh checkout.

Do not commit .env, credentials, personal memory, generated databases, unrestricted local paths, benchmark output containing private data, or provider responses containing secrets. Use synthetic identities in every proof.

Speculative fetch racing with deterministic cancellation (race_fetch)

1. User-visible capability. The existing fetch mode always waits for every URL in a prompt to reach a terminal state before answering — a wave-wide barrier, even once one source has already produced everything needed. race_fetch adds a second mode for the same fetch_url skill: when a prompt names two or more URLs together with a racing phrase ("whichever", "race", "fastest to respond", "first to respond"), every candidate starts fetching in parallel, but the graph stops waiting the instant the first one returns real, non-empty content. That winner is connected to the answer step deterministically (first task_succeeded with usable text wins, ties broken by the executor's existing sorted event order — never an LLM guess), and every still-pending or still-running sibling is cancelled in the same patch rather than being left to finish and be silently ignored.

2. Exact API request (synthetic identities; run against a local S13Code on 127.0.0.1:8113):

curl -s http://127.0.0.1:8113/v1/agent/runs \
  -H 'Content-Type: application/json' \
  -d '{
    "tenant_id": "course",
    "project_id": "race-fetch-demo",
    "user_id": "student-01",
    "agent_id": "assistant",
    "prompt": "Whichever of these two pages loads first, https://www.python.org https://postman-echo.com/delay/4, fetch it and summarize what kind of page it is in one sentence."
  }'

postman-echo.com/delay/4 is a real endpoint that deliberately waits 4 seconds before responding — chosen so the "losing" fetch is a genuine, still-in-flight network request at the moment it gets cancelled, not a fixture that merely returns fast. The whole run below completed in 3.16 seconds, i.e. before the 4-second loser could ever have finished.

3. Graph and ordered event trace (live run run-3e46747a351a):

graph: fetch_1 ──▶ answer          (fetch_2 cancelled, no edge to answer)

1167. run_started
1168. graph_patched   add=[fetch_1, fetch_2] connect=[] cancel=[] finish=False
      reason="first frontier selected for race_fetch"
1169. task_started: fetch_1
1170. task_started: fetch_2
1171. task_succeeded: fetch_1
1172. task_cancelled: fetch_2
1173. graph_patched   add=[answer] connect=[[fetch_1, answer]] cancel=[fetch_2] finish=False
      reason="fetch_1 won the fetch race with real content; cancelling 1 still-racing siblings"
1174. task_started: answer
1175. task_succeeded: answer
1176. graph_patched   add=[] connect=[] cancel=[] finish=True
      reason="grounded answer produced"

4. Actual final result (verbatim answer field):

https://www.python.org is the official website for the Python programming language, providing resources for downloads, documentation, community news, and information on how to get started with the language.

5. Evidence and provider/agent assignments:

Node Skill State Result Provider / model
fetch_1 fetch_url succeeded https://www.python.org, HTTP 200, 7,143 chars — (raw tool, no LLM call)
fetch_2 fetch_url cancelled None —
answer answer_with_evidence succeeded grounded answer above gemini_1 / gemini-3.1-flash-lite-preview

Evidence handed to the answer LLM contained exactly one item — the fetch_1 page text, cited as web_page / https://www.python.org. fetch_2 contributed nothing: its stored result is None, not merely unused.

6. Adversarial failure and fix. Attacked this feature's own claim that "cancelled work does not leak a late result into the graph." answer()'s evidence-gathering scans every node's stored result unconditionally — it has no check that a node is graph-connected to answer — so the entire guarantee rests on one thing: the winning patch's cancel= actually reaching the store. A test (tests/test_race_fetch.py) wrapped the real GraphStore.apply_patch to silently strip cancel= from every patch (simulating a dropped kwarg or a tampered/partially-applied patch, not a reimplementation of the feature) and showed the failure directly: the "loser" ran to completion, its forbidden content was recorded as an ordinary succeeded result, and it passed answer()'s evidence filter verbatim. The same scenario against the real, unmodified code — no monkeypatch — showed the loser genuinely cancelled, result staying None, nothing to leak. Both outcomes are asserted in tests/test_race_fetch.py::test_ATTACK_defeated_cancellation_leaks_the_loser_into_the_graph and tests/test_race_fetch.py::test_the_same_attack_fails_against_the_real_shipped_cancellation.

7. Commands that reproduce this from a fresh checkout:

git clone <this-fork-url> S13Code && cd S13Code
uv sync

# Ollama and glc_v3 must already be running -- see "Install and run" above.
export GLC_BASE_URL=http://127.0.0.1:8111
export S13_SANDBOX_ROOT="$PWD/sandbox"
uv run s13code serve &

curl -s http://127.0.0.1:8113/healthz

curl -s http://127.0.0.1:8113/v1/agent/runs \
  -H 'Content-Type: application/json' \
  -d '{
    "tenant_id": "course",
    "project_id": "race-fetch-demo",
    "user_id": "student-01",
    "agent_id": "assistant",
    "prompt": "Whichever of these two pages loads first, https://www.python.org https://postman-echo.com/delay/4, fetch it and summarize what kind of page it is in one sentence."
  }'

# All race_fetch tests: 10 routing cases, 2 mechanism proofs, 1 store-level
# guarantee, 2 adversarial (attack + defended).
uv run pytest tests/test_race_fetch.py -v
uv run pytest s13code/core/live_graph/tests/test_live_graph.py -v

# Full suite
uv run pytest -q

License

MIT. See LICENSE.

Appendix: trace logs

Supporting evidence for the "Student contribution" section above — the full baseline ("floor") capability runs and the full race_fetch sample-query batch, both real live runs, pushed to the end of the file so the required write-up above stays the first thing a reviewer reads.

Trace logs — Part 1: reproduce the floor

Real, live runs against a local S13Code + glc_v3, one per baseline capability the graph already supports before this session's extension. Every prompt below is synthetic (no real personal data); the two birthday runs use a placeholder ~/.s13code/artifacts/... path in place of this machine's actual absolute path.

Populations of London, Berlin, Paris — parallel research + distill (parallel_search)

Prompt: Find the populations of London, Berlin, and Paris, and tell me which two are closest in size. Status: completed in 53.9s

run_started
task_started: search_1
task_started: search_2
task_started: search_3
task_succeeded: search_2
task_succeeded: search_1
task_succeeded: search_3
task_started: distill
task_succeeded: distill
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

Populations (latest estimates)

City Population estimate Source
London ≈ 9,188,200 worldpopulationreview.com
Berlin ≈ 3,775,697 worldpopulationreview.com
Paris ≈ 2,059,821 worldpopulationreview.com

The smallest gap is between Berlin and Paris (≈ 1.72 million), making them the two cities whose populations are closest in size.

"What is 5 plus 7?" — plain arithmetic, memory mode

Prompt: What is 5 plus 7? Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

5 plus 7 is 12. The provided evidence does not contain information regarding this calculation.

"What is the capital of France?" — memory mode

Prompt: What is the capital of France? Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The capital of France is Paris.

Lagos/Cairo/Kinshasa — before the population-routing fix (memory mode, insufficient evidence)

Prompt: For Lagos, Cairo, and Kinshasa, find current populations and growth rates and tell me which is growing fastest. Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: search
task_succeeded: search
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The provided evidence does not contain information regarding the current populations or growth rates for Lagos, Cairo, or Kinshasa. Consequently, there is insufficient evidence to determine which of these cities is growing the fastest.

This is the exact trace that motivated _extract_entity_list() — the city list is stated before "populations," so the original anchored regex never matched and the router fell to the generic, single-blind-search memory mode. See run below for the corrected behavior.

Lagos/Cairo/Kinshasa — after the fix (structured_population, full success)

Prompt: For Lagos, Cairo, and Kinshasa, find current populations and growth rates and tell me which is growing fastest. Status: completed

run_started
task_started: search_1
task_started: search_2
task_started: search_3
task_succeeded: search_3
task_succeeded: search_2
task_succeeded: search_1
task_started: distill
task_succeeded: distill
task_started: validate
task_succeeded: validate
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished
City Estimated Population (2026) Annual Growth Rate
Lagos 14,881,845 – 15,070,000 2.48%
Cairo 9,801,536 (City) / 23,535,000 (Metro) 2.00% (Metro)
Kinshasa 14,565,700 – 21,852,144 5.13%

A definitive ranking is not possible: Cairo's growth rate is tied to its metro area while Lagos's and Kinshasa's are not consistently defined by geographic scope, so validate (the coder_validator role) correctly refuses to force a "fastest" conclusion across incomparable definitions.

Fibonacci / prime / square-root / sum — no coder skill exists, LLM computes directly

Prompt: What is the 50th Fibonacci number? Use the coder skill to calculate it... Calculate the exact value of the 10,000th prime number. What is the square root of 10,000th prime number. Add the result of all three queries... Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: search
task_succeeded: search
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished
  1. 50th Fibonacci number: 12,586,269,025
  2. 10,000th prime number: 104,729
  3. √104,729 ≈ 323.6186028325
  4. Sum: 12,586,374,077.6186028325

The provided memory evidence regarding Claude Shannon and city populations does not contain information relevant to these mathematical calculations.

Worth being precise about what this proves: there is no coder_validator-style "coder skill" reachable from a plain-text request like this one — the prompt fell to plain memory mode, and the arithmetic was computed by the answer LLM itself, not by any sandboxed tool in this registry. Asking for a skill that doesn't exist doesn't error; it's silently ignored by the router, which only recognizes the fixed set of intents in _work_intent().

Hindi translation — no translator skill exists either, memory mode

Prompt: Please translate this phrase to Hindi using the translator skill: 'Hello world, the weather is beautiful today' Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: search
task_succeeded: search
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

"नमस्ते दुनिया, आज मौसम बहुत सुहावना है।"

Same pattern as the coder-skill run above: "using the translator skill" has no matching registry entry, so it falls to memory mode and the LLM translates directly rather than through any dedicated tool.

Fetch Claude Shannon's Wikipedia page — fetch mode + distill

Prompt: Fetch https://en.wikipedia.org/wiki/Claude_Shannon and tell me his birth date, death date, and three key contributions to information theory. Status: completed

run_started
task_started: fetch_1
task_succeeded: fetch_1
task_started: distill
task_succeeded: distill
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished
  • Birth Date: April 30, 1916
  • Death Date: February 24, 2001
  • Contributions: founded information theory (1948, A Mathematical Theory of Communication); introduced the term "bit"; developed information entropy.
Tokyo family-friendly activities + weather — parallel_search, two independent research roles

Prompt: Find 3 family-friendly things to do in Tokyo this weekend. Check Saturday's weather forecast there and tell me which one is most appropriate. Status: completed

run_started
task_started: search_activities
task_started: search_weather
task_succeeded: search_weather
task_succeeded: search_activities
task_started: distill
task_succeeded: distill
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

Three family-friendly activities: theme parks, interactive museums, themed cafés/shopping. Weather data for Saturday wasn't specific enough in the search results to recommend one activity over another — the answer says so honestly rather than guessing.

Birthday fact + calendar reminder — birthday_reminder mode, durable fact write

Prompt: My mom's birthday is 15 May 2026. Remember that and give me a calendar reminder for two weeks before and on the day. Status: completed

run_started
task_started: recall
task_started: remember
task_succeeded: recall
task_succeeded: remember
task_started: reminder
task_succeeded: reminder
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

You told me that your mom's birthday is 15 May 2026 [source: api://agent/runs].

Calendar reminders created for 1 May 2026 (two weeks before) and 15 May 2026 (the day of), written under ~/.s13code/artifacts/<run-id>/.

("Mom's birthday" here is the project's own standing synthetic fixture for the memory/fact-supersession demo, not real personal data — the same example already used throughout docs/session13_architecture.md.)

"When is mom's birthday?" — later turn, memory recall of the durable fact

Prompt: When is mom's birthday? Status: completed

run_started
task_started: recall
task_succeeded: recall
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

You told me that your mom's birthday is 15 May 2026. Reminders remain available under ~/.s13code/artifacts/<run-id>/.

Confirms the fact outlives the transcript it came from, and correctly cites the original fact record rather than its own prior answer — see docs/session13_architecture.md §2 for the full mechanism.

"Search for 'Python asyncio best practices'..." — search_fetch mode

Prompt: Search for 'Python asyncio best practices', read the top 3 results, and give me a short numbered list of the advice they agree on. Status: completed

run_started
task_started: search
task_succeeded: search
task_started: fetch_1
task_started: fetch_2
task_started: fetch_3
task_succeeded: fetch_2
task_succeeded: fetch_3
task_succeeded: fetch_1
task_started: distill
task_succeeded: distill
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished
  1. Use asyncio for I/O-bound tasks only.
  2. Avoid blocking the event loop.
  3. Use asyncio.run() as the entry point.
  4. Enable debug mode during development.
  5. Explicitly close asynchronous generators.
Index papers/attention.md — wrong extension, honest failure (not a bug)

Prompt: Index the file papers/attention.md and tell me what the three key contributions of the Transformer architecture are according to this paper. Status: completed

run_started
task_started: index_file
task_failed: index_file
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The file papers/attention.md could not be located in the provided index [source: graph://run-.../index_file]. Consequently, I cannot provide the three key contributions of the Transformer architecture from that source.

sandbox/papers/ only contains .txt fixtures — this prompt named a .md file that never existed. sandbox_path() correctly raised FileNotFoundError before anything was read; the answer worker reported the miss honestly rather than fabricating contributions.

Index papers/attention.txt — corrected path, full success

Prompt: Index the file papers/attention.txt and tell me what the three key contributions of the Transformer architecture are according to this paper. Status: completed

run_started
task_started: index_file
task_succeeded: index_file
task_started: recall
task_succeeded: recall
task_started: distill
task_succeeded: distill
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished
  1. Architecture based solely on attention — dispenses with recurrence and convolutions entirely.
  2. Superior performance and efficiency — state-of-the-art BLEU scores, more parallelizable, less training time.
  3. Generalizes to other tasks — successful application to English constituency parsing.

Full raw logs: data/testresult/old_queries_trace.txt.

Test results — Part 2: extend one boundary — Live graph track

The ten sample queries for race_fetch (see docs/new_skill.md), actually run live against a local S13Code rather than only checked for routing. This batch is what caught the gap already corrected in docs/new_skill.md: several of the original sample URLs (example.com/one, the .test TLD, httpstat.us) don't resolve to real content from this network, so those runs correctly reach the race's no-winner fallback instead of a working summary — that is the graph behaving honestly, not a defect. Two queries (real domains) show a genuine live winner-take-all race.

Query 1 — both URLs 404 (placeholder paths under example.com)

Prompt: Whichever of these two loads first, https://example.com/one https://example.com/two, fetch it and summarize what it says. Status: failed

run_started
task_started: fetch_1
task_started: fetch_2
task_failed: fetch_1
task_failed: fetch_2

No answer text returned in this particular run (see note below on status: failed runs). A separate live check in this session's own history got as far as a completed run reporting the 404s honestly instead — the mode routing and both 404s are consistent either way; only whether this specific run reached its terminal answer node differs.

Query 2 — .test TLD never resolves (RFC 2606), correct no-winner fallback

Prompt: Race these three pages and tell me about whichever responds first: https://a.test/page https://b.test/page https://c.test/page Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_started: fetch_3
task_failed: fetch_2
task_failed: fetch_3
task_failed: fetch_1
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The attempt to access the requested pages resulted in a connection failure for all three URLs — ConnectError: [Errno -2] Name or service not known for all three [source: graph://.../fetch_1, fetch_2, fetch_3]. Because none of the pages could be reached, there is insufficient evidence to rank them.

Query 3 — real domains, genuine live winner (news.ycombinator.com beats lobste.rs)

Prompt: Fetch whichever of https://news.ycombinator.com and https://lobste.rs loads first and summarize the front page. Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_succeeded: fetch_1
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The front page of Hacker News (https://news.ycombinator.com) was retrieved. [Summary of top stories — AI/ML, software tools, industry news, Show HN, general interest.]

Source: https://news.ycombinator.com

fetch_2 (lobste.rs) never reaches a task_succeeded/task_failed line in this trace — the graph moved to answer as soon as fetch_1 won, exactly the wave-barrier-avoidance this feature exists for.

Query 4 — .test mirrors, correct no-winner fallback

Prompt: I want the fastest to respond of these mirrors: https://mirror1.test/doc https://mirror2.test/doc https://mirror3.test/doc — fetch and summarize it. Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_started: fetch_3
task_failed: fetch_1
task_failed: fetch_2
task_failed: fetch_3
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

Evidence is insufficient to fulfill your request. All three mirrors resulted in connection errors due to unknown hostnames. Consequently, I cannot determine which mirror is fastest.

Query 5 — real domains, genuine live winner (api.github.com beats gitlab/bitbucket)

Prompt: First to respond wins: https://api.github.com https://api.gitlab.com https://bitbucket.org — fetch whichever answers first and describe it. Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_started: fetch_3
task_succeeded: fetch_1
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The resource at https://api.github.com responded first. This source provides a JSON directory of GitHub API endpoints — code search, commits, issues, repositories, users, gists, org info, rate limits.

A second genuine live race, this time with three racers instead of two — fetch_2/fetch_3 again show no terminal event of their own.

Query 6 — intentionally-fake domains, correct honest no-winner fallback (by design)

Prompt: Whichever of these loads first — https://this-domain-does-not-exist-12345.test/x https://this-other-domain-also-fake-98765.test/y — fetch it and summarize it. Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_failed: fetch_1
task_failed: fetch_2
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

The requested URLs could not be accessed — ConnectError: [Errno -2] Name or service not known for both. Evidence is insufficient to summarize either page.

This one is supposed to fail — it's the fallback-path sample query, and it behaves exactly as documented.

Query 7 — httpstat.us unreachable from this network (environment issue, not a race_fetch bug)

Prompt: Whichever of these responds first, https://httpstat.us/200 https://httpstat.us/500, fetch it and tell me the result Status: completed

run_started
task_started: fetch_1
task_started: fetch_2
task_failed: fetch_1
task_failed: fetch_2
task_started: answer
task_succeeded: answer
task_succeeded: answer · graph: finished

Both requests encountered a RemoteProtocolError — the server disconnected without sending a response. The evidence is insufficient to provide a result for either request.

Confirms the earlier finding in docs/new_skill.md: httpstat.us is not reliably reachable from this environment, independent of anything in this feature.

Queries 8–9 — negative controls (single URL / no racing word), both runs show status: failed

Prompts: Fetch https://example.com/only-one-url and tell me what it says. and Read these two pages and summarize both: https://example.com/a https://example.com/b

run_started
task_started: fetch_1
task_failed: fetch_1
run_started
task_started: fetch_1
task_started: fetch_2
task_failed: fetch_1
task_failed: fetch_2

Both correctly routed to plain fetch mode (confirmed separately via _work_intent() — see docs/new_skill.md), and both hit the same example.com/* 404s as query 1. Neither run's trace shows a task_started: answer line, unlike query 1's own completed sibling run and every .test DNS-failure case above, which do reach answer. Recorded here as an observed, not-yet-root-caused data point — worth investigating further (possibly the same transient provider/gateway hiccup pattern already documented in docs/prompt_ui.md §9), not claimed as diagnosed.

Query 10 (Lagos/Cairo/Kinshasa, no URLs) — transient failure, then a clean retry

Same scope-check query as the final "floor" entry above. First attempt:

run_started
task_started: search_1
task_started: search_2
task_started: search_3
task_succeeded: search_3
task_succeeded: search_1
task_failed: search_2

status: failed, no answer reached — one researcher call failed and the run didn't recover. Re-run moments later completed cleanly end-to-end (structured_population → distill → validate → answer, same structured table and validation-warning answer already shown in the Part 1 section above). Consistent with the transient-provider-hiccup pattern this project has already documented elsewhere (docs/prompt_ui.md §9) rather than a new finding.

Full raw logs: data/testresult/new_querie_trace.txt.

About

EAG V3 Session 13 live graph, memory, semantic chunking and A2A runtime

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages