Repository navigation
Conversation
A relay front lived until its route left, so a standing prefix claim kept a front, its driver task, and a session placeholder source for every path ever requested under it. A front now retires once every track is forgotten and no consumer holds its broadcast, and a session closes a minted source once nothing holds it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minted source's track handler was created only when its serve machine first ran, after the requester already held the source. A front that read a track in that gap found no handler and ended the track `NotFound`. Both subscribers now create the handler before the accept and hand it to the serve machine. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…backwards A relay's lingering copy, re-subscribed for a returning reader, took the route's answer as current even when its largest group fell below what the copy cached. Under a prefix claim the relay never sees a worker close an output, so a path served again from group 0 handed the reader the old output's group, then nothing until the new sequence passed it. On moq-lite 07 and moq-transport, whose answers carry the largest position, a copy on a route without an epoch now takes a lower largest group as new content under the old name: it withholds its cache, closes its source so the front ends, and fails with UNROUTABLE, and the reader re-requests a fresh front. A route with an epoch may be serving a lagging replica, so its copy keeps waiting as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-ups from this PR. A new m0 quest makes every relay in a chain withhold a refused idle copy's cache, so a viewer behind an edge relay never gets a restarted publisher's old group; claim-served epochs requires it. The cold relay Largest quest now covers moq-lite too, removing the accepted false positive where a relay under-reports its upstream's largest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Collaborator
|
Rebase notes from the 2026-10-08 quest audit (#5058, #5063):
(Written by Claude Opus 5.5) |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #5054. Until it merges, review only the last commit,
fix(net): a relay copy fails loud when upstream's largest group goes backwards.Implements
quest/m0/largest-regression.md.Problem
A relay keeps a track's copy for the linger after its last reader leaves, and a returning reader re-subscribes it from the newest group it cached. Under a prefix claim the relay never sees a worker close an output, so when the worker serves the path again from group 0, the answer's largest group is below the cache. The copy took the answer as current and served the old output's group, then nothing until the new sequence passed it.
Approach
On versions whose answer carries the largest position (moq-lite 07, moq-transport), an idle copy on a route without an epoch takes a largest group below what it still caches as new content under the old name. It then:
UNROUTABLE, as a superseded broadcast does, so the reader requests a fresh front.A route with an epoch may be serving a replica that lags the copy, so its copy keeps waiting as before. moq-lite 05 and 06 carry no largest and keep waiting on the floor, as does an answer with no largest (a cold moq-transport relay).
The plan's open check, other legitimate lower answers on a route without an epoch: an upstream relay answers from its own copy's cache rather than its upstream's largest, so a copy it re-created cold that holds only older groups (another subscriber's backfill arriving first) under-reports. That needs the downstream copy to outlive the upstream's plus a concurrent lower-start subscriber upstream. It costs a spurious
UNROUTABLEand a re-request, never stale data, and the fix belongs at the reporting relay, so the split stands. Equal group numbers can't be told apart and still splice.Tests, mocked time:
tests/largest_regression.rsrestarts a claim worker's output within the linger on moq-lite 07 and moq-transport 14, 16 and 22. The returning reader getsUNROUTABLE, then the new output from group 0, never the old group; without the fix every version gets the old group. A same-epoch standby behind the cached copy keeps the copy and resumes it; that test fails if epoch routes are judged too. Atrack.rsunit test covers the judgment and the withheld cache.Impact
track::Producer::regressesandwithhold_cacheare crate-private.UNROUTABLEin this case.Alternatives
Decisions
The contributor's proposals, for the maintainer's review. ✅ marks the chosen option.
UNROUTABLE, as a superseded broadcast gets.Follow-ups
Planned in this PR (last commit, quest files only), as the contributor's proposals for the maintainer's review:
quest/m0/refused-copy.md[M]: behind any chain of relays, a returning reader never gets a restarted publisher's old group. An idle copy refused before it was answered withholds its cache and ends, and no hop answers from a withheld cache, on lite-07 and moq-transport.claim-epochs.mdnow requires it, since a refused epoch reuses the same handler.quest/m1/ietf-cold-largest.md, retitled "A relay reports its upstream's largest": moq-lite folds in as a second publisher of the same model change, which removes the false positive accepted above.Planning decisions:
Edge-relay goal:
Trigger:
route_failederror fails over ✅ (right for every refusal source today, and no wire change)lite-07 datagram-only tracks (no START to wait for):
Claim-served epochs:
Lite upstream largest:
ietf-cold-largest✅ (one model change, two publishers; the high-watermark dies with a withheld copy)(Written by Claude Opus 5.5)
🤖 Generated with Claude Code