Skip to content

fix(net): a relay copy fails loud when upstream's largest group goes backwards - #5057

Draft
Dryvnt wants to merge 4 commits into
moq-dev:mainfrom
Dryvnt:quest/m0/largest-regression
Draft

Dryvnt wants to merge 4 commits into
moq-dev:mainfrom
Dryvnt:quest/m0/largest-regression

Conversation

@Dryvnt

@Dryvnt Dryvnt commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #5054. Until it merges, review only the last commit, fix(net): a relay copy fails loud when upstream's largest group goes backwards.

Implements quest/m0/largest-regression.md.

Problem

A relay keeps a track's copy for the linger after its last reader leaves, and a returning reader re-subscribes it from the newest group it cached. Under a prefix claim the relay never sees a worker close an output, so when the worker serves the path again from group 0, the answer's largest group is below the cache. The copy took the answer as current and served the old output's group, then nothing until the new sequence passed it.

Approach

On versions whose answer carries the largest position (moq-lite 07, moq-transport), an idle copy on a route without an epoch takes a largest group below what it still caches as new content under the old name. It then:

  • withholds its cache, since an aborted track otherwise lets readers drain it;
  • closes its minted source first, so the front ends and leaves the origin's table (fix(net): an unread front ends after its linger #5054) before any reader sees the error;
  • fails with UNROUTABLE, as a superseded broadcast does, so the reader requests a fresh front.

A route with an epoch may be serving a replica that lags the copy, so its copy keeps waiting as before. moq-lite 05 and 06 carry no largest and keep waiting on the floor, as does an answer with no largest (a cold moq-transport relay).

The plan's open check, other legitimate lower answers on a route without an epoch: an upstream relay answers from its own copy's cache rather than its upstream's largest, so a copy it re-created cold that holds only older groups (another subscriber's backfill arriving first) under-reports. That needs the downstream copy to outlive the upstream's plus a concurrent lower-start subscriber upstream. It costs a spurious UNROUTABLE and a re-request, never stale data, and the fix belongs at the reporting relay, so the split stands. Equal group numbers can't be told apart and still splice.

Tests, mocked time: tests/largest_regression.rs restarts a claim worker's output within the linger on moq-lite 07 and moq-transport 14, 16 and 22. The returning reader gets UNROUTABLE, then the new output from group 0, never the old group; without the fix every version gets the old group. A same-epoch standby behind the cached copy keeps the copy and resumes it; that test fails if epoch routes are judged too. A track.rs unit test covers the judgment and the withheld cache.

Impact

  • Public API: none. track::Producer::regresses and withhold_cache are crate-private.
  • Wire: no format change. A relay now resets a downstream subscription with UNROUTABLE in this case.

Alternatives

  • End the front on the copy's error instead of closing the source: a failed copy is the front's to replace, so it would re-query the same source and stall on the old floor.
  • Judge every route: a same-epoch standby that lags would cut viewers who should resume.

Decisions

The contributor's proposals, for the maintainer's review. ✅ marks the chosen option.

  1. Relays further from the worker, which still hand a returning reader the old group:
    • ✅ A new m0 quest right after this PR, covering both protocols. The gap where a group arrives before the answer folds into it.
    • Extend this PR.
    • Put only the moq-lite part in this PR.
  2. The false positive from a relay under-reporting its largest:
    • ✅ Accept it, with an m1 follow-up quest: a moq-lite relay reports its upstream's largest.
    • Accept it with no follow-up.
    • Block this PR on the fix.
  3. The error a regressed copy ends with:
    • ✅ Keep UNROUTABLE, as a superseded broadcast gets.
    • A different error.

Follow-ups

Planned in this PR (last commit, quest files only), as the contributor's proposals for the maintainer's review:

  • New quest/m0/refused-copy.md [M]: behind any chain of relays, a returning reader never gets a restarted publisher's old group. An idle copy refused before it was answered withholds its cache and ends, and no hop answers from a withheld cache, on lite-07 and moq-transport. claim-epochs.md now requires it, since a refused epoch reuses the same handler.
  • quest/m1/ietf-cold-largest.md, retitled "A relay reports its upstream's largest": moq-lite folds in as a second publisher of the same model change, which removes the false positive accepted above.

Planning decisions:

Edge-relay goal:

  • Confirm: lite-07 and moq-transport, with lite-05/06 unchanged ✅
  • Include lite-05/06
  • Narrow to moq-lite

Trigger:

  • Any refusal before an answer withholds the cache. A verdict error closes the source; a route_failed error fails over ✅ (right for every refusal source today, and no wire change)
  • UNROUTABLE only
  • A distinct regression error code

lite-07 datagram-only tracks (no START to wait for):

  • Accept and document the gap ✅
  • Send START for datagram tracks

Claim-served epochs:

  • Required on the new quest ✅ (one handler, not two)
  • Related only

Lite upstream largest:

  • Fold into ietf-cold-largest ✅ (one model change, two publishers; the high-watermark dies with a withheld copy)
  • A separate m1 quest

(Written by Claude Opus 5.5)

🤖 Generated with Claude Code

Dryvnt and others added 4 commits October 8, 2026 13:59
A relay front lived until its route left, so a standing prefix claim kept a
front, its driver task, and a session placeholder source for every path ever
requested under it. A front now retires once every track is forgotten and no
consumer holds its broadcast, and a session closes a minted source once
nothing holds it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A minted source's track handler was created only when its serve machine
first ran, after the requester already held the source. A front that read
a track in that gap found no handler and ended the track `NotFound`. Both
subscribers now create the handler before the accept and hand it to the
serve machine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…backwards

A relay's lingering copy, re-subscribed for a returning reader, took the
route's answer as current even when its largest group fell below what the
copy cached. Under a prefix claim the relay never sees a worker close an
output, so a path served again from group 0 handed the reader the old
output's group, then nothing until the new sequence passed it.

On moq-lite 07 and moq-transport, whose answers carry the largest position,
a copy on a route without an epoch now takes a lower largest group as new
content under the old name: it withholds its cache, closes its source so the
front ends, and fails with UNROUTABLE, and the reader re-requests a fresh
front. A route with an epoch may be serving a lagging replica, so its copy
keeps waiting as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-ups from this PR. A new m0 quest makes every relay in a chain
withhold a refused idle copy's cache, so a viewer behind an edge relay never
gets a restarted publisher's old group; claim-served epochs requires it. The
cold relay Largest quest now covers moq-lite too, removing the accepted false
positive where a relay under-reports its upstream's largest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kixelated

Copy link
Copy Markdown
Collaborator

Rebase notes from the 2026-10-08 quest audit (#5058, #5063):

  • quest: audit the whole tree (2026-10-08) #5058 moved largest-regression and claim-epochs to m1: delete quest/m1/largest-regression.md and edit quest/m1/claim-epochs.md, not the m0 paths.
  • quest/m1/ietf-cold-largest.md is deleted on main: cold-relay Largest stays INVALID_RANGE (quest/m0/README.md). Decided 2026-10-08: it is not reopened, so accept the false-regression cost rather than relying on that quest.
  • Decided 2026-10-08: the new refused-copy quest goes in m1 beside its parents, not m0.

(Written by Claude Opus 5.5)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants