Skip to content

feat(net): linger upstream requests past their last reader - #4931

Closed
kixelated wants to merge 3 commits into
mainfrom
quest/m1/request-linger
Closed

kixelated wants to merge 3 commits into
mainfrom
quest/m1/request-linger

Conversation

@kixelated

Copy link
Copy Markdown
Collaborator

Problem

A session's upstream SUBSCRIBE is canceled the instant its last subscriber leaves, and an upstream FETCH is cut short (all the way to the publisher) the instant no fetcher waits and no reader holds its group. A reader that re-subscribes, seeks, or blips therefore churns a cancel and a fresh request upstream, hop by hop, for what is a sub-second gap.

Approach

Every "nobody wants this any more" point on the session side now waits out a fixed track::REQUEST_LINGER (1 s) first, restarting whenever demand returns. A reader back within it rides the request still in flight. Per the quest's decisions, demand decides per request type, unsplit: a FETCH watches its fetch callers and group readers, a subscription watches its subscribers, and group demand is not divided between them.

  • Rust model. Fetching::drop only withdraws an attempt still queued (it has cost nothing). An attempt a handler took stays joinable, so a returning fetch_group joins it. The handler gives it up with the new crate-private group::Request::reject_unused, atomic with a join under the fetch lock (mirroring track::Request::reject_unused). poll_abandon wraps that with the linger.
  • Rust sessions. A crate-private time::Linger (a Deadline that re-arms while demand is present) now drives every linger, including the existing 30 s copy linger in both subscribers, which drops their hand-rolled arm/disarm code.
    • lite: begin_subscription returns Begin::Idle instead of canceling, and ServeLoop cancels once the linger runs out. A returning subscriber turns into a SUBSCRIBE_UPDATE on the live subscription. The TRACK_INFO wait and FetchServeRun (before the answer and mid-ingest) linger too.
    • moq-transport: the pre-SUBSCRIBE_OK wait, the established subscription, and the group FETCH (before and after FETCH_OK) linger.
  • JS (js/net) has the same lifecycle, so it mirrors it with util/linger.ts (REQUEST_LINGER_MS, abandoned(demand)) in both subscribers' serve loops, the IETF pre-SUBSCRIBE_OK wait, and the lite fetch setup and response pump. Its fetch entry already stayed joinable until the group closed.
  • The constant. It is 1 s. The quest asked for a measured value. I sized it from reasoning instead: a re-subscribe or seek gap is one downstream round trip plus scheduling, well under a second on paths we serve, and staying far below the 30 s cache linger bounds what an abandoned request pulls. I did not take a production measurement. That is open decision 1 below.

Impact

  • Public API: none. REQUEST_LINGER, time::Linger, group::Request::reject_unused and poll_abandon, Requests::is_queued are crate-private, and REQUEST_LINGER_MS / abandoned are internal to @moq/net.
  • Behavior: group::Request::demand() can now go used again after going unused (its doc says so). A third-party Dynamic handler that drops a request the moment demand leaves still works, but a caller that joined in that gap reads Dropped. Handlers should give up with the linger (crate-private for now; see follow-ups).
  • Behavior: a publisher behind a relay sees its track's demand drop about 1 s after the last viewer leaves, not at once. A route change keeps the old route's upstream request for the same second.
  • Wire: none. The same messages go out, only later or not at all.

Tests

  • rs/moq-net/tests/request_linger.rs: through a relay on lite-06, lite-07 and moq-transport-19, a re-subscribe and a re-fetch inside the linger ride the request already upstream (the publisher sees no new request and no cancel), and the subscription is canceled only once the linger runs out. All six fail with REQUEST_LINGER set to zero.
  • Model unit tests cover a popped fetch staying joinable and reject_unused racing a returning caller. Existing cancel tests now assert the request is still held inside the linger before advancing simulated time past it.
  • moq-tokio's broadcast_rejoin_replays_a_current_warm_cache runs on the wall clock and waits out the linger three rounds per version, which pushed it past nextest's 60 s timeout. Its versions now run concurrently (about 4 s).
  • just check passes apart from moq-uring worker tests, which fail locally with ENOMEM on RLIMIT_MEMLOCK. That limit is shared by every process of this user, other agents included, and the crate is untouched here.
  • js/net/src/util/linger.test.ts covers the helper on fake timers. The JS teardown tests (integration, lite, IETF) now check the request is still held inside the linger, then run it out on fake timers instead of the wall clock.

Alternatives

  • Linger in the model (defer Fetching::drop's withdrawal with a timer there). The model has no clock today, and every handler would inherit a delay it did not ask for. The session already owns the clock and the upstream request.
  • Split group demand into fetch and subscription halves, so a subscription's groups could linger too. Rejected in planning: a front's reader should still give up a stale group a fetcher happens to hold.

Follow-ups

  • Measure the real re-request gap (relay logs pairing "subscribe canceled (idle)" with the next "subscribe started" for the same track, or a seek in <moq-watch>), and retune REQUEST_LINGER if 1 s is off.
  • Expose the linger to in-process Dynamic handlers (moq-ffi, moq-c) if any start canceling their own upstream work on demand.

(Written by Claude Opus 5.5)

🤖 Generated with Claude Code

kixelated and others added 3 commits October 6, 2026 09:32
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An upstream SUBSCRIBE or FETCH now outlives its last reader by a one second
linger, so a reader that re-subscribes, seeks, or blips rides the request
still in flight instead of churning a cancel and a fresh request upstream.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kixelated

Copy link
Copy Markdown
Collaborator Author

Closing unmerged: the maintainer decided to abandon the request linger (2026-10-06). The linger lives in the session subscriber, so it also holds client subscriptions open after the app drops them, costing bandwidth on every rendition switch, and no relay churn has been measured that would justify it. The quest is deleted in #4930; revisit if relay logs show re-request churn.

(Written by Claude Opus 5.5)

@kixelated kixelated closed this Oct 6, 2026
@kixelated
kixelated deleted the quest/m1/request-linger branch October 8, 2026 18:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant