diff --git a/quest/m0/README.md b/quest/m0/README.md index 24856199b2..a426ea0d4b 100644 --- a/quest/m0/README.md +++ b/quest/m0/README.md @@ -41,9 +41,12 @@ Published API or wire breaks still land on dev; each quest's Plan says so. ## Required - [IETF FIN semantics](/quest/m0/ietf-fin-not-cancel.md) - a request stream FIN stops updates without cancelling, and REQUEST_UPDATE on a subscribe is parsed +- [IETF early streams](/quest/m0/ietf-early-streams.md) - a moq-transport stream that arrives before SETUP is held until SETUP lands, never aborted +- [SUBSCRIBE_TRACKS refusal](/quest/m0/ietf-subscribe-tracks.md) - a draft-18+ SUBSCRIBE_TRACKS gets NOT_SUPPORTED on its stream, not a session close - [Request caps](/quest/m0/request-caps.md) - lite message sizes, IETF request IDs, and per-session announces and subscriptions are bounded - [noq reassembly cap](/quest/m0/noq-reassembly-cap.md) - noq carries quinn's stream reassembly cap and the connection receive window is finite by default - [qmux reset race](/quest/m0/qmux-reset-race.md) - qmux handles RESET_STREAM under one lock instead of panicking +- [qmux credit](/quest/m0/qmux-credit.md) - qmux returns connection credit for dropped and stopped streams and delivers its close frame, on both lines - [Shared fronts](/quest/m0/shared-fronts.md) - viewer sessions share a front, so fronts scale with peers, not viewers - [Revalidate overflow](/quest/m0/revalidate-overflow.md) - no auth duration can overflow a deadline and abort the relay - [Wildcard](/quest/m0/wildcard/README.md) - a relay resolves subscriptions against advertised prefixes, a service claims the prefix it could serve and refuses the rest instead of enumerating broadcasts, and the browser player treats a covering claim as availability diff --git a/quest/m0/ietf-early-streams.md b/quest/m0/ietf-early-streams.md new file mode 100644 index 0000000000..8c2af0a245 --- /dev/null +++ b/quest/m0/ietf-early-streams.md @@ -0,0 +1,34 @@ +# [S] IETF streams that arrive before SETUP are held + +## Goal + +A moq-transport stream that arrives before SETUP completes is held and +handled once SETUP lands, never aborted and never fatal, in Rust and JS, on +every draft the tree negotiates. + +## Plan + +Facts (`origin/main`, 2026-09-30): + +- Rust server, drafts 17+: `accept_setup` aborts every uni stream that is not + SETUP (0x2F00) with `UnexpectedStream`, which reaches the wire as + INTERNAL_ERROR. Bidi streams wait in the transport queue. +- JS: `receiveSetup` reads the first incoming uni stream and fails the whole + handshake if it is not SETUP. +- Drafts 16 to 21 say early data SHOULD be buffered until the control + streams arrive, and permit resetting only bidi streams. Drafts 18+ add that + parameters needing negotiation SHOULD NOT be used before the peer's SETUP. +- Group and fetch decoding needs only the draft version, which ALPN fixes; + what needs SETUP state is PUBLISH_NAMESPACE (cluster), solicit, and hidden. + +Decided (2026-09-30): queue non-SETUP uni streams during the handshake and +hand them to the normal classifier (padding, group, fetch, unknown, per +the uni-stream classifier from moq-dev/moq#4603) once SETUP lands. +QUIC stream credit already bounds the queue, so no extra cap. Do the same in +the JS handshake. Bidi streams keep waiting as today. Check draft 22's text +when it is reachable. + +Tests in both languages: a padding stream and a group stream that arrive +before SETUP are handled after it, and the session opens. + +Public API: none. Wire: none; fixes conformance. diff --git a/quest/m0/ietf-subscribe-tracks.md b/quest/m0/ietf-subscribe-tracks.md new file mode 100644 index 0000000000..ca015fa251 --- /dev/null +++ b/quest/m0/ietf-subscribe-tracks.md @@ -0,0 +1,23 @@ +# [S] SUBSCRIBE_TRACKS is refused per request + +## Goal + +A SUBSCRIBE_TRACKS (0x51) on drafts 18 and later gets REQUEST_ERROR with +NOT_SUPPORTED on its own stream, in Rust and JS, and the session stays open. + +## Plan + +Today Rust closes the session: drafts 17+ hit the `_` arm in +`rs/moq-net/src/ietf/session.rs` (`UnexpectedStream`, PROTOCOL_VIOLATION). +JS aborts the stream without a REQUEST_ERROR (`js/net/src/ietf/connection.ts` +default arm), and its `control.ts` check for drafts 18+ looks unreachable; +delete it if so. The drafts say limited endpoints SHOULD answer unsupported +messages with NOT_SUPPORTED (0x3 in every draft's registry). On drafts before +18, 0x51 is not a defined message and stays fatal. moq-dev/moq#4610 applies the same refuse-per-request rule to +other messages. + +Decode enough of the message to reply on its stream, then refuse. Tests in +both languages: a draft-18+ SUBSCRIBE_TRACKS is refused NOT_SUPPORTED while +another subscription on the session keeps delivering. + +Public API: none. Wire: none; fixes conformance. diff --git a/quest/m0/qmux-credit.md b/quest/m0/qmux-credit.md new file mode 100644 index 0000000000..df04b21d68 --- /dev/null +++ b/quest/m0/qmux-credit.md @@ -0,0 +1,53 @@ +# [M] qmux returns every byte's credit and sends its close frame + +## Goal + +A qmux session returns connection-level credit for every byte it receives, +whether the app reads it, drops the stream unread, or it arrives after +STOP_SENDING, so a long-lived session never stalls on MAX_DATA. `close()` +delivers its APPLICATION_CLOSE frame before the transport drops, so a TCP or +WebSocket peer sees the close code whenever the transport stays writable +within the close bound. Both hold on the 0.5 line that +`main` pins and the 0.6 line `dev` uses. + +## Plan + +Facts from moq-dev/web-transport `rs/qmux/src/session.rs` (0.5.1 and main): + +- Credit returns only through `RecvStream::report_consumed` on reads. + `Drop for RecvStream` sends STOP_SENDING and returns stream-count credit, + but never consumes bytes still buffered or queued in `inbound_data`. +- On STOP_SENDING the writer removes the recv entry, so later STREAM data is + dropped before the connection-credit check, and a later RESET_STREAM + returns early, so its final-size gap is never consumed. The peer counts all + of it against MAX_DATA. +- A new stream's entry is inserted only after `accept_uni`/`accept_bi` hands + it over, so a stream dropped in that window leaves a stale entry that keeps + charging the connection window. +- `close()` queues APPLICATION_CLOSE and marks the session closed at once. A + writer that is mid-write treats that as interrupted and drops the frame. + 0.6 made `close()` idempotent and keeps the first reason, but not the order. + +Work: + +- Consume credit for unread bytes on drop, for STREAM data and the RESET + final-size gap after STOP_SENDING (keep enough of a retired stream's state + to account for its final size), and close the accept gap. +- Mark the session closed only once the writer has sent the close frame, with + a bound so a stalled transport still drops. +- Tests in `rs/qmux`: a session that drops many unread streams keeps + delivering past its initial window, and a peer reads the close code after a + close issued mid-write. +- Release on both lines (0.5.x after 0.5.2 from + [qmux reset race](/quest/m0/qmux-reset-race.md), and 0.6.x), then bump + `main`'s pin. `dev` picks up 0.6.x. + +Why m0: the WebSocket fallback and the planned edge-to-core `tls://` links +both run on qmux, and MoQ drops streams constantly. + +Public API: none. Wire: none. + +## Related + +- [qmux reset race](/quest/m0/qmux-reset-race.md) - same file and release train; lands first +- [qmux on noq-proto](/quest/m2/quic-qmux.md) - replaces these stream maps later diff --git a/quest/m1/README.md b/quest/m1/README.md index 3a3f8782d3..b258d96d61 100644 --- a/quest/m1/README.md +++ b/quest/m1/README.md @@ -27,6 +27,7 @@ QUIC studies there on that rule. - [Cluster routing](/quest/m1/cluster-routing/README.md) - edge and core tiers carry each broadcast into a region once over path vector, with the backbone hidden from end users - [Remove `--hop`](/quest/m1/hop-removal.md) - on `dev`, redundant publishers share an explicit `@`, and `--hop` and the publisher's Hop ID are gone - [One route cost](/quest/m1/route-cost.md) - on `dev`, Warm and Cold collapse to one static route cost +- [Delete removed cluster flags](/quest/m1/cluster-shims.md) - on `dev`, once a release has carried their refusals, `mesh` and `linger` leave `cluster::Config` - [Track tail interop](/quest/m1/track-tail-interop.md) - a Rust publisher ending a track with a group in flight is read to its end by the JS subscriber, and the reverse, in `just test interop` - [Merge queue](/quest/m1/merge-queue.md) - the required checks run on `merge_group`, so a stale green check can no longer break main - [Binding audio delay](/quest/m1/binding-surface.md) - moq-ffi and every wrapper configure and observe audio playout delay @@ -130,8 +131,11 @@ QUIC studies there on that rule. - [Plan: cache age-out](/quest/m1/cache-wall-eviction.md) - a swept benchmark decides whether the track cache ages groups out on wall time without a write - [Frame slot charge](/quest/m1/frame-slot-charge.md) - a group's frame slots past the first four count against the cache pool, including capacity a released group keeps - [Front deadlines](/quest/m1/front-deadline-index.md) - a front's per-event cost stops growing with its track count: an expiry index and per-track wakes, proven by a churn benchmark +- [Listener deadlines](/quest/m1/listener-deadlines.md) - io_uring, HTTP/2, and the internal listener bound slow handshakes and headers, and iroh honors `quic.keep_alive` +- [Papercuts](/quest/m1/papercuts.md) - JS refuses to serve a broadcast it did not produce, `just` skips `.scratch/`, and a uring test stops sleeping - [Front parking](/quest/m1/origin-front-parks.md) - an unroutable request waits on a front instead of re-asking on every route-table move - [Route wakes](/quest/m1/route-wakes.md) - a route change wakes only the fronts it can move, so pool churn stops scaling with served paths +- [Admission bench](/quest/m1/admission-bench.md) - the driver's per-event admission walk is benchmarked over tracks and copies - [Publish channel count](/quest/m1/publish-audio-channel-count.md) - forcing a channel count on an Audio.Capture stops costing the subscriber gaps of silence - [JS abandonment](/quest/m1/js-subscribe-abandonment.md) - a viewer returning during IETF subscribe setup keeps its track across microtasks - [Epoch primitive](/quest/m1/epoch.md) - one `Epoch` type in moq-net and @moq/net, carried as a trailing `@` path segment, shared by e2ee and broadcast epochs diff --git a/quest/m1/admission-bench.md b/quest/m1/admission-bench.md new file mode 100644 index 0000000000..84c0a2474f --- /dev/null +++ b/quest/m1/admission-bench.md @@ -0,0 +1,23 @@ +# [S] Admission walk benchmark + +## Goal + +A benchmark sweeps tracks per front against copies per track for the +driver's admission walk, so a per-event cost that grows with the front +shows up as a slope. + +## Plan + +On the wildcard line, `run_front` (`rs/moq-net/src/model/origin.rs`) walks +every track and every copy (`io.copies()`) and calls `Provenance::admit` after +each event whenever `front.admit()` returns an origin, even when nothing +changed; Abort and End walk the same way. #4279 left this unbenchmarked +(moq-dev/moq#4607 covered pool resolution). The walk only runs once a +lite-07 reply names an origin, so extend `rs/moq-net/benches/session.rs` over +lite-07 mock sessions. If the slope matters, say so rather than optimizing +here; [Front deadlines](/quest/m1/front-deadline-index.md) owns per-track +wakes. + +## Required + +- The wildcard line (moq-dev/moq#4403) lands on main, where the walk lives diff --git a/quest/m1/cluster-shims.md b/quest/m1/cluster-shims.md new file mode 100644 index 0000000000..2147ec27a8 --- /dev/null +++ b/quest/m1/cluster-shims.md @@ -0,0 +1,21 @@ +# [XS] Delete the removed cluster flags + +## Goal + +`cluster::Config` no longer carries the hidden `mesh` and `linger` fields, +and `--cluster-mesh` and `--cluster-linger` are unknown flags like any other. + +## Plan + +Both fields exist only to refuse their removed settings at startup with a +pointer to the replacement (`deprecated()` in `rs/moq-relay/src/cluster.rs`, +checked there and in `config.rs`). The linger refusal shipped in +moq-relay 0.15.0; the mesh refusal (#4601) ships in the next release. Once a +release has carried both, delete the fields, `deprecated()`, and its checks, +per the no-shim rule. Lands on `dev`: removing public fields is a break. + +Public API: removes two hidden `cluster::Config` fields. Wire: none. + +## Required + +- A moq-relay release carrying the `--cluster-mesh` refusal from #4601 diff --git a/quest/m1/listener-deadlines.md b/quest/m1/listener-deadlines.md new file mode 100644 index 0000000000..3283ecd6a3 --- /dev/null +++ b/quest/m1/listener-deadlines.md @@ -0,0 +1,40 @@ +# [M] Every listener bounds its handshake and headers + +## Goal + +Every relay listener bounds slow peers the way the default runtime does +after the handshake deadline (moq-dev/moq#4612): the io_uring +workers apply `listen.timeout`, the HTTPS listener's HTTP/2 path and the +`[internal]` listener drop a connection with no complete request in flight +for `listen.timeout` (covering slow headers and an idle keep-alive, even one +that ACKs every PING), and the iroh backend honors `quic.keep_alive`. + +## Plan + +- io_uring (`rs/moq-relay/src/uring.rs` `serve_connection`): bound the + WebTransport `Request::accept`, `respond`, and `accept_request_lite` by one + deadline from `listen::Config::resolved_timeout()`, on the worker's own + timer (`rs/moq-uring/src/timer.rs`), closing with the timeout code. Delete + the "There is no timeout here" note in `rs/moq-uring/src/quic/web.rs`. +- HTTP: `axum_server` builds hyper with no timer, and #4612's + `header_read_timeout` is HTTP/1 only. hyper's HTTP/2 PING keep-alive + (`http2().timer(..)`, `keep_alive_interval`, `keep_alive_timeout`) only + detects a dead peer: a client that ACKs every PING and never finishes a + request stays up. So the bound is a per-connection idle deadline: the + connection closes once it has had no complete request in flight for + `listen.timeout`, which covers both a trickled HEADERS block and an idle + keep-alive. hyper has no such knob, so it lives in the relay's serve path + (`rs/moq-relay/src/listener.rs` and `web.rs`, e.g. an in-flight counter in + the per-connection service driving that connection's graceful shutdown). + Set the PING keep-alive too, for dead peers. Apply both to the HTTPS + listener and the `[internal]` listener, which also gets the HTTP/1 header + timer. +- iroh (`rs/moq-tokio/src/iroh.rs`): set iroh 1.3's + `keep_alive_interval` from `quic.keep_alive`, and fix the docs that say iroh + has no knob (`rs/moq-tokio/src/quic.rs`, `doc/bin/relay/config.md`). +- Tests on a paused clock where the runtime allows: a stalled io_uring + handshake closes at the deadline, and an HTTP/2 client that ACKs PINGs but + never sends a request, or trickles one request's headers, is dropped at the + deadline while one with a request in flight is not. + +Public API: none beyond existing settings. Wire: none. diff --git a/quest/m1/papercuts.md b/quest/m1/papercuts.md new file mode 100644 index 0000000000..56a365d6e2 --- /dev/null +++ b/quest/m1/papercuts.md @@ -0,0 +1,26 @@ +# [S] Papercuts from the 2026-09-30 spawn + +## Goal + +Three small fixes found while landing the m0 quests: `@moq/net` refuses to +serve a broadcast it did not produce, `just fix` and `just check` leave the +gitignored `.scratch/` alone, and `remote_wake_unparks` passes under load. + +## Plan + +- JS republish: `@moq/net` publishes only what it produces (decided with + moq-dev/moq#4599), naming its own origin. Requests already skip received + routes (`#demand` resolves through `bestEntry(path, received)` in + `js/net/src/origin.ts`), so the gap is an app handing a consumed + `broadcast.Consumer` (one a session delivered) back to an origin for + serving, e.g. `Request.accept(consumer)`, which labels upstream content + with the local origin's hop. Refuse that loudly at the point it enters the + origin, not in the lite publisher's `local(..) ?? demand(..)` resolve. +- `.scratch/`: add it to `.taplo.toml`'s excludes and to `.remarkignore`, so + `taplo format` and `sh/markdown.sh` stop rewriting agents' scratch clones. + biome, nixfmt, just, and shfmt already skip it. +- `rs/moq-uring/src/worker.rs` `remote_wake_unparks`: observe that the worker + parked instead of sleeping 50 ms and asserting elapsed wall time, per the + rule that unit tests mock time. + +Public API: none. Wire: none.