diff --git a/doc/bin/relay/cluster.md b/doc/bin/relay/cluster.md index b1eb6ddbe7..6d126766e9 100644 --- a/doc/bin/relay/cluster.md +++ b/doc/bin/relay/cluster.md @@ -16,7 +16,10 @@ same way so the cluster converges instead of flapping. Both wire protocols carry it: natively on moq-lite, and via the [cluster extension](/draft/moq-cluster) on moq-transport 17+. -Failover routes must carry copies of the same broadcast. For each track, the +Failover routes must carry copies of the same broadcast. A relay moves a +subscription only between sources from the same origin: on moq-lite-07 the one a +source's SUBSCRIBE\_OK or FETCH\_OK names, otherwise the first hop of its route. +A change of origin ends the subscription and the viewer re-subscribes. For each track, the relay requires matching timescale, retention window, publisher priority, and group ordering. A source with different properties is refused before its groups are spliced in. If no compatible source remains, the track fails with @@ -43,18 +46,20 @@ link costs 1, which reproduces plain hop counting. Each relay adds the price of the link an announcement arrived on before forwarding it, so a route's cost is the sum of what it crossed. -Wildcard advertisements are forwarded and costed the same way as an exact-path +Prefix advertisements are forwarded and costed the same way as an exact-path route: each hop appends its identity, adds the link price, and passes the claim on. An advertisement must be contained by one of the publisher's granted -prefixes (`grant/**`); an over-wide pattern is refused rather than clamped. +prefixes (`grant/**`); an over-wide prefix is refused rather than clamped. -Routing prefers the most specific pattern, then a fully identified hop list +Routing prefers the longest covering prefix, then a fully identified hop list over one that holds a 0 (an anonymous hop) at any depth, then the lowest cost, -then the shortest hop list, breaking any remaining tie toward the newest -announcement so a reconnecting publisher isn't outranked by the session it -replaced. An assigned identity for an anonymous peer is local selection state -and is never written into the hop list. Resolving a non-prefix pattern into a -subscription is not implemented yet. +then the shortest hop list, then a hash of the requested path and the hop list, +breaking any remaining tie toward the newest announcement so a reconnecting +publisher isn't outranked by the session it replaced. Hashing the requested +path spreads equal-cost advertisers of one prefix, such as a transcode pool, +across its paths instead of sending every path to one of them, and every relay +picks the same one for a given path. An assigned identity for an anonymous peer +is local selection state and is never written into the hop list. ```toml [cluster] diff --git a/doc/concept/moq-lite.md b/doc/concept/moq-lite.md index c2947245ed..5f858f9962 100644 --- a/doc/concept/moq-lite.md +++ b/doc/concept/moq-lite.md @@ -161,7 +161,11 @@ member. The Rust consumer is a `Stream` and the TypeScript one an async iterable Announcements are hints; requests are the authority. When a subscriber asks for a covered path the advertiser will not serve, the advertiser refuses that -request rather than narrowing the claim, and no message narrows a route. Token +request rather than narrowing the claim, and no message narrows a route. A +refusal is final: a relay resolves a request against the longest covering +prefix and never retries another advertiser or a broader prefix. An advertiser +running out of capacity withdraws or re-prices its route instead, leaving +headroom for requests already in flight. Token scope is any pattern union; the session asks for each member's literal head on the prefix-only wire and filters locally. In Rust and TypeScript, `origin.scope(root, patterns)` narrows the handle's permissions and presents paths diff --git a/drafts/draft-lcurley-moq-cluster.md b/drafts/draft-lcurley-moq-cluster.md index 912be90e32..7712a07968 100644 --- a/drafts/draft-lcurley-moq-cluster.md +++ b/drafts/draft-lcurley-moq-cluster.md @@ -241,10 +241,8 @@ A refusal never falls through to a less specific tier. Within that tier, a receiver SHOULD prefer a HOP_PATH that contains no 0 entry over one that does, then the lowest ROUTE_COST, breaking ties toward the shorter HOP_PATH and then toward the most recently received. This is advisory: a receiver MAY apply local policy, such as measured RTT, instead. -NO_CAPACITY ({{iana}}) refuses a request the publisher could serve but has no capacity for now. -It permits ONE re-resolution within the same tier, excluding the refusing advertiser: every route with its non-zero first Hop ID, or its session when that ID is 0. -A receiver that has spent its retry, or has no other candidate, MUST refuse downstream with a code other than NO_CAPACITY, so retries cannot compound hop by hop. -Every other refusal, including an unrecognized code, is terminal. +Every refusal, including an unrecognized code, is terminal. +A publisher signals capacity through its advertisement alone: it withdraws or re-prices it before it runs out, leaving headroom for requests already in flight, since a withdrawal and a request for the slot it gave away can cross. A receiver SHOULD NOT cache refusals. A relay MUST NOT advertise a namespace merely because it resolved it: the covering advertisement stays the only one until the publisher advertises the concrete namespace, which it SHOULD do once producing, so a later request finds the running content by its exact namespace instead of resolving a second producer. @@ -266,7 +264,7 @@ One rule for advertisement and dispatch keeps advertised paths truthful and prev Under this extension an advertisement is a path, so a session advertises a namespace at most once, a relay forwards only the best path it knows ({{selection}}), and a subscription is served from one source at a time. A receiver MAY still hold paths to several publishers of one namespace and choose between them as it sees fit: serve from the cheapest and move to the next when it fails. -A refusal moves to another publisher only as {{selection}} allows: once, and only for NO_CAPACITY. +A refusal never moves to another publisher ({{selection}}). The advertised path and the served source stay the same publisher: a relay that moves to another MUST withdraw its advertisement and advertise the new path ({{updating}}), so the first Hop ID downstream always names the publisher whose Objects flow. Moving between distinct publishers is a discontinuity: their groups are not one sequence, so a subscriber sees an unrelated Location, and a FETCH that succeeds against one may fail against the other. @@ -283,7 +281,7 @@ Because a relay only appends to HOP_PATH, it cannot make a competing path look s ROUTE_COST has no such protection: it is a single value the sender chooses, so a relay can advertise 0 for content it is not carrying and attract subscriptions it then has to fetch. Both cost only a suboptimal path choice, and the latter is self-limiting, since the traffic won this way must then be served. -Implementations SHOULD bound the work started by requests beneath a broad advertisement, using NO_CAPACITY when capacity is exhausted. +Implementations SHOULD bound the work started by requests beneath a broad advertisement, withdrawing it before capacity is exhausted. A receiver MUST NOT make security decisions based on Hop IDs, and a deployment spanning a trust boundary SHOULD treat a peer's ROUTE_COST as a hint to clamp or ignore rather than an accounting figure. @@ -314,21 +312,13 @@ Both are carried in PUBLISH_NAMESPACE, in REQUEST_UPDATE of a PUBLISH_NAMESPACE The Key-Value-Pair parity is load-bearing: HOP_PATH is odd, so its value is a length-prefixed byte string, while HOP_ID, RELAY_COST, and ROUTE_COST are even, so their values are bare varints. -## MOQT Error Codes - -This document requests one registration in the "REQUEST_ERROR Codes" registry. - -| Value | Name | Reference | -|:--------|:------------|:--------------| -| 0x40B5A | NO_CAPACITY | This Document | - --- back # Appendix A: Changelog ## moq-cluster-02 -- Defined request resolution against the longest covering prefix and the NO_CAPACITY refusal with its single re-resolution; any other refusal is terminal, including between several publishers of one namespace. +- Defined request resolution against the longest covering prefix; every refusal is terminal, including between several publishers of one namespace, and capacity is signaled only by withdrawing or re-pricing the advertisement. - A relay does not advertise a namespace because it resolved it; the publisher advertises the concrete namespace once producing. ## moq-cluster-01 diff --git a/drafts/draft-lcurley-moq-lite.md b/drafts/draft-lcurley-moq-lite.md index 36bef0581f..26000264e9 100644 --- a/drafts/draft-lcurley-moq-lite.md +++ b/drafts/draft-lcurley-moq-lite.md @@ -305,8 +305,6 @@ Sent when resetting a stream (RESET_STREAM), or when refusing to receive one (ST | ------- | ------------- | ----------- | | 0x12 | MALFORMED_TRACK | The track's content could not be parsed. | | ------- | ------------- | ----------- | -| 0x30 | NO_CAPACITY | The publisher could serve this request but has no capacity for it now. Permits one re-resolution (see [Resolution](#resolution)); elsewhere it is terminal like any refusal. Bridges to NO_CAPACITY in {{I-D.lcurley-moq-cluster}}. | -| ------- | ------------- | ----------- | | 0x31 | CONTROL_TIMEOUT | The peer took too long to answer a control request. Distinct from DELIVERY_TIMEOUT, which is content that missed its deadline; it has no moq-transport value and bridges to INTERNAL_ERROR. | | ------- | ------------- | ----------- | | 0x32 | GROUP_TOO_LARGE | The group grew past the publisher's cache budget and was aborted. | @@ -418,12 +416,21 @@ The per-subscriber winner changing travels as an ANNOUNCE_UPDATE; the last quali When serving a subscription, a publisher MUST select the source by that same exclusion; if only excluded sources remain, the subscription is unroutable. Applying one rule to both advertisement and dispatch keeps advertised paths truthful, which is what prevents subscription cycles of any length. -When resolving a path covered by several routes (across any number of streams), the subscriber SHOULD prefer the most specific covering route (see [Resolution](#resolution)), then a path that contains no 0 Hop ID over one that does, then the lowest Warm Route Cost after adding each arriving link's cost (see [Cost Parameter](#cost-parameter)), breaking ties toward the lowest Cold Route Cost, then toward the shortest path, and then toward the most recently received, so a reconnecting publisher is not outranked by the stale session it replaced. +When resolving a path covered by several routes (across any number of streams), the subscriber SHOULD prefer the most specific covering route (see [Resolution](#resolution)), then a path that contains no 0 Hop ID over one that does, then the lowest Warm Route Cost after adding each arriving link's cost (see [Cost Parameter](#cost-parameter)), breaking ties toward the lowest Cold Route Cost, then toward the shortest path, then toward the lowest Spread Hash, and then toward the most recently received, so a reconnecting publisher is not outranked by the stale session it replaced. + +The Spread Hash is the 64-bit FNV-1a hash, with offset basis `0x420C0DECB00B` and the standard FNV-64 prime, of the requested path's UTF-8 bytes followed by each Hop ID of the route's path, oldest first, as 8 little-endian bytes. +It is keyed on the requested path rather than the route's prefix, so equal-cost advertisers of one prefix share its paths instead of the first one taking them all, while one path resolves to the same advertiser on every relay that holds the same routes. +When choosing which route to advertise for a prefix, the requested path is the prefix itself. -A route's identity is its first hop: the endpoint that originated it (see [ANNOUNCE_START](#announce-start)). -Two routes covering one path with the same non-zero first hop are the same origin reached different ways, and a relay MAY move a live subscription between them, resuming at a group boundary, so a route change the identity survives (a reconnect, a cheaper path, a draining session) is invisible to the subscriber. -Across differing first hops, or where either is 0, the routes promise nothing about each other's content: a relay MUST NOT splice a live subscription across them, and when the serving session ends, in-flight subscriptions end with it (a reset) and the subscriber re-requests through the best remaining route. -Equal first hops promise the same origin, not interchangeable bytes; what a resuming relay serves next is whatever that origin publishes next at the group boundary. +A subscription's identity is the origin serving it, named by the `Origin` field of the reply that carries its content: [SUBSCRIBE_OK](#subscribe-ok) for a subscription and [FETCH_OK](#fetch-ok) for a fetch. +A relay learns it from the reply rather than the route: a route promises only that paths under its prefix are servable, and an advertiser serving a prefix from several origins advertises one route for all of them. +Two sources of one track whose replies name the same non-zero Origin are the same origin reached different ways, and a relay MAY move a live subscription between them, resuming at a group boundary, so a change the identity survives (a reconnect, a draining session, a relay failing over within a pool) is invisible to the subscriber. +Across differing Origins, or where either is 0, the sources promise nothing about each other's content: a relay MUST NOT splice a live subscription across them, and instead ends it (a reset) so the subscriber re-requests through the best remaining route. +Group Streams are not ordered with the Subscribe Stream, so a relay that splices on Origin MUST hold a subscription's groups until its SUBSCRIBE_OK names their origin, and discard them if that origin is not the one it serves. +Datagrams cannot be held, so such a relay MUST drop a subscription's datagrams until then; a publisher sends SUBSCRIBE_OK before a subscription's first datagram as well as its first Group Stream. +A relay learns a replacement's Origin only once it replies, so it SHOULD keep serving from a live source when a better route appears rather than trade it for a source that may end the subscription. +A source reached over an earlier version, whose replies carry no Origin, is identified by its route's first hop: the endpoint that originated the route (see [ANNOUNCE_START](#announce-start)). +Equal Origins promise the same origin, not interchangeable bytes; what a resuming relay serves next is whatever that origin publishes next at the group boundary. #### Resolution {#resolution} A SUBSCRIBE, FETCH, or TRACK request names a path, and the receiver resolves it against the routes covering that path, after the per-subscriber exclusion above. @@ -444,11 +451,8 @@ A seed of `2^32` is RECOMMENDED only when `H * C < 2^32`; otherwise the deployme Unknown, out-of-budget, and saturated routes are outside this guarantee; a receiver MUST NOT infer that they outrank standby capacity merely because they might already carry content. An advertiser that will not serve a resolved request resets the request stream with a typed code (see [Error Codes](#error-codes)). -A NO_CAPACITY reset permits the receiver ONE re-resolution, within the same tier and excluding every route whose advertiser is the refusing one. -The exclusion is what makes the retry safe, not the retraction arriving first: an advertiser's capacity and a receiver's view of it are at least half a round trip apart, so a retraction and a request for the slot it gave away necessarily cross. -Re-resolution may find no other route, and the request is then unroutable; that is a correct outcome, not a fallback list. -A receiver that has spent its re-resolution, or has nothing to spend it on, MUST reset the downstream request with a code other than NO_CAPACITY, so the single retry cannot compound hop by hop. -Every other code, and any unrecognized one, is terminal and propagates without re-resolution, so probing unserved paths costs one round trip per path. +Every refusal is terminal and propagates without re-resolution, so probing unserved paths costs one round trip per path. +An advertiser signals capacity through its route alone: it withdraws or re-prices the route before it runs out, leaving headroom for requests already in flight, since a withdrawal and a request for the slot it gave away can cross. A receiver SHOULD NOT cache refusals; rate limiting is the advertiser's concern. ### Subscribe @@ -472,9 +476,9 @@ A subscriber opens a Fetch Stream (0x3) to request a single Group from a Track. The subscriber sends a FETCH message containing the broadcast path, track name, priority, group sequence, and the frame range within that group. Unlike SUBSCRIBE, FETCH works on both live and ended broadcasts; it is the only way to read an ended one. -The publisher responds with FRAME messages directly on the same bidirectional stream — there is no response header. +The publisher responds with a FETCH_OK naming the origin serving the group, followed by FRAME messages on the same bidirectional stream. The Subscribe ID, Group Sequence, and index of the first returned frame are implicit, taken from the original FETCH request. -Because there is no response header, a publisher that cannot serve the requested frame range in full MUST reset the stream rather than return a shorter run; the subscriber has no way to learn where a truncated response actually started. +Because the response carries no position, a publisher that cannot serve the requested frame range in full MUST reset the stream rather than return a shorter run; the subscriber has no way to learn where a truncated response actually started. As with a subscription, the subscriber MUST already have the track's [TRACK_INFO](#track-info) to parse the returned frames; because the properties are immutable, a single Track Stream lookup is reused across every FETCH of that track (group-by-group fetches do not re-fetch it). The publisher FINs the stream after the last frame, or resets the stream on error. @@ -1124,13 +1128,14 @@ Common values include `1000` (milliseconds), `1000000` (microseconds), `48000` ( A SUBSCRIBE_OK message confirms a subscription and resolves its absolute start position. It is the first message the publisher sends on the Subscribe Stream, once the start position is known. -This is the trimmed-down counterpart of MoqTransport's SUBSCRIBE_OK: it retains the name and the role of the publisher's positive response, but carries only the resolved start position (all other per-track properties live in [TRACK_INFO](#track-info)). +This is the trimmed-down counterpart of MoqTransport's SUBSCRIBE_OK: it retains the name and the role of the publisher's positive response, but carries only the resolved start position and who serves it (all other per-track properties live in [TRACK_INFO](#track-info)). ~~~ SUBSCRIBE_OK Message { Type (i) = 0x0 Message Length (i) Group (i) + Origin (i) } ~~~ @@ -1151,6 +1156,12 @@ The subscriber derives the start frame from `Group` and its own request: The second case is easy to get wrong, so to be explicit: a subscriber that requested group 5 frame 15 and receives `Group` = 6 starts at **frame 0** of group 6, not frame 15. The frame offset belonged to group 5 and is gone along with the rest of it; it does not carry forward to whichever group the publisher resolved to. +**Origin**: +The Hop ID of the origin serving the subscription, which relays splice failover on (see [Routing](#routing)). +An endpoint names the Origin its own source's reply named, or for a source reached over an earlier version, the first hop of that source's route, and for content it produces, its own Hop ID. +Where none of these identifies anyone (0, or an endpoint without a stable Hop ID of its own), it SHOULD generate a random Hop ID for that content and name it for as long as the content lasts, so a relay downstream can still resume within it. +A value of 0 names nobody, and a relay never splices across it. + ## SUBSCRIBE_END {#subscribe-end} A SUBSCRIBE_END message is sent by the publisher to signal that no group at or after a given sequence will be produced. @@ -1217,13 +1228,27 @@ The last frame to return (inclusive), encoded as the absolute frame index + 1. A value of 0 means through the end of the group (default). A `Frame End` below `Frame Start` once decoded is a protocol violation; equal bounds are a legal single-frame range. -The publisher responds with FRAME messages directly on the same stream — there is no response header. -The subscriber parses them using the track's [TRACK_INFO](#track-info), which it MUST already have (see the [Track Stream](#track-stream)); the group sequence and the index of the first frame are implicit from the FETCH request. +The publisher responds with a [FETCH_OK](#fetch-ok) followed by FRAME messages on the same stream. +The subscriber parses the frames using the track's [TRACK_INFO](#track-info), which it MUST already have (see the [Track Stream](#track-stream)); the group sequence and the index of the first frame are implicit from the FETCH request. The publisher FINs the stream after the last frame, or resets on error. There is no FETCH_ERROR message — the publisher signals failure by resetting the stream. A publisher holding fewer frames than requested MUST reset rather than truncate, since a short response is indistinguishable from one that started elsewhere. A group that ends before `Frame End` is not a truncation: the publisher FINs after the last frame it has, provided the group is complete and it served everything from `Frame Start` onward. +## FETCH_OK {#fetch-ok} +FETCH_OK is the publisher's answer on a Fetch Stream, sent once the group is resolved and before its first FRAME. + +~~~ +FETCH_OK Message { + Message Length (i) + Origin (i) +} +~~~ + +**Origin**: +The Hop ID of the origin serving the group, as in [SUBSCRIBE_OK](#subscribe-ok). +A relay that splices on Origin MUST NOT deliver the frames of a fetch whose Origin is not the one it serves. + ## PROBE PROBE is used to measure the available bitrate of the connection. @@ -1339,6 +1364,8 @@ The `Message Length` describes the payload size on the wire. - Added `Stream Count` to SUBSCRIBE_END: the number of Group Streams opened for the subscription. SUBSCRIBE_END is now sent once every counted Group Stream has opened, rather than as soon as the final group is known. - Removed SUBSCRIBE_DROP and its type 0x2; a group without a Group Stream is not counted. - The Subscribe Stream FIN now follows once every counted Group Stream has finished or been reset. +- Added `Origin` to SUBSCRIBE_OK and the FETCH_OK message ahead of a fetch's frames: the origin serving the request. A subscription's identity is now the Origin its reply names rather than its route's first hop, so a relay splices a failover only between sources naming the same non-zero Origin, including within a pool advertised by one route, and holds a subscription's groups (dropping its datagrams) until its SUBSCRIBE_OK names their origin. SUBSCRIBE_OK now precedes a subscription's first datagram too. +- Added the Spread Hash tie-break after the shortest path: a hash of the requested path and the route's Hop IDs, so equal-cost advertisers of one prefix share its paths. - Added announce compression: ANNOUNCE_START gains `Path Base` and `Path Keep` to copy the head of a live advertisement's suffix, and ANNOUNCE_START and ANNOUNCE_UPDATE gain `Hop Base` and `Hop Keep` to copy the tail of a live advertisement's Hop ID list. ## moq-lite-06 @@ -1352,7 +1379,7 @@ The `Message Length` describes the payload size on the wire. - Corrected SUBSCRIBE_END `Group` to an exclusive bound: the first sequence that will never be delivered, with 0 meaning no groups were produced. It was previously specified as the inclusive last group, which could not distinguish an empty track from one whose only group was 0. - Split ANNOUNCE_BROADCAST into three typed messages: ANNOUNCE_START (0x0), ANNOUNCE_END (0x1), and ANNOUNCE_UPDATE (0x2), each prefixed with a Type discriminator like the subscribe stream's responses. - Specified resolution of a request against the routes covering its path: only the most specific tier, the longest prefix, is consulted, so a concrete path shadows every broader route; a refusal never falls through, and cost and the existing tie-breaks order the tier. A route is always a prefix, and no message narrows one: an advertiser that serves only some of the paths beneath its prefix refuses the rest. A relay never announces a path because it resolved it; the advertiser announces the concrete path once producing. Standby ordering requires enforced deployment bounds on charged link count and cost; 2^32 is a recommended seed only when it exceeds their product. -- Split the reserved stream error range: 32 through 47 stays reserved, and 48 through 63 is moq-lite's own, assigned by the tables and mapped rather than forwarded across a bridge. Assigned 0x30 NO_CAPACITY there: it permits one re-resolution within the tier excluding the refusing advertiser, and a receiver that has spent or lacks that retry resets downstream with another code. Assigned 0x32 GROUP_TOO_LARGE: a group that grew past the publisher's cache budget is aborted. Every other code is terminal. +- Split the reserved stream error range: 32 through 47 stays reserved, and 48 through 63 is moq-lite's own, assigned by the tables and mapped rather than forwarded across a bridge. Assigned 0x32 GROUP_TOO_LARGE: a group that grew past the publisher's cache budget is aborted. Every code is terminal. - Assigned 0x33 NOT_FOUND, 0x34 OLD, and 0x35 EVICTED in the stream error table: a group the publisher cannot serve because it was never here, has been superseded, or was dropped under memory pressure. - Assigned 0x36 UNROUTABLE, 0x37 WRONG_SIZE, 0x38 FRAME_TOO_LARGE, and 0x39 TIMESTAMP_MISMATCH in the stream error table, moving them out of the reserved 32 through 47 range, which no longer carries provisional placeholders. - Assigned 0x31 CONTROL_TIMEOUT in the stream error table: a request stream torn down because the peer never answered, which DELIVERY_TIMEOUT described as late content. It has no moq-transport value and bridges to INTERNAL_ERROR. @@ -1514,7 +1541,7 @@ GOAWAY carries an optional New Session URI that asks the peer to reconnect elsew Hop IDs (see [ANNOUNCE_OK](#announce-ok) and [ANNOUNCE_START](#announce-start)) expose the relay path of a broadcast, which may reveal internal topology. A relay that does not wish to disclose its position MAY use the reserved value 0 ("unknown") instead of a stable identifier, at the cost of losing loop detection through itself (see [Routing](#routing)). The Hop ID announcement filter (see [Hop Parameter](#hop-parameter)) exists for loop avoidance, not access control: a subscriber cannot verify that a publisher honored it, so it MUST NOT be relied upon to hide a broadcast from a peer that declared its Hop ID. ## Resource Exhaustion -A peer can open many streams (subscriptions, announcements, fetches), request large announce prefixes, or advertise broad routes. Implementations SHOULD bound the number of concurrent subscriptions, announce matches, and cached groups, and SHOULD rely on QUIC flow control and stream limits to backpressure a misbehaving peer (see [ANNOUNCE_REQUEST](#announce-request)). Expiration (see [Expiration](#expiration)) bounds how long stale groups consume memory and flow control. A broad route invites a request for any covered path, each of which may start work: an advertiser SHOULD bound the work it starts and refuse beyond that with NO_CAPACITY, and a receiver re-resolves at most once per request, so a flood of requests costs the mesh one round trip each rather than a search (see [Resolution](#resolution)). +A peer can open many streams (subscriptions, announcements, fetches), request large announce prefixes, or advertise broad routes. Implementations SHOULD bound the number of concurrent subscriptions, announce matches, and cached groups, and SHOULD rely on QUIC flow control and stream limits to backpressure a misbehaving peer (see [ANNOUNCE_REQUEST](#announce-request)). Expiration (see [Expiration](#expiration)) bounds how long stale groups consume memory and flow control. A broad route invites a request for any covered path, each of which may start work: an advertiser SHOULD bound the work it starts, withdrawing its route before it runs out, and every refusal is terminal, so a flood of requests costs the mesh one round trip each rather than a search (see [Resolution](#resolution)). ## Datagram Injection Datagrams are routed to a subscription solely by Subscribe ID and carry no per-group authentication beyond that of the QUIC connection. On an unmodified QUIC/WebTransport connection this is sufficient, since datagrams are protected by the transport. A subscriber MUST silently drop any datagram with an unknown Subscribe ID and MUST deduplicate against groups received on streams (see [Datagrams](#datagrams)). diff --git a/js/net/src/broadcast.ts b/js/net/src/broadcast.ts index b744b042bb..0494138546 100644 --- a/js/net/src/broadcast.ts +++ b/js/net/src/broadcast.ts @@ -6,7 +6,7 @@ import { type GetPromise, Once, Signal } from "@moq/signals"; import { NotFound } from "./error.ts"; import type { Consumer as GroupConsumer } from "./group.ts"; -import { Route } from "./hop.ts"; +import { type Hop, Route, randomHop } from "./hop.ts"; import { hooks, type TrackSequence } from "./internal.ts"; import * as track from "./track.ts"; import { registerWire, trackOf, type Broadcast as Wire } from "./wire.ts"; @@ -31,6 +31,17 @@ class BroadcastState { // Live consumer handles sharing this state (see {@link Consumer.clone}). The broadcast // closes once the last one closes, so a shared consumer can be handed to several callers. consumers = 0; + // The origin serving this broadcast (see the wire's `origin`): named by an upstream + // reply, or generated on first use for content nobody named. + origin?: Hop; +} + +// The origin a peer is told serves this broadcast: proxied from upstream, or a random +// one for content originating here, stable for the broadcast's life and shared by every +// session serving it. +function origin(state: BroadcastState): Hop { + state.origin ??= randomHop(); + return state.origin; } function dequeueRequest(state: BroadcastState): track.Request | undefined { @@ -239,6 +250,10 @@ export class Producer { resolveTrackInfo: (name) => resolveTrackInfo(this.#state, name), fetchGroup: (name, sequence, options) => fetchGroup(this.#state, name, sequence, options), requested: () => this.#requested(), + origin: () => origin(this.#state), + name: (named) => { + this.#state.origin = named; + }, }; } @@ -304,6 +319,10 @@ export class Consumer { resolveTrackInfo: (name) => resolveTrackInfo(this.#state, name), fetchGroup: (name, sequence, options) => fetchGroup(this.#state, name, sequence, options), requested: () => this.#requested(), + origin: () => origin(this.#state), + name: (named) => { + this.#state.origin = named; + }, }); } diff --git a/js/net/src/error.test.ts b/js/net/src/error.test.ts index 61bb89b056..1093b8bf68 100644 --- a/js/net/src/error.test.ts +++ b/js/net/src/error.test.ts @@ -222,11 +222,11 @@ test("the code tables match the spec", () => { // carries nothing, so no code sits there. const assignedLite: StreamCode[] = [ StreamCode.ControlTimeout, - StreamCode.NoCapacity, StreamCode.GroupTooLarge, StreamCode.NotFound, StreamCode.Old, StreamCode.Evicted, + StreamCode.Unroutable, StreamCode.FrameTooLarge, ]; for (const code of Object.values(StreamCode)) { @@ -239,7 +239,6 @@ test("the code tables match the spec", () => { } // The values the Rust `StreamError` sends for the same conditions. expect(Number(StreamCode.ControlTimeout)).toBe(0x31); - expect(Number(StreamCode.NoCapacity)).toBe(0x30); expect(Number(StreamCode.GroupTooLarge)).toBe(0x32); expect(Number(StreamCode.NotFound)).toBe(0x33); expect(Number(StreamCode.Old)).toBe(0x34); @@ -361,7 +360,6 @@ test("toStreamCode and fromTransport agree on what a code means", () => { StreamCode.TooFarBehind, StreamCode.MalformedTrack, StreamCode.ControlTimeout, - StreamCode.NoCapacity, StreamCode.GroupTooLarge, StreamCode.NotFound, StreamCode.Old, diff --git a/js/net/src/error.ts b/js/net/src/error.ts index 04ad2b84c2..7e1ee4c81e 100644 --- a/js/net/src/error.ts +++ b/js/net/src/error.ts @@ -92,8 +92,8 @@ export const StreamCode = Object.freeze( Evicted: 0x35 as StreamCode, /** A frame declared a payload larger than the receiver accepts. */ FrameTooLarge: 0x38 as StreamCode, - /** The publisher could serve this request but has no capacity for it now. */ - NoCapacity: 0x30 as StreamCode, + /** The broadcast is neither announced nor served, so there is no route to it. */ + Unroutable: 0x36 as StreamCode, /** A group grew past its cache budget and was aborted. */ GroupTooLarge: 0x32 as StreamCode, } as const), diff --git a/js/net/src/ietf/publisher.ts b/js/net/src/ietf/publisher.ts index ee12f52118..432bb903da 100644 --- a/js/net/src/ietf/publisher.ts +++ b/js/net/src/ietf/publisher.ts @@ -224,7 +224,9 @@ export class Publisher { } catch (err: unknown) { const e = error(err); const condition = - e instanceof StreamError && e.code === StreamCode.NotFound ? "does_not_exist" : "internal"; + e instanceof StreamError && (e.code === StreamCode.NotFound || e.code === StreamCode.Unroutable) + ? "does_not_exist" + : "internal"; refusal = { errorCode: toRequestCode(condition, "subscribe", version), reasonPhrase: reason(e) }; } diff --git a/js/net/src/integration.test.ts b/js/net/src/integration.test.ts index c59f511f1e..43e44d2bdf 100644 --- a/js/net/src/integration.test.ts +++ b/js/net/src/integration.test.ts @@ -463,10 +463,18 @@ for (const [protocol, carriesOptIn] of [ }); } -test("integration: lite draft-05 datagram delivery", async () => { +// Draft-07 holds a subscription's content until SUBSCRIBE_START names its origin, so a +// datagram-only track must still send one for any datagram to arrive. +for (const protocol of [Lite.ALPN_05, Lite.ALPN_07_WIP]) { + test(`integration: ${protocol} datagram delivery`, async () => { + await datagramDelivery(protocol); + }); +} + +async function datagramDelivery(protocol: string) { const enc = new TextEncoder(); const dec = new TextDecoder(); - const pair = createMockTransportPair(Lite.ALPN_05); + const pair = createMockTransportPair(protocol); const origin = new OriginProducer(); const [client, server] = await Promise.all([ @@ -503,7 +511,7 @@ test("integration: lite draft-05 datagram delivery", async () => { remote.close(); client.close(); server.close(); -}); +} test("integration: lite draft-05 datagrams not sent on a non-datagram transport", async () => { const enc = new TextEncoder(); diff --git a/js/net/src/internal.ts b/js/net/src/internal.ts index 0f00c90848..2d9e0836ec 100644 --- a/js/net/src/internal.ts +++ b/js/net/src/internal.ts @@ -220,3 +220,23 @@ export const hooks: { throw new Error("broadcast.ts not loaded"); }, }; + +/** + * Spreads equal routes across paths: FNV-1a 64 of `path` then each hop, oldest first, as 8 + * little-endian bytes. Keyed on the requested path so an equal-cost pool advertising one + * prefix shares its paths, and every node holding the same routes picks the same member. + * Mirrors `fnv_key` in `rs/moq-net`; the seed is the draft's Spread Hash offset basis. + */ +export function spreadHash(path: string, hops: readonly bigint[]): bigint { + const prime = 0x100000001b3n; + let hash = 0x420c0decb00bn; + for (const byte of new TextEncoder().encode(path)) { + hash = BigInt.asUintN(64, (hash ^ BigInt(byte)) * prime); + } + for (const hop of hops) { + for (let shift = 0n; shift < 64n; shift += 8n) { + hash = BigInt.asUintN(64, (hash ^ ((hop >> shift) & 0xffn)) * prime); + } + } + return hash; +} diff --git a/js/net/src/lite/fetch.test.ts b/js/net/src/lite/fetch.test.ts index 0eb06dca94..f99f2d522b 100644 --- a/js/net/src/lite/fetch.test.ts +++ b/js/net/src/lite/fetch.test.ts @@ -1,7 +1,8 @@ import { expect, test } from "bun:test"; +import { HopSchema } from "../hop.ts"; import * as Path from "../path.ts"; import { Reader, Writer } from "../stream.ts"; -import { Fetch } from "./fetch.ts"; +import { Fetch, FetchOk } from "./fetch.ts"; import { Version } from "./version.ts"; function concat(chunks: Uint8Array[]): Uint8Array { @@ -44,3 +45,19 @@ test("Fetch round-trips on draft-03/04/05", async () => { expect(got.group).toBe(42); } }); + +test("FetchOk names the origin on draft-07 only", async () => { + const written: Uint8Array[] = []; + const writer = new Writer( + new WritableStream({ write: (chunk) => void written.push(new Uint8Array(chunk)) }), + ); + await new FetchOk(HopSchema.parse(42n)).encode(writer, Version.DRAFT_07); + writer.close(); + await writer.closed; + const buf = concat(written); + expect(buf).toEqual(new Uint8Array([1, 42])); + const got = await FetchOk.decode(new Reader(undefined, buf), Version.DRAFT_07); + expect(got.origin).toBe(HopSchema.parse(42n)); + + await expect(new FetchOk(HopSchema.parse(42n)).encode(writer, Version.DRAFT_06)).rejects.toThrow(); +}); diff --git a/js/net/src/lite/fetch.ts b/js/net/src/lite/fetch.ts index 35fe2e17de..e937b3c0b1 100644 --- a/js/net/src/lite/fetch.ts +++ b/js/net/src/lite/fetch.ts @@ -1,7 +1,8 @@ +import { type Hop, HopSchema } from "../hop.ts"; import * as Path from "../path.ts"; import type { Reader, Writer } from "../stream.ts"; import * as Message from "./message.ts"; -import { hasFrameBounds, Version } from "./version.ts"; +import { hasFrameBounds, hasOrigin, Version } from "./version.ts"; function guardFetch(version: Version) { switch (version) { @@ -102,3 +103,26 @@ export class Fetch { return Message.decode(r, (r) => Fetch.#decode(r, version)); } } + +/** + * FETCH_OK: the publisher's answer on a Fetch Stream, ahead of the frames, naming the + * origin serving the group (the Hop ID a relay stitches failover on). Draft-07+ only; + * older versions answer with the frames alone. + */ +export class FetchOk { + origin: Hop; + + constructor(origin: Hop) { + this.origin = origin; + } + + async encode(w: Writer, version: Version): Promise { + if (!hasOrigin(version)) throw new Error("FETCH_OK not supported for this version"); + return Message.encode(w, (w) => w.u62(this.origin)); + } + + static async decode(r: Reader, version: Version): Promise { + if (!hasOrigin(version)) throw new Error("FETCH_OK not supported for this version"); + return Message.decode(r, async (r) => new FetchOk(HopSchema.parse(await r.u62()))); + } +} diff --git a/js/net/src/lite/origin.test.ts b/js/net/src/lite/origin.test.ts new file mode 100644 index 0000000000..6518cb44ef --- /dev/null +++ b/js/net/src/lite/origin.test.ts @@ -0,0 +1,76 @@ +import { expect, test } from "bun:test"; +import * as broadcast from "../broadcast.ts"; +import { HopSchema, randomHop, UNKNOWN_HOP } from "../hop.ts"; +import { createMockTransportPair } from "../mock.ts"; +import * as Path from "../path.ts"; +import { Reader, Stream } from "../stream.ts"; +import { wireOf } from "../wire.ts"; +import { Group as GroupMessage } from "./group.ts"; +import { StreamId } from "./stream.ts"; +import { encodeSubscribeResponse, Subscribe, SubscribeStart } from "./subscribe.ts"; +import { Subscriber } from "./subscriber.ts"; +import { TrackInfo, Track as TrackMessage } from "./track.ts"; +import { ALPN_07_WIP, Version } from "./version.ts"; + +const VERSION = Version.DRAFT_07; + +/** Whether `promise` settles within `ms`. */ +async function settlesWithin(promise: Promise, ms: number): Promise { + let timer: ReturnType | undefined; + const pending = new Promise((resolve) => { + timer = setTimeout(() => resolve(false), ms); + }); + try { + return await Promise.race([promise.then(() => true), pending]); + } finally { + clearTimeout(timer); + } +} + +test("a broadcast originating here names one random origin for its life", () => { + const producer = new broadcast.Producer(); + const origin = wireOf(producer).origin(); + expect(origin).not.toBe(UNKNOWN_HOP); + expect(wireOf(producer).origin()).toBe(origin); + // Every handle, and so every session serving it, names the same one. + expect(wireOf(producer.consume()).origin()).toBe(origin); + expect(wireOf(new broadcast.Producer()).origin()).not.toBe(origin); +}); + +test("a draft-07 group waits for SUBSCRIBE_START, whose origin the broadcast then names", async () => { + const pair = createMockTransportPair(ALPN_07_WIP); + const subscriber = new Subscriber(pair.client, VERSION, randomHop()); + const consumer = subscriber.consume(Path.from("room")); + const reader = consumer.track("video").subscribe({}); + + const info = await Stream.accept(pair.server); + if (!info) throw new Error("the subscriber never asked for TRACK_INFO"); + expect(await info.reader.u53()).toBe(StreamId.Track); + await TrackMessage.decode(info.reader, VERSION); + await new TrackInfo({ maxAge: 60_000 }).encode(info.writer, VERSION); + info.close(); + + const sub = await Stream.accept(pair.server); + if (!sub) throw new Error("the subscriber never subscribed"); + expect(await sub.reader.u53()).toBe(StreamId.Subscribe); + await Subscribe.decode(sub.reader, VERSION); + + // The group's stream races ahead of the subscribe stream. + let controller!: ReadableStreamDefaultController; + const readable = new ReadableStream({ start: (c) => (controller = c) }); + void subscriber.runGroup(new GroupMessage({ subscribe: 0n, sequence: 0 }), new Reader(readable)); + controller.enqueue(new Uint8Array([0, 1, 120])); + controller.close(); + + const next = reader.recvGroup(); + expect(await settlesWithin(next, 50)).toBe(false); + + const origin = HopSchema.parse(42n); + await encodeSubscribeResponse(sub.writer, { start: new SubscribeStart(0, origin) }, VERSION); + const group = await next; + expect(group?.sequence).toBe(0); + expect(await group?.readString()).toBe("x"); + + // Republishing the broadcast proxies the origin upstream named. + expect(wireOf(consumer).origin()).toBe(origin); +}); diff --git a/js/net/src/lite/publisher.ts b/js/net/src/lite/publisher.ts index ed2a854dd8..98719ca354 100644 --- a/js/net/src/lite/publisher.ts +++ b/js/net/src/lite/publisher.ts @@ -13,7 +13,7 @@ import { type Advertised, type Advertisements, wireOf } from "../wire.ts"; import { AnnounceInit, AnnounceOk, type AnnounceRequest, encodeAnnounceBroadcast } from "./announce.ts"; import { Datagram as DatagramMessage } from "./datagram.ts"; import * as DatagramStream from "./datagram_stream.ts"; -import type { Fetch } from "./fetch.ts"; +import { type Fetch, FetchOk } from "./fetch.ts"; import { Group as GroupMessage } from "./group.ts"; import { Priority, sendOrder } from "./priority.ts"; import { Probe } from "./probe.ts"; @@ -31,6 +31,7 @@ import { hasAnnounceId, hasAnnounceOk, hasDatagrams, + hasOrigin, hasProbeRtt, hasStreamCount, resolvesStart, @@ -257,6 +258,40 @@ class SubscriptionControls { } } +/** + * A lite-05+ subscription's SUBSCRIBE_START and SUBSCRIBE_END, written in order. The group + * and datagram loops both serve the subscription, so START goes out ahead of whichever serves + * first: it resolves the start and names the origin, which a draft-07 subscriber needs before + * it delivers either. + */ +class SubscribeResponses { + #controls: SubscriptionControls; + #start: (sequence: number) => Promise; + #writes: Promise = Promise.resolve(true); + #started?: Promise; + + constructor(controls: SubscriptionControls, start: (sequence: number) => Promise) { + this.#controls = controls; + this.#start = start; + } + + /** Sends SUBSCRIBE_START at `sequence`, once; false when peer departure superseded it. */ + start(sequence: number): Promise { + this.#started ??= this.#write(() => this.#start(sequence)); + return this.#started; + } + + /** Sends SUBSCRIBE_END after any START still being written; false as for {@link start}. */ + end(write: () => Promise): Promise { + return this.#write(write); + } + + #write(write: () => Promise): Promise { + this.#writes = this.#writes.then((ok) => ok && this.#controls.response(write())); + return this.#writes; + } +} + // A microtask is too short: decoding one framed update crosses several awaits, each of which // can requeue behind the serving continuation. A task boundary lets the decoder finish whatever // the transport already delivered before the next group pop. Updates are rare, so groups do not @@ -599,12 +634,6 @@ export class Publisher { console.debug(`publish ok: broadcast=${msg.broadcast} track=${track.name}`); - // Serve datagrams concurrently with groups whenever the transport carries them - // (the writer exists iff so). No group fallback: otherwise they simply aren't sent. - if (this.#datagramWriter) { - datagrams = this.#runDatagrams(msg.id, track, timescale); - } - controls = new SubscriptionControls({ reader: stream.reader, writer: stream.writer, @@ -621,16 +650,40 @@ export class Publisher { }); }, }); + + const bounds: FrameBounds = { + startGroup: msg.startGroup, + startFrame: msg.startFrame, + endGroup: msg.endGroup, + endFrame: msg.endFrame, + }; + const responses = new SubscribeResponses(controls, async (sequence) => { + // SUBSCRIBE_START promises nothing below this sequence will be delivered. + // Arrival-order serving could later surface a straggler below the first + // group, so pin the floor to what was announced. + hooks.replaceGroups(track, { + start: { included: sequence }, + end: bounds.endGroup === undefined ? undefined : { included: bounds.endGroup }, + }); + // Read once content flowed: an upstream's SUBSCRIBE_START named the origin + // before any of its content did. + const start = new SubscribeStart(sequence, wireOf(front).origin()); + await encodeSubscribeResponse(stream.writer, { start }, this.version); + }); + + // Serve datagrams concurrently with groups whenever the transport carries them + // (the writer exists iff so, and only on lite-05+). No group fallback: otherwise + // they simply aren't sent. + if (this.#datagramWriter) { + datagrams = this.#runDatagrams(msg.id, track, timescale, responses); + } + await this.#runTrack(track, stream.writer, controls, { sub: msg.id, broadcast: msg.broadcast, timescale, - bounds: { - startGroup: msg.startGroup, - startFrame: msg.startFrame, - endGroup: msg.endGroup, - endFrame: msg.endFrame, - }, + bounds, + responses, }); console.debug(`publish done: broadcast=${msg.broadcast} track=${track.name}`); @@ -683,6 +736,9 @@ export class Publisher { // come off the same front, so the metadata and the frames are one generation. const info = await this.#resolveTrackInfo(front, msg.track); group = await wireOf(front).fetchGroup(msg.track, msg.group, { priority: msg.priority }); + if (hasOrigin(this.version)) { + await new FetchOk(wireOf(front).origin()).encode(stream.writer, this.version); + } await this.#runFetchGroup(group, stream.writer, { timescale: Timescale(info.timescale), start: msg.startFrame, @@ -714,13 +770,18 @@ export class Publisher { track: track.Subscriber, stream: Writer, controls: SubscriptionControls, - serving: { sub: bigint; broadcast: Path.Valid; timescale: Timescale; bounds: FrameBounds }, + serving: { + sub: bigint; + broadcast: Path.Valid; + timescale: Timescale; + bounds: FrameBounds; + responses: SubscribeResponses; + }, ) { - const { sub, broadcast, timescale, bounds } = serving; + const { sub, broadcast, timescale, bounds, responses } = serving; // Lite-05+ resolves the range on the subscribe stream: SUBSCRIBE_START once the // first group is known, SUBSCRIBE_END when the track finishes. const emitRange = supportsTrackStream(this.version); - let startSent = false; let endSent = false; // Lite-07+ counts the group streams in SUBSCRIBE_END, so it goes out only once every @@ -743,14 +804,12 @@ export class Publisher { const sendEnd = async (): Promise => { endSent = true; if (!emitRange) return true; - return controls.response( - (async () => { - // A group that gives up before its stream opens is never counted. - if (countStreams) while (opening.size > 0) await Promise.all(opening); - const end = new SubscribeEnd(boundary(), streams); - await encodeSubscribeResponse(stream, { end }, this.version); - })(), - ); + return responses.end(async () => { + // A group that gives up before its stream opens is never counted. + if (countStreams) while (opening.size > 0) await Promise.all(opening); + const end = new SubscribeEnd(boundary(), streams); + await encodeSubscribeResponse(stream, { end }, this.version); + }); }; // One ranking for the whole subscription, shared by every group it serves. @@ -844,26 +903,7 @@ export class Publisher { const group = recv.group; const range = frameRange(bounds, group.sequence); - if (emitRange && !startSent) { - startSent = true; - // SUBSCRIBE_START promises nothing below this sequence will be delivered. - // Arrival-order serving could later surface a straggler below the first - // group, so pin the floor to what was announced. - hooks.replaceGroups(track, { - start: { included: group.sequence }, - end: bounds.endGroup === undefined ? undefined : { included: bounds.endGroup }, - }); - if ( - !(await controls.response( - encodeSubscribeResponse( - stream, - { start: new SubscribeStart(group.sequence) }, - this.version, - ), - )) - ) - return; - } + if (emitRange && !(await responses.start(group.sequence))) return; const options: RunGroup = { sub, @@ -957,7 +997,7 @@ export class Publisher { * * @internal */ - async #runDatagrams(sub: bigint, track: track.Subscriber, timescale: Timescale) { + async #runDatagrams(sub: bigint, track: track.Subscriber, timescale: Timescale, responses: SubscribeResponses) { const writer = this.#datagramWriter; if (!writer) return; // Only reached with a writer (see the #datagramWriter gate). const maxSize = DatagramStream.maxDatagramSize(this.#quic); @@ -966,6 +1006,7 @@ export class Publisher { for (;;) { const datagram = await track.recvDatagram(); if (!datagram) return; // Track finished; #runTrack tears the subscription down. + if (!(await responses.start(datagram.sequence))) return; // Convert the timestamp to the track's advertised timescale, matching #serveGroup. const ts = Math.round(datagram.timestamp.as(timescale)); diff --git a/js/net/src/lite/subscribe.test.ts b/js/net/src/lite/subscribe.test.ts index afd4c7c0a6..a185678a37 100644 --- a/js/net/src/lite/subscribe.test.ts +++ b/js/net/src/lite/subscribe.test.ts @@ -1,4 +1,5 @@ import { expect, test } from "bun:test"; +import { HopSchema, UNKNOWN_HOP } from "../hop.ts"; import * as Path from "../path.ts"; import { Reader, Writer } from "../stream.ts"; import { @@ -151,6 +152,19 @@ test("SubscribeStart round-trips on draft-05", async () => { expect(got.start.group).toBe(42); }); +test("SubscribeStart names the origin on draft-07", async () => { + const start = new SubscribeStart(7, HopSchema.parse(42n)); + // Type, length, group, origin; draft-06 has no room for the origin. + expect(await encode(Version.DRAFT_07, { start })).toEqual(new Uint8Array([0, 2, 7, 42])); + expect(await encode(Version.DRAFT_06, { start })).toEqual(new Uint8Array([0, 1, 7])); + const got = await responseRoundtrip(Version.DRAFT_07, { start }); + if (!("start" in got)) throw new Error("expected start"); + expect([got.start.group, got.start.origin]).toEqual([7, HopSchema.parse(42n)]); + const old = await responseRoundtrip(Version.DRAFT_06, { start }); + if (!("start" in old)) throw new Error("expected start"); + expect(old.start.origin).toBe(UNKNOWN_HOP); +}); + test("SubscribeEnd round-trips on draft-05", async () => { // Type, length, group: no stream count before draft-07. expect(await encode(Version.DRAFT_05, { end: new SubscribeEnd(7, 3) })).toEqual(new Uint8Array([1, 1, 7])); diff --git a/js/net/src/lite/subscribe.ts b/js/net/src/lite/subscribe.ts index 3d1a1f2711..81e3beabea 100644 --- a/js/net/src/lite/subscribe.ts +++ b/js/net/src/lite/subscribe.ts @@ -1,7 +1,8 @@ +import { type Hop, HopSchema, UNKNOWN_HOP } from "../hop.ts"; import * as Path from "../path.ts"; import type { Reader, Writer } from "../stream.ts"; import * as Message from "./message.ts"; -import { hasFrameBounds, hasGroupOrder, hasStreamCount, resolvesStart, Version } from "./version.ts"; +import { hasFrameBounds, hasGroupOrder, hasOrigin, hasStreamCount, resolvesStart, Version } from "./version.ts"; /** * Encode the `Group Start` field shared by SUBSCRIBE and SUBSCRIBE_UPDATE. @@ -432,19 +433,30 @@ export class SubscribeOk { */ export class SubscribeStart { group: number; + /** + * The origin serving the subscription: the Hop ID a relay stitches failover on. + * {@link UNKNOWN_HOP} names nobody. Draft-07+; older versions decode it as unknown. + */ + origin: Hop; - constructor(group: number) { + constructor(group: number, origin: Hop = UNKNOWN_HOP) { this.group = group; + this.origin = origin; } - async encode(w: Writer): Promise { + async encode(w: Writer, version: Version): Promise { return Message.encode(w, async (w) => { await w.u53(this.group); + if (hasOrigin(version)) await w.u62(this.origin); }); } - static async decode(r: Reader): Promise { - return Message.decode(r, async (r) => new SubscribeStart(await r.u53())); + static async decode(r: Reader, version: Version): Promise { + return Message.decode(r, async (r) => { + const group = await r.u53(); + const origin = hasOrigin(version) ? HopSchema.parse(await r.u62()) : UNKNOWN_HOP; + return new SubscribeStart(group, origin); + }); } } @@ -557,7 +569,7 @@ export async function encodeSubscribeResponse(w: Writer, resp: SubscribeResponse // Draft-05+: SUBSCRIBE_OK is gone; START/END/DROP carry the resolved range. if ("start" in resp) { await w.u53(0x0); - await resp.start.encode(w); + await resp.start.encode(w, version); } else if ("end" in resp) { await w.u53(0x1); await resp.end.encode(w, version); @@ -592,7 +604,7 @@ export async function decodeSubscribeResponse(r: Reader, version: Version): Prom const typ = await r.u53(); switch (typ) { case 0x0: - return { start: await SubscribeStart.decode(r) }; + return { start: await SubscribeStart.decode(r, version) }; case 0x1: return { end: await SubscribeEnd.decode(r, version) }; case 0x2: diff --git a/js/net/src/lite/subscriber.test.ts b/js/net/src/lite/subscriber.test.ts index f9dec8b901..77a7e79ca9 100644 --- a/js/net/src/lite/subscriber.test.ts +++ b/js/net/src/lite/subscriber.test.ts @@ -1,5 +1,6 @@ import { expect, spyOn, test } from "bun:test"; import { Signal } from "@moq/signals"; +import { Producer as BroadcastProducer } from "../broadcast.ts"; import type { Probe as ProbeStats } from "../connection/stats.ts"; import { error, reason, StreamCode, StreamError } from "../error.ts"; import { HopSchema, isAnonymous, MAX_HOPS, Route, UNKNOWN_HOP } from "../hop.ts"; @@ -753,8 +754,8 @@ function expectCut(err: unknown, cause: Error | undefined) { } } -// Lite has no FETCH_OK, so a publisher that never answers holds each setup stage until the -// subscriber closes. The stream that stage opened is reset, even one opening after the close. +// Before lite-07 there is no FETCH_OK, so a publisher that never answers holds each setup stage +// until the subscriber closes. The stream that stage opened is reset, even one opening after the close. test.each([ ["the TRACK_INFO", "track", undefined], ["the FETCH", "fetch", undefined], @@ -765,7 +766,7 @@ test.each([ const subscriber = new Subscriber(quic, Version.DRAFT_05, HopSchema.parse(1n)); let settled = false; - const fetch = subscriber.fetchGroup(Path.from("room"), "video", 0).then( + const fetch = subscriber.fetchGroup(new BroadcastProducer().consume(), Path.from("room"), "video", 0).then( () => { settled = true; return undefined; @@ -799,7 +800,9 @@ test("a fetch started after the subscriber closes rejects without opening a stre const subscriber = new Subscriber(quic, Version.DRAFT_05, HopSchema.parse(1n)); subscriber.close(); - const err = await subscriber.fetchGroup(Path.from("room"), "video", 0).catch((err: unknown) => err); + const err = await subscriber + .fetchGroup(new BroadcastProducer().consume(), Path.from("room"), "video", 0) + .catch((err: unknown) => err); expectCut(err, undefined); expect(streams.length).toBe(0); }); diff --git a/js/net/src/lite/subscriber.ts b/js/net/src/lite/subscriber.ts index 61c3e7aaf2..582e5365ca 100644 --- a/js/net/src/lite/subscriber.ts +++ b/js/net/src/lite/subscriber.ts @@ -23,7 +23,7 @@ import { } from "./announce.ts"; import { Datagram as DatagramMessage } from "./datagram.ts"; import * as DatagramStream from "./datagram_stream.ts"; -import { Fetch as FetchMessage } from "./fetch.ts"; +import { Fetch as FetchMessage, FetchOk } from "./fetch.ts"; import { frameDecoder, type Group as GroupMessage, readFrames } from "./group.ts"; import { sendOrder } from "./priority.ts"; import { Probe } from "./probe.ts"; @@ -44,6 +44,7 @@ import { hasAnnounceId, hasAnnounceOk, hasDatagrams, + hasOrigin, hasProbeRtt, hasStreamCount, restartSupported, @@ -86,6 +87,12 @@ interface SubscribeEntry { // (SUBSCRIBE_END), once it declares them. start?: number; end?: number; + // The broadcast the subscription belongs to, which records the origin SUBSCRIBE_START + // names so a session republishing it names the same one. + broadcast: broadcast.Consumer; + // Whether SUBSCRIBE_START arrived. On draft-07, which names the serving origin there, + // group streams wait for it: until then nobody knows whose content they carry. + started: Signal; // Group streams opened by the publisher, when SUBSCRIBE_END carries the count. streams?: number; } @@ -504,14 +511,14 @@ export class Subscriber { for (;;) { const request = await wireOf(consumer).requested(); if (!request) break; - void this.#runSubscribe(path, request); + void this.#runSubscribe(consumer, path, request); } })(); return consumer; } - async #runSubscribe(broadcast: Path.Valid, request: track.Request) { + async #runSubscribe(consumer: broadcast.Consumer, broadcast: Path.Valid, request: track.Request) { const id = this.#subscribeNext++; const subscription = request.subscription; const initialBounds = groupBounds(subscription.groups); @@ -540,7 +547,7 @@ export class Subscriber { // Open the stream under a timeout. The stream handle flows back via `state` // so the timeout path can abort it if it finishes opening after the deadline. const state: { stream?: Stream } = {}; - const setup = this.#openSubscribe(state, msg, request, id, timescale); + const setup = this.#openSubscribe(state, msg, request, id, timescale, consumer); let opened: { stream: Stream; entry: SubscribeEntry }; try { @@ -635,6 +642,7 @@ export class Subscriber { request: track.Request, id: bigint, timescale: Signal, + consumer: broadcast.Consumer, ): Promise<{ stream: Stream; entry: SubscribeEntry }> { let producer: track.Producer; let drainOk = false; @@ -652,7 +660,13 @@ export class Subscriber { } // Register before opening SUBSCRIBE so a racing GROUP stream finds the entry. - const entry: SubscribeEntry = { track: producer, timescale, tail: new Tail() }; + const entry: SubscribeEntry = { + track: producer, + timescale, + tail: new Tail(), + broadcast: consumer, + started: new Signal(!hasOrigin(this.version)), + }; this.#subscribes.set(id, entry); state.stream = await Stream.open(this.#quic); @@ -726,6 +740,7 @@ export class Subscriber { // Open a FETCH stream for one group and stream its bare frames into a group, for the // ConsumeBroadcast backing track.Consumer.fetchGroup() (lite-05+). async fetchGroup( + front: broadcast.Consumer, broadcast: Path.Valid, track: string, sequence: number, @@ -737,7 +752,7 @@ export class Subscriber { let entry = this.#fetches.get(key); if (!entry || entry.group.isClosed) { const group = new netGroup.Producer(sequence); - entry = { group, accepted: this.#runFetch(broadcast, track, sequence, options, group) }; + entry = { group, accepted: this.#runFetch(front, broadcast, track, sequence, options, group) }; this.#fetches.set(key, entry); void group.closed.then(() => { if (this.#fetches.get(key)?.group === group) this.#fetches.delete(key); @@ -759,6 +774,7 @@ export class Subscriber { // Open the FETCH stream and pump the response into the shared group. Setup errors close the // group, evict the entry, and reject every caller waiting for acceptance. async #runFetch( + front: broadcast.Consumer, broadcast: Path.Valid, track: string, sequence: number, @@ -770,10 +786,10 @@ export class Subscriber { throw new Error("fetch group requires moq-lite-05 or newer"); } - // Lite has no FETCH_OK, so a publisher that never answers would hold the setup forever. - // Subscriber.close() closing the group releases every caller at any stage, and resets - // the streams the setup opened. - const setup = this.#fetchSetup(broadcast, track, sequence, options); + // A publisher that never answers would hold the setup forever. Subscriber.close() + // closing the group releases every caller at any stage, and resets the streams the + // setup opened. + const setup = this.#fetchSetup(front, broadcast, track, sequence, options); let accepted: { stream: Stream; info: TrackInfo }; try { accepted = await untilClosed(group, setup); @@ -795,6 +811,7 @@ export class Subscriber { // Resolve the track's timescale, then open the FETCH stream and wait for it to be accepted. async #fetchSetup( + front: broadcast.Consumer, broadcast: Path.Valid, track: string, sequence: number, @@ -805,9 +822,16 @@ export class Subscriber { return this.#exchange({ sendOrder: sendOrder({ priority }) }, async (stream) => { await stream.writer.u53(StreamId.Fetch); await new FetchMessage({ broadcast, track, priority, group: sequence }).encode(stream.writer, this.version); - // A byte or an empty-group FIN accepts the fetch; a reset rejects it. - // done() buffers that byte so the response pump can decode it normally. - await stream.reader.done(); + if (hasOrigin(this.version)) { + // Draft-07 accepts with FETCH_OK naming the serving origin before any frame; + // record it so a session republishing the broadcast names the same one. + const ok = await FetchOk.decode(stream.reader, this.version); + wireOf(front).name(ok.origin); + } else { + // A byte or an empty-group FIN accepts the fetch; a reset rejects it. + // done() buffers that byte so the response pump can decode it normally. + await stream.reader.done(); + } return { stream, info }; }); } @@ -872,6 +896,10 @@ export class Subscriber { if ("start" in resp) { entry.start = resp.start.group; + if (hasOrigin(this.version)) { + wireOf(entry.broadcast).name(resp.start.origin); + entry.started.set(true); + } } else if ("end" in resp) { if (entry.end !== undefined) throw new ProtocolViolation("duplicate SUBSCRIBE_END"); entry.end = resp.end.group; @@ -1000,11 +1028,22 @@ export class Subscriber { return; } - const { track, timescale, tail } = entry; + const { track, timescale, tail, started } = entry; const producer = new netGroup.Producer(group.sequence); const read = tail.open(group.sequence); try { + // Hold the group until SUBSCRIBE_START names whose content it is: the group's + // stream can arrive before the subscribe stream's. + while (!started.peek()) { + if (track.closed.peek() !== undefined) { + producer.close(); + stream.stop(new StreamError(StreamCode.Cancel, { message: "cancel" })); + return; + } + await Signal.race(started, track.closed); + } + track.writeGroup(producer); // Block until the timescale is known; the group's stream can arrive before @@ -1087,6 +1126,10 @@ export class Subscriber { const entry = this.#subscribes.get(dg.subscribe); if (!entry) return; // Unknown or already-closed subscription. + // A datagram cannot wait for SUBSCRIBE_START to name its origin like a group stream + // does, so on draft-07 it is dropped until then. + if (!entry.started.peek()) return; + // Datagrams are lite-05+, which always negotiates a timescale; if it hasn't resolved // yet (the datagram raced ahead of TRACK_INFO), drop rather than guess. const scale = entry.timescale.peek(); @@ -1209,7 +1252,7 @@ class ConsumeBroadcast extends broadcast.Consumer { super(state); overrideBroadcastWire(this, { resolveTrackInfo: (name) => subscriber.resolveTrackInfo(path, name), - fetchGroup: (name, sequence, options) => subscriber.fetchGroup(path, name, sequence, options), + fetchGroup: (name, sequence, options) => subscriber.fetchGroup(this, path, name, sequence, options), }); this.#subscriber = subscriber; this.#path = path; diff --git a/js/net/src/lite/version.ts b/js/net/src/lite/version.ts index 2d83114b6f..96ef8efca5 100644 --- a/js/net/src/lite/version.ts +++ b/js/net/src/lite/version.ts @@ -249,6 +249,22 @@ export function hasStreamCount(version: Version): boolean { } } +/** Whether SUBSCRIBE_OK and FETCH_OK name the origin serving the request, which relays stitch failover on. Added in lite-07, with FETCH_OK itself. */ +export function hasOrigin(version: Version): boolean { + // Explicitly list older versions so future versions keep the lite-07+ behavior. + switch (version) { + case Version.DRAFT_01: + case Version.DRAFT_02: + case Version.DRAFT_03: + case Version.DRAFT_04: + case Version.DRAFT_05: + case Version.DRAFT_06: + return false; + default: + return true; + } +} + /** Whether ANNOUNCE_START and ANNOUNCE_UPDATE may copy a path head or hop-chain tail from a live announcement. Added in lite-07. */ export function hasAnnounceCompression(version: Version): boolean { // Explicitly list older versions so future versions keep the lite-07+ behavior. diff --git a/js/net/src/origin.test.ts b/js/net/src/origin.test.ts index 5ed445a314..4d3ae57a78 100644 --- a/js/net/src/origin.test.ts +++ b/js/net/src/origin.test.ts @@ -3,6 +3,7 @@ import { getter } from "@moq/signals"; import { type Consumer as BroadcastConsumer, Producer as BroadcastProducer } from "./broadcast.ts"; import { StreamCode, StreamError } from "./error.ts"; import { HopSchema, Route } from "./hop.ts"; +import { spreadHash } from "./internal.ts"; import type { Consumer, Table } from "./origin.ts"; import { Producer } from "./origin.ts"; import * as Path from "./path.ts"; @@ -1166,7 +1167,7 @@ test("reject surfaces the error from demand", async () => { const it = handle.requested(); const pending = wireOf(consumer).demand(Path.from("live/cam")); const { value: req } = await it.next(); - const err = new StreamError(StreamCode.NoCapacity, { message: "full" }); + const err = new StreamError(StreamCode.NotFound, { message: "not here" }); req?.reject(err); await expect(pending).rejects.toBe(err); @@ -1262,7 +1263,7 @@ test("advancing requested without settling rejects the previous request", async const second = await it.next(); expect(second.value?.path).toBe(Path.from("live/other")); - await expect(firstDemand).rejects.toMatchObject({ code: StreamCode.NoCapacity }); + await expect(firstDemand).rejects.toMatchObject({ code: StreamCode.Unroutable }); const produced = new BroadcastProducer(); second.value?.accept(produced); await expect(secondDemand).resolves.toBeDefined(); @@ -1387,7 +1388,7 @@ test("a refusal from a superseded route does not end the request", async () => { origin.close(); }); -test("close rejects queued requests with NoCapacity", async () => { +test("close rejects queued requests as unroutable", async () => { const origin = new Producer(); const handle = origin.dynamic(Path.from("live")); const waiting = handle.requested().next(); @@ -1403,7 +1404,6 @@ test("close rejects queued requests with NoCapacity", async () => { req?.accept(new BroadcastProducer()); await settle(); expect(request.active.peek()).toBeUndefined(); - expect(Number(new StreamError(StreamCode.NoCapacity).code)).toBe(0x30); request.close(); origin.close(); @@ -1424,6 +1424,48 @@ test("announced filters by arbitrary patterns and reports captures", async () => origin.close(); }); +test("the spread hash matches rs/moq-net byte for byte", () => { + expect(spreadHash("pool/job-0", [10n])).toBe(0xefb5e20a66101c32n); + expect(spreadHash("pool/job-0", [11n])).toBe(0x0eb0a91370ff6653n); +}); + +test("an equal-cost pool spreads its paths the same way on every node", async () => { + const workers = [10n, 11n, 12n, 13n].map((id) => HopSchema.parse(id)); + const paths = Array.from({ length: 64 }, (_, i) => Path.from(`pool/job-${i}`)); + + // The worker each path resolves to on an origin whose pool arrived in `order`. + async function winners(order: typeof workers): Promise { + const origin = new Producer(); + const served = new Map(); + const handles = order.map((hop) => { + const handle = wireOf(origin).receive(Path.from("pool"), { hops: [hop], cost: 3n }); + void (async () => { + for await (const request of handle.requested()) { + served.set(request.path, hop); + request.accept(new BroadcastProducer()); + } + })(); + return handle; + }); + const requests = paths.map((path) => origin.request(path)); + await settle(); + for (const request of requests) request.close(); + for (const handle of handles) handle.close(); + origin.close(); + return paths.map((path) => served.get(path) ?? -1n); + } + + const forward = await winners(workers); + const reverse = await winners([...workers].reverse()); + expect(reverse).toEqual(forward); + + // Not an assertion about any two paths, which a correct hash may put on one worker: + // only that the set does not pile onto a few. + for (const worker of workers) { + expect(forward.filter((hop) => hop === worker).length).toBeGreaterThanOrEqual(paths.length / 16); + } +}); + test("a rooted producer shares the table and enforces its pattern union", async () => { const origin = new Producer(); const scoped = origin.scope(Path.from("tenant"), new Path.Patterns([Path.Pattern.parse("room/*")])); diff --git a/js/net/src/origin.ts b/js/net/src/origin.ts index f5b9438238..bb41fbcccb 100644 --- a/js/net/src/origin.ts +++ b/js/net/src/origin.ts @@ -15,7 +15,7 @@ import * as announce from "./announced.ts"; import * as broadcast from "./broadcast.ts"; import { StreamCode, StreamError } from "./error.ts"; import { isAnonymous, Route, routesEqual } from "./hop.ts"; -import { hiddenBelow, hooks, scopeCaptures, scopeHead, scopeOverlaps } from "./internal.ts"; +import { hiddenBelow, hooks, scopeCaptures, scopeHead, scopeOverlaps, spreadHash } from "./internal.ts"; import * as Path from "./path.ts"; import { type Advertised, type Advertisements, registerWire, wireOf } from "./wire.ts"; @@ -225,8 +225,16 @@ function compareRoutes(a: Route, b: Route): number { return 0; } -/** The preferred of `entries` (newest first) not skipped: the best route, then fewest hops, then newest. */ -function preferredEntry(entries: readonly RouteEntry[], skip?: (entry: RouteEntry) => boolean): RouteEntry | undefined { +/** + * The preferred of `entries` (newest first) for resolving `path`, not skipped: the best route, + * then fewest hops, then the lowest {@link spreadHash}, then newest. `path` is the requested + * path for a request, or the prefix itself for an advertisement. + */ +function preferredEntry( + path: Path.Valid, + entries: readonly RouteEntry[], + skip?: (entry: RouteEntry) => boolean, +): RouteEntry | undefined { let best: RouteEntry | undefined; for (const entry of entries) { if (skip?.(entry)) continue; @@ -236,7 +244,13 @@ function preferredEntry(entries: readonly RouteEntry[], skip?: (entry: RouteEntr } const a = entry.route.peek(); const b = best.route.peek(); - const order = compareRoutes(a, b) || a.hops.length - b.hops.length; + let order = compareRoutes(a, b) || a.hops.length - b.hops.length; + // Hashed only on a tie, so the common single-route prefix never pays for it. + if (order === 0) { + const ha = spreadHash(path, a.hops); + const hb = spreadHash(path, b.hops); + order = ha < hb ? -1 : ha > hb ? 1 : 0; + } if (order < 0) best = entry; } return best; @@ -247,8 +261,8 @@ function received(entry: RouteEntry): boolean { return !entry.originated; } -function noCapacity(): StreamError { - return new StreamError(StreamCode.NoCapacity, { message: "no capacity" }); +function unroutable(): StreamError { + return new StreamError(StreamCode.Unroutable, { message: "unroutable" }); } /** A served route from {@link Producer.dynamic}: the queue a handler drains. */ @@ -320,7 +334,7 @@ class ServeState { close(abort?: Error): void { if (this.closed.peek() !== undefined) return; - const err = abort ?? noCapacity(); + const err = abort ?? unroutable(); this.closed.set(err); const queued = [...this.pending.values()]; this.pending.clear(); @@ -542,6 +556,7 @@ class OriginState { for (const [prefix, entries] of this.routes.peek() ?? []) { if (!Path.hasPrefix(prefix, path)) continue; const entry = preferredEntry( + path, entries, (candidate) => !candidate.scope.matches(path) || (skip?.(candidate) ?? false), ); @@ -1521,7 +1536,7 @@ export class Consumer { * requests beneath it. * * Drop it (or {@link close}) to retract the route and reject anything still waiting - * with {@link StreamCode.NoCapacity}. {@link update} re-prices it in place. + * with {@link StreamCode.Unroutable}. {@link update} re-prices it in place. * * @public */ @@ -1567,7 +1582,7 @@ export class Dynamic { if (!server) return; let current: Request | undefined; const drop = () => { - current?.reject(noCapacity()); + current?.reject(unroutable()); current = undefined; }; try { diff --git a/js/net/src/wire.ts b/js/net/src/wire.ts index 9065dc3561..bb9af83628 100644 --- a/js/net/src/wire.ts +++ b/js/net/src/wire.ts @@ -10,7 +10,7 @@ import type { Dispose, Getter } from "@moq/signals"; import type * as broadcast from "./broadcast.ts"; import type { Consumer as GroupConsumer } from "./group.ts"; -import type { Route } from "./hop.ts"; +import type { Hop, Route } from "./hop.ts"; import type * as origin from "./origin.ts"; import type * as Path from "./path.ts"; import type * as track from "./track.ts"; @@ -21,6 +21,13 @@ export interface Broadcast { resolveTrackInfo(name: string): Promise; fetchGroup(name: string, sequence: number, options?: track.FetchGroupOptions): Promise; requested(): Promise; + /** + * The origin serving this broadcast, named in SUBSCRIBE_OK and FETCH_OK: the one an + * upstream reply named, or else a random one generated once for the broadcast. + */ + origin(): Hop; + /** Record the origin an upstream reply named for this broadcast. */ + name(origin: Hop): void; } /** The protocol-facing operations behind an origin producer. */ diff --git a/js/watch/src/broadcast.test.ts b/js/watch/src/broadcast.test.ts index 4debf75790..cd7d19d9c7 100644 --- a/js/watch/src/broadcast.test.ts +++ b/js/watch/src/broadcast.test.ts @@ -343,6 +343,109 @@ describe("cross-broadcast renditions", () => { }); }); +// A derived rendition produced only on demand by a service that claims a covering prefix (a +// wildcard) instead of announcing each path. Nothing announces the rendition until something +// subscribes to it, so the player must list it from the claim alone. The rendition is a sibling: +// one beneath its source is already covered by the source's own announcement. The service prefix +// is hidden, as a deployment keeps it out of listings, so the claim is only seen by opting in. +describe("wildcard renditions", () => { + const name = Path.from("live/foo.hang"); + const derived = Path.from(".pro/transcode/foo.hang"); + const rel = Path.normalizeRelative("../.pro/transcode/foo.hang"); + + const watch = (owner: Origin.Producer) => + new Broadcast({ + origin: owner, + name, + enabled: true, + catalogFormat: "manual", + catalog: { + video: { + renditions: { + source: video("avc1.64001e"), + transcode: video("avc1.640028", rel), + }, + }, + } as Catalog.Root, + }); + + it("lists and demands a rendition only a wildcard covers", async () => { + const owner = new Origin.Producer(); + const main = publish(owner, name); + const source = watch(owner); + const effect = new Effect(); + + try { + await settle(); + expect(videoRenditions(source)).toEqual(["source"]); + + const worker = owner.dynamic(Path.from(".pro/transcode")); + const requests = worker.requested(); + await settle(); + expect(videoRenditions(source)).toEqual(["source", "transcode"]); + + // Selecting the rendition is what asks the worker to start, before anything announces it. + let active: Moq.Broadcast.Consumer | undefined; + effect.run((nested) => { + active = source.relativeBroadcast(nested, rel); + }); + const { value: request } = await requests.next(); + expect(request?.path).toBe(derived); + expect(active).toBeUndefined(); + + const produced = new Moq.Broadcast.Producer(); + produced.createTrack("video").writeString("frame"); + request?.accept(produced); + await settle(); + expect(active).toBeDefined(); + const track = active?.track("video").subscribe().ordered(); + expect(await track?.readString()).toBe("frame"); + track?.close(); + + // The worker announcing the path it now serves leaves the rendition where it was. + const concrete = owner.createBroadcast(derived); + concrete.announce(); + await settle(); + expect(videoRenditions(source)).toEqual(["source", "transcode"]); + expect(active).toBeDefined(); + + concrete.close(); + worker.close(); + } finally { + effect.close(); + source.close(); + main.close(); + owner.close(); + } + }); + + it("hides the rendition once the last covering wildcard is withdrawn", async () => { + const owner = new Origin.Producer(); + const main = publish(owner, name); + const source = watch(owner); + + try { + // A catch-all (an archive) and a narrower pool both cover the derived path. + const archive = owner.dynamic(Path.empty()); + const worker = owner.dynamic(Path.from(".pro/transcode")); + await settle(); + expect(videoRenditions(source)).toEqual(["source", "transcode"]); + + worker.close(); + await settle(); + expect(videoRenditions(source)).toEqual(["source", "transcode"]); + + archive.close(); + await settle(); + expect(videoRenditions(source)).toEqual(["source"]); + } finally { + source.close(); + main.close(); + owner.close(); + } + }); +}); + describe("manual catalog", () => { it("republishes a manual catalog mutated in place", async () => { const owner = new Origin.Producer(); diff --git a/js/watch/src/broadcast.ts b/js/watch/src/broadcast.ts index 1064a0708c..99e8c9b801 100644 --- a/js/watch/src/broadcast.ts +++ b/js/watch/src/broadcast.ts @@ -194,7 +194,9 @@ export class Broadcast { const origin = effect.get(this.in.origin); if (!origin) return; - const announced = origin.announced(); + // Hidden routes count: a service claim under a `.`-named prefix is kept out of listings, but + // it still covers the renditions it would produce, and this set is never shown to anyone. + const announced = origin.announced(Path.Pattern.all(), { hidden: true }); effect.cleanup(() => announced.close()); this.#announced.set(new Set()); @@ -223,7 +225,9 @@ export class Broadcast { // Whether `path` is covered by an announced route, for `relativeBroadcast`'s // cross-broadcast refs. Announcements are prefix routes, so a route at "room/" covers - // "room/alice/cam.hang" without naming it. Opens the announcement stream on first use. + // "room/alice/cam.hang" without naming it. That is how a rendition produced only on demand + // gets selected: its service claims a covering prefix, and nothing announces the exact path + // until this subscribes. Opens the announcement stream on first use. // The blind cases (announcement gate off, no discovery) never reach here; see `#relativeTarget`. #isPathAnnounced(effect: Effect, path: Moq.Path.Valid): boolean { this.#wantAnnounced.set(true); diff --git a/quest/m0/wildcard/README.md b/quest/m0/wildcard/README.md index 6bc8b6153b..3a55d3e74a 100644 --- a/quest/m0/wildcard/README.md +++ b/quest/m0/wildcard/README.md @@ -1,20 +1,21 @@ -# [L] Wildcard advertisements +# [S] Wildcard advertisements ## Goal A service claims the prefix it could serve rather than enumerating every broadcast under it; the client library filters that claim against the pattern interest the caller asked for, so nothing on the wire spells a -wildcard. A claim is priced at what starting the work would cost. Specificity wins first: a concrete -claim shadows a wildcard regardless of cost, and prices compete within the -same specificity tier. A terminal concrete refusal does not fall through to -a catch-all; its claim must be withdrawn. Retracting a wildcard stops new -work without shedding what is already running. - -Three workloads need this, and they are the three pattern shapes. A transcode +wildcard. A claim is priced at what starting the work would cost. The longest +covering prefix wins first: a concrete announcement shadows a broader claim +regardless of cost, and prices compete only among claims of the same prefix. +A terminal concrete refusal does not fall through to a catch-all; its claim +must be withdrawn. Retracting a claim stops new work without shedding what is +already running. + +Three workloads need this. A transcode worker today announces a standby derivative for every matching live broadcast, -so announcements scale as workers times broadcasts; with a suffix pattern -`**/transcode.pro` it advertises once for the whole fleet. A chat backend +so announcements scale as workers times broadcasts; claiming one service +prefix, it advertises once for the whole fleet. A chat backend cannot enumerate at all: rooms exist independently of any broadcast, and the subtree pattern `/chat/**` expresses them. An archive serving recordings over FETCH wants to say "if nobody is publishing this live, I have it", which @@ -33,15 +34,12 @@ across the fleet in resident memory. Decided in [#3770](https://github.com/moq-dev/moq/pull/3770): publishing is prefix-only on every wire and patterns never leave the token or the client library. `dynamic(prefix, route)` -advertises a prefix; a suffix or catch-all claim is expressed as the -widest prefix that covers it (`**` is the root) and the request is the -authority, so the advertise half of this questline is re-scoped to prefix -claims resolved against pattern interest. The three workloads above still -hold: the transcoder claims the root and refuses what it will not serve. -Resolve and Demand are additive and land on main. Both are done on the line -branch (#4050 re-resolves on a refusal; 9d059b1b9 lists and demands covered -renditions in the browser player), so main no longer lists them; the line -branch moves to `quest/m0/wildcard/README` to match this path. +advertises a prefix; the catch-all claim is the root prefix, and the request +is the authority, so the advertise half of this questline is re-scoped to +prefix claims resolved against pattern interest. The three workloads above +still hold: the transcoder claims its service prefix +([Where derived output lives](#where-derived-output-lives)), and the archive +claims the root and refuses what it does not have. ### What already exists, and what does not @@ -56,33 +54,31 @@ same exact containment check. `Cost { warm, cold }` [#2925](https://github.com/moq-dev/moq/pull/2925). [moq#3225](https://github.com/moq-dev/moq/pull/3225) moved a long way toward -this. An announcement carries a `Pattern` covering a set of paths. Rust -`announce::Update.pattern` and TypeScript `Announce.Update.pattern` use the -matcher directly, so callers explicitly select prefix-shaped claims when -they need a concrete broadcast path. - -The routing table exists too. `Consumer::request_broadcast` resolves a local -broadcast first, then `best_server`: the longest covering prefix, filtered by -the requester's excluded hop, ordered by `route_order` -(`rs/moq-net/src/model/origin.rs:633`), served on demand by the session that -announced it and cached per prefix in `ServeState.served` (`:764`). That is the split-horizon-safe -lookup the old `origin::Dynamic` could not provide, and it is what -resolve extends rather than replaces. - -Request resolution, by contrast, is still prefix-only (`best_server` in -`rs/moq-net/src/model/origin.rs`). The pattern matcher itself exists: +this, and #3770 settled the wire: an announcement carries a path prefix on +every protocol, and a consumer filters announced paths against its pattern +interest locally. + +Request resolution exists too. `Consumer::request_broadcast` mints a front per +path (`rs/moq-net/src/model/front.rs`) that selects through `best_route`: a +local broadcast first, then the longest covering prefix, filtered by the +requester's excluded hop and ordered by `route_order`, whose hash is keyed on +the requested path so one prefix's pool shares its paths. A refusal from that +tier is final, a front resumes only onto a source whose SUBSCRIBE_OK names the +same origin (the route's first hop on wires older than lite-07), and FETCH +resolves the same way. The pattern matcher itself exists: `moq_net::{Pattern, Patterns, Segment}` and `Path.Pattern` / `Path.Patterns` in `js/net/src/path.ts` own the shared matching, containment, -specificity, and rebasing advertisements reuse. +specificity, and rebasing tokens and filters reuse. What is genuinely missing, beyond patterns themselves, is content identity. Announcement `Epoch` was specified into lite-06 by [#2611](https://github.com/moq-dev/moq/pull/2611), never implemented, and removed from the draft by #3225, which retired `draft-lcurley-moq-broadcast` with it. [moq#3312](https://github.com/moq-dev/moq/pull/3312) restored per-path identity -from the route's first hop, reversing #3225's no-splice rule, and this -questline builds its collision handling on that rather than on a generation -field. +from the route's first hop, reversing #3225's no-splice rule, and lite-07 moved +it to the origin a SUBSCRIBE_OK or FETCH_OK names, since a pool's one route labels +many origins. This questline builds its collision handling on that rather than +on a generation field. ### Decisions @@ -91,101 +87,71 @@ field. dialect is what tokens and the consume-side filter use, matched by the shared matcher, so nothing resembles a second grammar and nothing on the wire spells a wildcard. -- **Most specific pattern wins, and its refusal is final.** This is the rule - routing already follows: `best_server` filters to the longest covering prefix - before it compares cost, and the lite draft says the same, matching - longest-prefix-match wherever it appears. When several patterns match one - path, only the tier selected by the matcher's shared structural specificity is - consulted; equal-specificity patterns - form one pool that cost and the request hash order. A terminal refusal from - the winning tier IS the answer and never falls through to a less specific - pattern, so a transcoder refusing a path does not leak the request to the - archive's catch-all, and one unserved path still costs one round trip. The - capacity re-resolution below stays within the tier, refuser excluded. The - accepted consequence: an offline derivative (a recording of - `foo.hang/transcode.pro`) is not reachable through the catch-all, because - the more specific transcode pattern shadows it. -- **A wildcard is a POOL, not a competitor.** Several advertisers of one - pattern is the normal state, not a hazard: every transcode worker advertises - `**/transcode.pro` and takes a share. What distributes them is a +- **Longest prefix wins, and its refusal is final.** `best_route` filters to + the longest covering prefix before it compares cost, and the lite and + cluster drafts say the same. Advertisers of one prefix form one pool that + cost and the request hash order. A refusal from the winning tier IS the + answer and never falls through to a shorter prefix or to another + advertiser, so a transcoder refusing a path does not leak the request to + the archive's catch-all, and one unserved path costs one round trip. The + accepted consequence: an offline derivative is not reachable through the + catch-all while a longer prefix covers it. +- **A claim is a POOL, not a competitor.** Several advertisers of one + prefix is the normal state, not a hazard: every transcode worker claims the + same prefix and takes a share. What distributes them is a deterministic hash of the REQUESTED path against each advertiser, so distinct - paths spread rather than one advertiser winning the whole pattern. + paths spread rather than one advertiser winning the whole prefix. Distribution is the requirement, not any particular pair: a correct hash may legitimately rank the same advertiser first for two given paths, so what must hold is that a large path set spreads and that one path always resolves the same way. Cost orders the pool first, which keeps work local and makes a distant advertiser the overflow rather than an equal peer. -- **A wildcard is priced, not special-cased.** Within a tier, route selection - stays one comparison on one metric. Concrete-versus-wildcard is not decided - by price at all: a concrete claim is maximally specific, so "most specific - wins" above already shadows every pattern behind it at any cost. The +- **A claim is priced, not special-cased.** Within a tier, route selection + stays one comparison on one metric. Concrete-versus-claim is not decided + by price at all: a concrete announcement is the longest prefix, so "longest + prefix wins" above already shadows every claim behind it at any cost. The accepted consequence follows from that rule's finality: a live session's - concrete claim shadows a healthy wildcard pool even when its service is + concrete claim shadows a healthy pool even when its service is broken, its terminal refusal does not fall through, and the shadow lasts exactly as long as the claiming session that carries it. - The seed still has a floor, because standby and running claims of equal - specificity do meet: a standby concrete claim (`with_cost(1000)` is the + The seed still has a floor, because standby and running claims of the same + prefix do meet: a standby concrete claim (`with_cost(1000)` is the existing per-broadcast convention) shares a tier with a running publisher's - concrete announcement. The + concrete announcement and with warm-advertise's exact-path warm routes. The floor MUST exceed the deployment's enforced maximum charged-link count times its enforced maximum link cost (32 links at cost at most 5 gives a bound of 160, with producing origins seeded at 0), or a nearby standby outranks a distant running copy and the mesh starts a second encode of a stream it is already serving. That floor replaces the ad-hoc standby bias the moq.pro (downstream) transcode worker carries today. -- **One cost varint, not the pair.** `Cost` is `{ warm, cold }` because a relay - that is carrying a broadcast discounts the warm half. A wildcard carries - nothing and can never be warm, so the two halves are provably equal and the - message carries one value. `From` already means exactly this. -- **A wildcard is a capability, not an inventory.** It advertises what the +- **A claim is a capability, not an inventory.** It advertises what the sender could serve, never that a given path exists. Refusal is how a specific path is denied. This is why an over-claiming advertisement is not a defect: - the catch-all `**` is legal, and answering "not that one" is the mechanism. -- **Containment against the publish scope is what authorization checks.** An - advertised pattern MUST be contained by the sender's granted patterns (the - matcher's containment check). This handles literal-headed and leading-star - patterns identically and refuses any attempted widening rather than clamping - it. Fleet-wide services use the cluster identity; a customer service may - advertise only the exact set its own v1 grant contains. -- **Wildcards are visible to subscribers.** A subscriber sees every pattern - matching under its scope, rebased by the matcher's exact set-valued operation, - and duplicates combine into one. That is the point: it tells a client it may subscribe to - matching paths, and its withdrawal tells the client the capability is gone. - This is what makes a lazily-produced rendition discoverable without the - composer waiting for an announcement that only demand would produce. The - browser player currently enforces the opposite (`js/watch`'s - `#isPathAnnounced` hides a catalog rendition with no exact-path - announcement); demand makes a covering wildcard count as - availability there. -- **Refusal is a typed stream reset, with no negative cache.** An advertiser - resets a subscribe it will not serve, and the reset carries which KIND of - refusal it is (`Error::to_code` already puts a typed code on the wire; the - capacity code is NO_CAPACITY, 0x30, in moq-lite's own 48-63 range). - - A capacity refusal is unavoidable: an advertiser's capacity and a relay's view - of it are separated by at least half a round trip, so a retraction and a - request for the slot it just gave away WILL cross, at a rate of request rate - times retraction rate times RTT. Only that code permits ONE re-resolution, - with the refusing advertiser excluded from it. The exclusion is what makes the - retry safe, NOT the retraction arriving first: the reset and the retraction - travel independently, so re-resolution may pick another advertiser, and may - equally find none and return unroutable. Both are correct outcomes. - - Every other refusal is terminal and propagates: a path no rule covers, an - unauthorized one, one that does not exist. Scanning unserved paths therefore - still costs one round trip per path, and this is not a fallback list: there - is no walk down the candidates, only one re-resolution. Classification is - EXPLICIT, per the repository's retry policy: an unrecognized or bare reset is - permanent, so a new refusal mode surfaces instead of quietly joining a retry - loop. No negative cache either way; rate limiting stays with the advertiser - and the per-project auth gate. + claiming the root is legal, and answering "not that one" is the mechanism. +- **Claims are visible to subscribers.** A claim is an ordinary prefix + announcement, so it tells a client it may subscribe beneath it, and its + withdrawal tells the client the capability is gone. The browser player's + gate (`js/watch`'s `#isPathAnnounced`) lists a catalog rendition under any + covering prefix, so a lazily-produced rendition is discoverable without an + announcement that only demand would produce. +- **Every refusal is terminal, and capacity lives in the route.** An + advertiser resets a request it will not serve, and the relay propagates it + without retrying another advertiser, on moq-lite and moq-transport alike, + so the protocols need no capacity code. An advertiser sheds load by + withdrawing or re-pricing its route before it runs out, leaving headroom + for requests already in flight, since a withdrawal and a request for the + slot it gave away can cross. A request that still loses the race fails, + and the subscriber re-requests against the updated table. No negative + cache; rate limiting stays with the advertiser and the per-project auth + gate. - **A double claim is settled by route identity, not by a lease.** Two relays can hash one path to different workers before either concrete announcement propagates, and both land at the SAME literal path. Whichever route wins selection serves it, and a consumer moves between them only when the winner's - identity is preserved, per the first-hop resume rule ([moq#3312](https://github.com/moq-dev/moq/pull/3312)); two distinct workers are two identities, + identity is preserved, per the resume rule (the origin a SUBSCRIBE_OK + names on lite-07, the route's first hop before it); two distinct workers are two identities, so the loser's subscribers end and resubscribe rather than being spliced onto - another worker's frames mid-group. Wildcard routing invents neither a lease + another worker's frames mid-group. Claim routing invents neither a lease nor a generation. This is weaker than the retired `Epoch` design, which could declare two workers' output interchangeable and splice between them; a service that needs that guarantee has to carry it in its own media contract, not in @@ -199,71 +165,40 @@ field. ### Where derived output lives -Suffix matching lets a contribution be published where it is addressed, a -descendant of its source. This is the moq.pro (downstream) deployment shape, -and it is what the suffix pattern form exists for: - -```text -pid/foo.hang source -pid/foo.hang/catalog.pro combined catalog, edge-composed -pid/foo.hang/transcode.pro the transcode contribution -pid/foo.hang/transcribe.pro the transcription contribution -``` - -The `.pro` segment suffix is both the routed pattern and the platform-output -marker: `**/transcode.pro` routes every project's transcode demand to the -worker pool, and a segment ending in `.pro` is the one predicate every source -rule matcher excludes, so platform output is never recursively transcoded or -recorded. - -The rejected alternative publishes contributions at mirrored paths in reserved -namespaces (`.transcode//...`) hidden by an origin-consumer overlay, -because prefix-only matching needs the variable part of a path trailing. That -overlay is not a view transform: `pid/foo` and `.transcode/pid/foo` are -separate tree leaves with separate broadcast fronts, so it has to build a -logical front across roots that re-owns route selection, content identity, the -split-horizon guard, and splicing. The suffix pattern needs none of it while -keeping what the mirror buys: - -- **The grant needs no transform.** The customer addresses - `foo.hang/transcode.pro`, a descendant of `foo.hang`, so an existing grant - covers it by ordinary segment-aware prefix. No companion-grant rule, no - atomic `.pro/` scope, and nothing minted differently, which matters because - customer-issued tokens are minted by integrations the platform does not - control. -- **Metering is untouched.** The published path is rooted at `pid`, so the - platform's egress metering sees the customer path with no special case at - all. -- **The wildcard is fleet-wide.** The suffix is project-agnostic, so a worker - advertises once for every project rather than once per project, which is what - removes project discovery entirely. -- **Takeover is single-front.** A worker's concrete announcement lands at the - literal path the wildcard served, so wildcard-versus-concrete and - worker-versus-worker collisions are ordinary route selection at one tree node, - not a cross-root front. - -What the mirror buys and this deliberately gives up: a customer holding -`publish: ["pid/"]` CAN publish `foo.hang/transcode.pro` themselves, competing -with or forging platform output. Both then resolve at one path, cost decides, -and a live customer broadcast beats the worker's standby seed. That is confined -to their own namespace, self-sabotage of their own catalog, never another -project's, and is cheaper to allow and document than a reserved-name registry -or a token transform. The mirror's SUBSCRIBE-only overlay asymmetry existed to -prevent exactly this and goes with it. - -The archive is the same shape at the source path itself: a recording IS the -broadcast, served from storage through the catch-all pattern. A wildcard names -no generation, so a client that must distinguish recording generations reads -the catalog's archive entry ([archive](/quest/m1/archive/README.md)) rather -than announce state. +A prefix claim needs the variable part of a path trailing, so a fleet-wide +service claims its own prefix and mirrors the source path beneath it +(`.pro/transcode//foo.hang`, moq.pro's convention) rather than publishing +beneath the source. The source's catalog reaches the contribution through a +cross-broadcast reference. + +The leading `.` is deliberate. Existing customers on moq-lite-06 or older must +never see `.pro/` broadcasts, which could confuse their business logic. Those +versions cannot opt into hidden routes, so the relay never announces them +there. Hidden routes are a moq-lite-07 feature, so the player's covering check +opts into them and sees a claim only when lite-07 is negotiated; the token's +scope must still reach them. A customer who wants transcodes upgrades, or +subscribes to the explicit `.pro//...` path, which works on any +version. Grants and metering are the deployment's; moq.pro's are in its +[wildcard questline](https://github.com/moq-dev/moq.pro/blob/main/quest/m2/wildcard/README.md). + +The archive serves the source path itself: a recording IS the broadcast, +served from storage through the root claim, and a live publisher's concrete +announcement shadows it. A claim names no generation, so a client that must +distinguish recording generations reads the catalog's archive entry +([archive](/quest/m1/archive/README.md)) rather than announce state. + +## Required + +- moq-lite-07 is finalized and negotiated by default, no longer + `moq-lite-07-wip`, so the player's covering check sees hidden claims ## Related - [path-patterns](/quest/m1/path-patterns.md) - owns the pattern dialect - and the shared matcher advertisements reuse -- [archive](/quest/m1/archive/README.md) - an archive advertises the catch-all - pattern, and its catalog names the generations a wildcard cannot + and the shared matcher tokens and filters reuse +- [archive](/quest/m1/archive/README.md) - an archive claims the root, and its + catalog names the generations a claim cannot - [Cluster routing](/quest/m1/cluster-routing.md) - origin selection by cost - with an HRW tie-break, built on this line's specificity -- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - derived output moves under the - source's `@` segment, which the suffix patterns still match + with an HRW tie-break, built on this line's longest-prefix rule +- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - derived output + mirrors the source path, `@` segment included diff --git a/quest/m1/README.md b/quest/m1/README.md index 93cdfc7aad..816f4d0644 100644 --- a/quest/m1/README.md +++ b/quest/m1/README.md @@ -65,6 +65,8 @@ transport, benchmark tooling); worktrees isolate commits, not semantics. - [#2991](/quest/m1/2991-net-coalesce-dynamic-tracks-and-preserve-sequences-across.md) - one dynamic producer per track name in both languages, with the sequence namespace surviving a replacement - [JavaScript FETCH](/quest/m1/js-fetch.md) - generic on-demand group serving and IETF FETCH for browser publishers - [Archive](/quest/m1/archive/README.md) - record selected tracks to any object_store and replay them over FETCH or derived HLS; the catalog entry and format may break in place, since no archives exist +- [Cluster origin reply](/quest/m1/cluster-origin.md) - a moq-transport downstream of a spreading relay never splices two pool members, and the cluster draft says how it tells them apart +- [Verified route upgrade](/quest/m1/front-upgrade.md) - a relay moves a live subscription to a cheaper route once that route's reply names the same origin - [Tooling](/quest/m1/tooling/README.md) - justfiles become a one-line menu over `sh/`, one impact map scopes CI, and every workflow step runs a recipe - [Path patterns](/quest/m1/path-patterns.md) - one matcher for every predicate over broadcast paths: tokens, origins, interest - [In-band auth](/quest/m1/auth/README.md) - a session tells its peer what it may publish and subscribe to, unions tokens presented in band, and fails loud on an out-of-scope publish diff --git a/quest/m1/broadcast-epoch/README.md b/quest/m1/broadcast-epoch/README.md index 9e0ac40d40..75cd10e878 100644 --- a/quest/m1/broadcast-epoch/README.md +++ b/quest/m1/broadcast-epoch/README.md @@ -40,10 +40,10 @@ Decided: bare `foo`, since a route covers its descendants, not its parent. A bare-name viewer behind one needs a publisher that opts out with the raw prefix route. Document this rather than promise it works. -- Derived output lives under the epoch it came from - (`pid/foo.hang/@e/transcode.pro`), so nested epochs must parse. This moves - the [wildcard](/quest/m0/wildcard/README.md) line's derived-output example - down one segment, and its suffix patterns still match. +- Derived output mirrors the epoch it came from + (`.pro/transcode/pid/foo.hang/@e`, per the + [wildcard](/quest/m0/wildcard/README.md) line's derived-output layout), so + the service's prefix claim still covers it. This README owns: diff --git a/quest/m1/cluster-origin.md b/quest/m1/cluster-origin.md new file mode 100644 index 0000000000..c887da5b0e --- /dev/null +++ b/quest/m1/cluster-origin.md @@ -0,0 +1,38 @@ +# [M] Cluster origin reply + +## Goal + +A relay that spreads a pool's paths never lets a moq-transport downstream +splice one pool member's content onto another's, and +`draft-lcurley-moq-cluster` says how a downstream learns which member serves a +path. + +## Plan + +Spreading keys route selection's hash on the requested path, so equal-cost +advertisers of one prefix share its paths, and a relay advertises one route for +the whole pool. On moq-lite-07 the SUBSCRIBE_OK and FETCH_OK replies name the +serving origin, and a relay splices a failover only between replies naming the +same one. The cluster extension (moq-transport 17+) has no such field, so a +downstream relay still pins failover to the advertisement's first Hop ID, which +labels the pool rather than the member serving the path. The draft's +"Several Publishers of One Namespace" section still requires that the first +Hop ID downstream name the publisher whose Objects flow, which a spreading +relay breaks. + +Two ways out, to settle before implementing: + +- Name the serving origin in a moq-transport reply as a cluster-extension + parameter, mirroring lite-07, and stitch on it. Recommended: one identity + model on every wire. +- Keep the label truthful instead: toward a downstream that cannot learn the + origin, do not spread (or advertise per path). + +Either way the test is the one lite-07 has: two workers behind a pool relay +behind a moq-transport downstream relay, and killing the serving worker ends +the downstream subscription rather than splicing the survivor's objects. + +## Related + +- [Wildcard](/quest/m0/wildcard/README.md) - spreading and the lite-07 reply + origin landed there diff --git a/quest/m1/front-upgrade.md b/quest/m1/front-upgrade.md new file mode 100644 index 0000000000..b8dff150e4 --- /dev/null +++ b/quest/m1/front-upgrade.md @@ -0,0 +1,37 @@ +# [M] Verified route upgrade + +## Goal + +A relay whose front learned its origin from a reply moves a live subscription +to a cheaper route once that route's source proves it serves the same origin, +instead of staying on the costlier route until it fails. + +## Plan + +On moq-lite-07 a front's identity is the origin its source's SUBSCRIBE_OK or +FETCH_OK names, not the route's first hop, because one route can lead to many +pool members. The front only learns a new route's origin once that source +replies, so today it stays on its live source when a better route appears +(`Pin::Stay` in `rs/moq-net/src/model/front.rs`) and only moves on failover. +That keeps content from being spliced across origins but leaves traffic on a +route that may no longer be the cheapest. + +The move to make: when a cheaper route appears, ask it before letting go of the +live source, and splice over at a group boundary only once its reply names the +front's origin. A reply naming another origin leaves the front where it is, +with nothing delivered from the candidate. Things to watch: + +- The candidate's content is held until admitted (the copy's provenance), so + asking early must not deliver anything or start demand the live source + already covers longer than needed. +- Churn: a route that keeps flapping should not keep opening and dropping + upstream subscriptions. +- Measure the cost of a verification round trip against staying put, and + benchmark it across routes and subscribers. + +## Related + +- [Wildcard](/quest/m0/wildcard/README.md) - the reply-named identity landed + there +- [Cluster origin reply](/quest/m1/cluster-origin.md) - the same identity on + moq-transport diff --git a/rs/moq-gst/src/sink/pad.rs b/rs/moq-gst/src/sink/pad.rs index f9b68dd37b..cd044ae8e2 100644 --- a/rs/moq-gst/src/sink/pad.rs +++ b/rs/moq-gst/src/sink/pad.rs @@ -29,12 +29,12 @@ enum PadState { /// Where a pad's buffers land: a codec importer, a subtitle track, or opaque data. /// -/// Both payloads are large (a codec importer, a container producer), so each is boxed to keep the -/// enum small. +/// Every payload is large (a codec importer, a container producer, a track producer), so +/// each is boxed to keep the enum small. enum Sink { Media(Box), Text(Box), - Opaque(moq_net::track::Producer), + Opaque(Box), } /// An audio or video pad, published through a codec importer. @@ -318,7 +318,7 @@ impl Pad { .with_context(|| format!("cannot reserve track {name}"))?; // Followed at the live edge, so it keeps the default retention the media helper raises. let info = moq_net::track::Info::default().with_timescale(moq_net::Timescale::MICRO); - self.track = Some(Sink::Opaque(request.accept(info))); + self.track = Some(Sink::Opaque(Box::new(request.accept(info)))); self.caps = Some(caps.clone()); return Ok(name); } diff --git a/rs/moq-net/src/error.rs b/rs/moq-net/src/error.rs index 148a857f64..8a2f844f02 100644 --- a/rs/moq-net/src/error.rs +++ b/rs/moq-net/src/error.rs @@ -197,7 +197,6 @@ impl StreamError { Self::GoingAway => 0x4, Self::TooFarBehind => 0x5, Self::MalformedTrack => 0x12, - // 0x30 NO_CAPACITY is assigned by other work in this range. Do not reuse it. Self::ControlTimeout => 0x31, Self::GroupTooLarge => 0x32, Self::NotFound => 0x33, @@ -682,8 +681,7 @@ mod tests { assert_eq!(StreamError::from_code(err.to_code()), err, "{err:?} did not round trip"); } - // moq-lite's own 48-63 range, pinned to the draft's table. They stay off 0x30 - // (NO_CAPACITY), which other work assigns. + // moq-lite's own 48-63 range, pinned to the draft's table. for (err, code) in [ (StreamError::ControlTimeout, 0x31), (StreamError::GroupTooLarge, 0x32), diff --git a/rs/moq-net/src/lite/fetch.rs b/rs/moq-net/src/lite/fetch.rs index c098c5990a..c0c3b6d44e 100644 --- a/rs/moq-net/src/lite/fetch.rs +++ b/rs/moq-net/src/lite/fetch.rs @@ -1,7 +1,7 @@ use std::borrow::Cow; use crate::{ - Path, + Hop, Path, coding::{Decode, DecodeError, Encode, EncodeError}, }; @@ -81,6 +81,33 @@ impl Message for Fetch<'_> { } } +/// The publisher's answer on a Fetch Stream, ahead of the FRAME messages: the +/// origin serving the group, which a relay stitches failover on. +/// +/// Lite07+ only; older versions answer with the frames alone. +#[derive(Clone, Debug)] +pub struct FetchOk { + /// [`Hop::UNKNOWN`] names nobody. + pub origin: Hop, +} + +impl Message for FetchOk { + fn decode_msg(r: &mut R, version: Version) -> Result { + if !version.has_origin() { + return Err(DecodeError::Version); + } + let origin = Hop::from_wire(u64::decode(r, version)?)?; + Ok(Self { origin }) + } + + fn encode_msg(&self, w: &mut W, version: Version) -> Result<(), EncodeError> { + if !version.has_origin() { + return Err(EncodeError::Version); + } + self.origin.id().encode(w, version) + } +} + #[cfg(test)] mod test { use super::*; @@ -103,6 +130,21 @@ mod test { Fetch::decode_msg(&mut slice, version).unwrap() } + #[test] + fn fetch_ok_names_the_origin_on_lite07() { + let msg = FetchOk { + origin: Hop::new(42).unwrap(), + }; + let mut buf = Vec::new(); + msg.encode(&mut buf, Version::Lite07).unwrap(); + assert_eq!(buf, [1, 42]); + let got = FetchOk::decode(&mut buf.as_slice(), Version::Lite07).unwrap(); + assert_eq!(got.origin, msg.origin); + + // Older versions answer with the frames alone. + assert!(msg.encode(&mut Vec::new(), Version::Lite06).is_err()); + } + #[test] fn fetch_roundtrips() { for version in [Version::Lite03, Version::Lite04, Version::Lite05] { diff --git a/rs/moq-net/src/lite/publisher.rs b/rs/moq-net/src/lite/publisher.rs index df6a175afc..625c033705 100644 --- a/rs/moq-net/src/lite/publisher.rs +++ b/rs/moq-net/src/lite/publisher.rs @@ -1,5 +1,5 @@ use crate::runtime::Timers as _; -use crate::{SessionError, announce, frame, group, origin, track}; +use crate::{SessionError, announce, broadcast, frame, group, origin, track}; use std::{ collections::HashMap, ops::Bound, @@ -998,6 +998,7 @@ enum SubscribeState { /// Waiting for the model subscription to be confirmed. Confirm { msg: lite::Subscribe<'static>, + broadcast: broadcast::Consumer, subscribing: track::Subscribing, }, /// Streaming groups and datagrams. Boxed: by far the largest state, and the enum @@ -1102,11 +1103,15 @@ impl SubscribeServe { // duplicate demand). let track_consumer = broadcast.track(&msg.track)?; let subscribing = track_consumer.subscribe(subscription).into_inner(); - self.state = SubscribeState::Confirm { msg, subscribing }; + self.state = SubscribeState::Confirm { + msg, + broadcast, + subscribing, + }; } SubscribeState::Confirm { subscribing, .. } => { let track = ready!(subscribing.poll_ok(waiter))?; - let SubscribeState::Confirm { msg, .. } = + let SubscribeState::Confirm { msg, broadcast, .. } = std::mem::replace(&mut self.state, SubscribeState::Decode) else { unreachable!() @@ -1154,6 +1159,10 @@ impl SubscribeServe { version: self.shared.version, timescale, opens: Default::default(), + served: Served { + broadcast: Some(broadcast), + here: self.shared.self_origin, + }, }; let run = TrackRun::new(sub, track, Bounds::from(&msg), track_priority_tx); @@ -1225,6 +1234,7 @@ enum FetchState { /// Waiting for the fetched group. Fetch { msg: lite::Fetch<'static>, + broadcast: broadcast::Consumer, fetching: track::Fetching, }, /// Streaming the group's frames in order. The delta-timestamp baseline @@ -1329,12 +1339,34 @@ impl FetchServe { }, ) .into_inner(); - self.state = FetchState::Fetch { msg, fetching }; + self.state = FetchState::Fetch { + msg, + broadcast, + fetching, + }; } - FetchState::Fetch { msg, fetching } => { + FetchState::Fetch { + msg, + broadcast, + fetching, + } => { let mut group = ready!(kio::Pollable::poll(fetching, waiter))?; - // The response carries no header, so a short run is indistinguishable + // Lite-07 names the origin serving the group ahead of its frames, read + // now that the group resolved (its content flowed, so the front's + // origin is settled). + if self.shared.version.has_origin() { + let served = Served { + broadcast: Some(broadcast.clone()), + here: self.shared.self_origin, + }; + let stream = self.stream.as_mut().expect("stream present"); + stream.writer.buffer(&lite::FetchOk { + origin: served.origin(), + })?; + } + + // The response carries no position, so a short run is indistinguishable // from one that started elsewhere: only serve a range we can cover // exactly. `fetch_group` already positions the consumer, so this is a // belt-and-braces check on a promise the wire can't restate. @@ -2123,6 +2155,37 @@ struct Subscription { timescale: Option, /// The group streams this subscription opened, shared by every group it serves. opens: Arc, + /// Who SUBSCRIBE_OK names as serving it (lite-07+). + served: Served, +} + +/// Names the origin serving a request, for SUBSCRIBE_OK and FETCH_OK: the one the +/// broadcast's front serves, read when the reply goes out (a copy's content only +/// reaches the front once its origin is admitted), or ours for a broadcast that is +/// not route-fed. +#[derive(Clone)] +struct Served { + broadcast: Option, + here: Hop, +} + +#[cfg(test)] +impl Default for Served { + fn default() -> Self { + Self { + broadcast: None, + here: Hop::UNKNOWN, + } + } +} + +impl Served { + fn origin(&self) -> Hop { + self.broadcast + .as_ref() + .and_then(|broadcast| broadcast.origin()) + .unwrap_or(self.here) + } } /// Counts a subscription's group streams for lite-07's SUBSCRIBE_END. @@ -2316,22 +2379,7 @@ impl TrackRun { tracing::debug!(subscribe = self.ctx.id, track = %self.ctx.track_name, sequence, "skipping group with a missing head"); continue; } - if self.emit_range && !self.start_sent { - self.start_sent = true; - // Only the group: the subscriber derives the start frame from - // its own request (see `lite::SubscribeStart`). - stream - .writer - .buffer(&lite::SubscribeResponse::Start(lite::SubscribeStart { - group: sequence, - }))?; - // SUBSCRIBE_OK is an implicit drop of everything below the - // resolved start (the subscriber records it as a permanent - // miss), so a lower group arriving late must not be served - // after all. A widening SUBSCRIBE_UPDATE re-lowers the floor, - // renegotiating the resolved start along with the demand. - self.track.start_at(sequence); - } + self.start(stream, sequence)?; let frame_start = group.index(); tracing::debug!(subscribe = self.ctx.id, track = %self.ctx.track_name, sequence, "serving group"); @@ -2348,7 +2396,10 @@ impl TrackRun { self.children .push(GroupServe::new(self.ctx.clone(), sequence, frame_start, handle, group)); } - Recv::Datagram(datagram) => self.ctx.serve_datagram(datagram), + Recv::Datagram(datagram) => { + self.start(stream, datagram.sequence)?; + self.ctx.serve_datagram(datagram); + } Recv::Boundary(group) => { // The track declared its exclusive final sequence. Forward it now, // even if trailing groups (below `group`) are still in flight, then @@ -2381,6 +2432,30 @@ impl TrackRun { return Poll::Pending; } } + + /// Send SUBSCRIBE_START ahead of the first group served, by stream or datagram: it + /// resolves the start and names the origin, which a lite-07 subscriber needs before + /// it delivers either. + fn start(&mut self, stream: &mut Stream, sequence: u64) -> Result<(), Error> { + if !self.emit_range || self.start_sent { + return Ok(()); + } + self.start_sent = true; + // Only the group: the subscriber derives the start frame from its own request + // (see `lite::SubscribeStart`). + stream + .writer + .buffer(&lite::SubscribeResponse::Start(lite::SubscribeStart { + group: sequence, + origin: self.ctx.served.origin(), + }))?; + // SUBSCRIBE_OK is an implicit drop of everything below the resolved start (the + // subscriber records it as a permanent miss), so a lower group arriving late + // must not be served after all. A widening SUBSCRIBE_UPDATE re-lowers the floor, + // renegotiating the resolved start along with the demand. + self.track.start_at(sequence); + Ok(()) + } } /// Serves one group on its own unidirectional stream: the header, then every @@ -2800,6 +2875,7 @@ mod serve_group_test { version: Version::Lite06, timescale: Some(crate::Timescale::default()), opens: Default::default(), + served: Served::default(), }; let track = track::Producer::new(Arc::new(broadcast::Info::default()), "test", None); @@ -2841,6 +2917,7 @@ mod serve_group_test { version: Version::Lite06, timescale: Some(crate::Timescale::default()), opens: Default::default(), + served: Served::default(), }; let track = track::Producer::new(Arc::new(broadcast::Info::default()), "test", None); @@ -2886,6 +2963,7 @@ mod serve_group_test { version: Version::Lite06, timescale: Some(crate::Timescale::default()), opens: Default::default(), + served: Served::default(), }; let track = track::Producer::new(Arc::new(broadcast::Info::default()), "test", None); @@ -2949,6 +3027,7 @@ mod serve_group_test { version: Version::Lite06, timescale: Some(crate::Timescale::default()), opens: Default::default(), + served: Served::default(), }; let track = track::Producer::new(Arc::new(broadcast::Info::default()), "test", None); @@ -3019,6 +3098,7 @@ mod serve_group_test { version: Version::Lite06, timescale: Some(crate::Timescale::default()), opens: Default::default(), + served: Served::default(), }; let track = track::Producer::new(Arc::new(broadcast::Info::default()), "test", None); @@ -3065,6 +3145,11 @@ mod serve_group_test { version: Version::Lite07, timescale: Some(crate::Timescale::default()), opens: Default::default(), + // Content originating here: SUBSCRIBE_START names our own hop. + served: Served { + broadcast: None, + here: Hop::new(9).unwrap(), + }, }; let bounds = Bounds { start_group: Some(0), @@ -3098,8 +3183,8 @@ mod serve_group_test { track.finish().unwrap(); assert!(matches!(run.await.unwrap(), TrackEnd::Finished)); - // SUBSCRIBE_START at 0, then SUBSCRIBE_END at 3 with 2 streams. - assert_eq!(*log.writes.lock().unwrap(), [0, 1, 0, 1, 2, 3, 2]); + // SUBSCRIBE_START at 0 from origin 9, then SUBSCRIBE_END at 3 with 2 streams. + assert_eq!(*log.writes.lock().unwrap(), [0, 2, 0, 9, 1, 2, 3, 2]); } /// A track that ends without a group still ends the subscription, with no stream owed. @@ -3135,7 +3220,7 @@ mod serve_group_test { write_group(&mut track, 1, 1000); track.finish().unwrap(); assert!(futures::poll!(run.as_mut()).is_pending(), "group 1 is still opening"); - assert_eq!(*log.writes.lock().unwrap(), [0, 1, 0], "only SUBSCRIBE_START so far"); + assert_eq!(*log.writes.lock().unwrap(), [0, 2, 0, 9], "only SUBSCRIBE_START so far"); let Ok(mut open) = gate.write() else { panic!("transport gate closed"); @@ -3143,7 +3228,7 @@ mod serve_group_test { *open = true; drop(open); run.await.unwrap(); - assert_eq!(*log.writes.lock().unwrap(), [0, 1, 0, 1, 2, 2, 1]); + assert_eq!(*log.writes.lock().unwrap(), [0, 2, 0, 9, 1, 2, 2, 1]); } } diff --git a/rs/moq-net/src/lite/subscribe.rs b/rs/moq-net/src/lite/subscribe.rs index 4656e94521..a3513f8019 100644 --- a/rs/moq-net/src/lite/subscribe.rs +++ b/rs/moq-net/src/lite/subscribe.rs @@ -1,7 +1,7 @@ use std::borrow::Cow; use crate::{ - Path, + Hop, Path, coding::{Decode, DecodeError, Encode, EncodeError, Sizer}, }; @@ -295,6 +295,9 @@ impl Message for SubscribeOk { #[derive(Clone, Debug)] pub struct SubscribeStart { pub group: u64, + /// The origin serving the subscription: the Hop ID a relay stitches failover on. + /// [`Hop::UNKNOWN`] names nobody. Lite07+; older versions decode it as unknown. + pub origin: Hop, } impl Message for SubscribeStart { @@ -302,16 +305,23 @@ impl Message for SubscribeStart { if !version.has_track_stream() { return Err(DecodeError::Version); } - Ok(Self { - group: u64::decode(r, version)?, - }) + let group = u64::decode(r, version)?; + let origin = match version.has_origin() { + true => Hop::from_wire(u64::decode(r, version)?)?, + false => Hop::UNKNOWN, + }; + Ok(Self { group, origin }) } fn encode_msg(&self, w: &mut W, version: Version) -> Result<(), EncodeError> { if !version.has_track_stream() { return Err(EncodeError::Version); } - self.group.encode(w, version) + self.group.encode(w, version)?; + if version.has_origin() { + self.origin.id().encode(w, version)?; + } + Ok(()) } } @@ -578,7 +588,10 @@ mod test { #[test] fn subscribe_start_roundtrips_on_lite05() { - let resp = SubscribeResponse::Start(SubscribeStart { group: 42 }); + let resp = SubscribeResponse::Start(SubscribeStart { + group: 42, + origin: Hop::UNKNOWN, + }); let mut buf = Vec::new(); resp.encode(&mut buf, Version::Lite05).unwrap(); let mut slice = buf.as_slice(); @@ -602,6 +615,31 @@ mod test { } } + #[test] + fn subscribe_start_carries_the_origin_on_lite07() { + let resp = SubscribeResponse::Start(SubscribeStart { + group: 7, + origin: Hop::new(42).unwrap(), + }); + let mut buf = Vec::new(); + resp.encode(&mut buf, Version::Lite07).unwrap(); + // Type, length, group, origin. + assert_eq!(buf, [0, 2, 7, 42]); + match SubscribeResponse::decode(&mut buf.as_slice(), Version::Lite07).unwrap() { + SubscribeResponse::Start(start) => assert_eq!((start.group, start.origin), (7, Hop::new(42).unwrap())), + other => panic!("expected Start, got {other:?}"), + } + + // Older versions have no room for it. + let mut buf = Vec::new(); + resp.encode(&mut buf, Version::Lite06).unwrap(); + assert_eq!(buf, [0, 1, 7]); + match SubscribeResponse::decode(&mut buf.as_slice(), Version::Lite06).unwrap() { + SubscribeResponse::Start(start) => assert_eq!(start.origin, Hop::UNKNOWN), + other => panic!("expected Start, got {other:?}"), + } + } + #[test] fn subscribe_end_carries_the_stream_count_on_lite07() { let resp = SubscribeResponse::End(SubscribeEnd { group: 7, streams: 3 }); diff --git a/rs/moq-net/src/lite/subscriber.rs b/rs/moq-net/src/lite/subscriber.rs index 8a0259fec9..eb82982af0 100644 --- a/rs/moq-net/src/lite/subscriber.rs +++ b/rs/moq-net/src/lite/subscriber.rs @@ -94,6 +94,10 @@ struct TrackEntry { timescale: Option, /// The groups received so far, so the subscription's end can wait for the ones owed. tail: kio::Producer, + /// Whether this subscription's SUBSCRIBE_OK arrived. On a wire that names the + /// serving origin there, group streams wait for it: until then nobody knows + /// whose content they carry. + started: kio::Shared, } impl Subscriber { @@ -453,6 +457,12 @@ impl Subscriber { return Ok(()); }; + // A datagram cannot wait for its subscription's SUBSCRIBE_OK like a group stream + // does, so until the origin it names is admitted, the datagram is dropped. + if self.version.has_origin() && !entry.producer.provenance().is_admitted() { + return Ok(()); + } + // Datagrams are lite-05+, which always negotiates a timescale; default defensively. let scale = entry.timescale.unwrap_or_default(); let timestamp = @@ -738,6 +748,11 @@ struct GroupRecv { enum GroupRecvState { /// Reading the GROUP header. Header, + /// Holding the group until its subscription's SUBSCRIBE_OK names the serving origin + /// and the front admits it, on a wire that names one. + Hold { + header: lite::Group, + }, /// Filling the group, bailing if the track or group dies first. Serve { /// Guarded: dropping this machine mid-group is a cancellation, not a clean end. @@ -762,7 +777,31 @@ impl GroupRecv { match &mut self.state { GroupRecvState::Header => { let mut cx = std::task::Context::from_waker(waiter.waker()); - let hdr = ready!(self.reader.poll_decode::(&mut cx))?; + let header = ready!(self.reader.poll_decode::(&mut cx))?; + self.state = GroupRecvState::Hold { header }; + } + GroupRecvState::Hold { header } => { + if self.subscriber.version.has_origin() { + let entry = self + .subscriber + .subscribes + .lock() + .get(&header.subscribe) + .cloned() + .ok_or(Error::Cancel)?; + if let Poll::Ready(err) = entry.producer.poll_closed(waiter) { + return Poll::Ready(Err(err)); + } + ready!(entry.started.poll(waiter, |started| match **started { + true => Poll::Ready(()), + false => Poll::Pending, + })); + ready!(entry.producer.provenance().poll_admitted(waiter))?; + } + let GroupRecvState::Hold { header: hdr } = std::mem::replace(&mut self.state, GroupRecvState::Done) + else { + unreachable!() + }; let (group, track, timescale) = { let mut subs = self.subscriber.subscribes.lock(); @@ -1431,9 +1470,12 @@ mod tests { fn fin_responses(version: Version, started: bool, clean: bool) -> Vec { let mut responses = Vec::new(); if started { - lite::SubscribeResponse::Start(lite::SubscribeStart { group: 0 }) - .encode(&mut responses, version) - .unwrap(); + lite::SubscribeResponse::Start(lite::SubscribeStart { + group: 0, + origin: crate::Hop::UNKNOWN, + }) + .encode(&mut responses, version) + .unwrap(); } if clean { lite::SubscribeResponse::End(lite::SubscribeEnd { group: 0, streams: 0 }) @@ -1494,6 +1536,7 @@ mod tests { producer, timescale: Some(Timescale::default()), tail: Default::default(), + started: kio::Shared::new(true), }, ); @@ -1530,6 +1573,65 @@ mod tests { ); } + /// On lite-07 a datagram carries no origin of its own, so until its subscription's + /// origin is admitted it is dropped rather than risk splicing another origin's content. + #[test] + fn datagram_waits_for_the_admitted_origin() { + let version = Version::Lite07; + let origin = origin::Config::new(crate::Hop::new(1).unwrap()).produce(); + let subscriber = Subscriber::new(SubscriberConfig { + runtime: crate::time::Clock::tokio(), + session: SinkSession::default(), + origin, + recv_bandwidth: None, + version, + peer_setup: Default::default(), + peer_hop: None, + cost: None, + going_away: Default::default(), + }); + + let broadcast = crate::broadcast::Info::new().produce(); + let producer = broadcast.create_track("datagrams", None).unwrap(); + let provenance = producer.provenance(); + let mut received = producer.subscribe(None); + subscriber.subscribes.lock().insert( + 7, + TrackEntry { + producer, + timescale: Some(Timescale::default()), + tail: Default::default(), + started: kio::Shared::new(false), + }, + ); + + let payload = |sequence| { + lite::Datagram { + subscribe: 7, + sequence, + timestamp: sequence, + payload: bytes::Bytes::from_static(b"x"), + } + .encode_bytes(version) + .unwrap() + }; + + // Named but not admitted: the front serves another origin. + let named = crate::Hop::new(42).unwrap(); + provenance.name(named).unwrap(); + provenance.admit(crate::Hop::new(43).unwrap()); + subscriber.route_datagram(payload(1)).unwrap(); + assert!( + received.recv_datagram().now_or_never().is_none(), + "delivered before admission" + ); + + provenance.admit(named); + subscriber.route_datagram(payload(2)).unwrap(); + let datagram = received.recv_datagram().now_or_never().unwrap().unwrap().unwrap(); + assert_eq!(datagram.sequence, 2); + } + /// A lite-05 subscribe still waiting on TRACK_INFO is not in the subscribe map. /// Dropping that machine must reject the origin request with the session's error. #[tokio::test] @@ -2713,6 +2815,8 @@ struct SubStream { requested: Option, /// The groups received for this subscription, shared with its [`TrackEntry`]. tail: kio::Producer, + /// Whether SUBSCRIBE_OK arrived, shared with its [`TrackEntry`]. + started: kio::Shared, /// The first group the publisher serves (SUBSCRIBE_START), once declared. served: Option, /// The track's exclusive end and stream count (SUBSCRIBE_END), once declared. @@ -2733,7 +2837,7 @@ impl SubStream { enum Sub { None, - Active(SubStream), + Active(Box>), } /// Every advertisement the peer currently has live on one announce stream. @@ -3070,12 +3174,14 @@ impl TrackServe { tracing::info!(id, broadcast = %self.subscriber.log_path(&self.path), track = %self.name, "subscribe started"); let tail = kio::Producer::new(Tail::default()); + let started = kio::Shared::new(!self.subscriber.version.has_origin()); self.subscriber.subscribes.lock().insert( id, TrackEntry { producer: producer.clone(), timescale, tail: tail.clone(), + started: started.clone(), }, ); @@ -3087,6 +3193,7 @@ impl TrackServe { id, subscription, tail, + started, state: EstablishState::Open, } } @@ -3104,7 +3211,7 @@ impl TrackServe { let id = est.id; match kio::wait(move |waiter| est.poll(waiter)).await { Ok(active) => { - *sub = Sub::Active(active); + *sub = Sub::Active(Box::new(active)); Ok(()) } Err(err) => { @@ -3130,7 +3237,7 @@ impl TrackServe { let mut est = Box::new(est); let id = est.id; match kio::wait(move |waiter| est.poll(waiter)).await { - Ok(active) => *sub = Sub::Active(active), + Ok(active) => *sub = Sub::Active(Box::new(active)), Err(err) => { self.subscriber.remove_subscribe(id); return Err(err); @@ -3187,6 +3294,7 @@ struct Establish { id: u64, subscription: Subscription, tail: kio::Producer, + started: kio::Shared, state: EstablishState, } @@ -3272,6 +3380,7 @@ impl Establish { priority: self.subscription.priority, requested: self.subscription.start, tail: self.tail.clone(), + started: self.started.clone(), served: None, end: None, } @@ -3456,10 +3565,11 @@ impl TrackInfoFetch { // window matches what the upstream advertises (relays re-serve with // the same bound). `broadcast` is left at its default here; // `track::Request::accept` stamps the track's real broadcast. - let model = track::Info::default() + let mut model = track::Info::default() .with_timescale(info.timescale) .with_max_age(info.max_age) .with_priority(info.priority); + model.names_origin = serve.subscriber.version.has_origin(); return Poll::Ready(Ok(model)); } } @@ -3545,7 +3655,7 @@ impl ServeLoop { let id = est.id; self.mode = ServeMode::Select; match res { - Ok(active) => self.sub = Sub::Active(active), + Ok(active) => self.sub = Sub::Active(Box::new(active)), Err(err) => { // Opening the upstream failed (usually the session dying): hand // the track back for another route to resume. @@ -3602,8 +3712,12 @@ impl ServeLoop { match self.dynamic.poll_requested_group(waiter) { Poll::Ready(Ok(req)) => { if self.supports_fetch { - self.fetches - .push(FetchServeRun::new(serve.clone(), req, self.timescale)); + self.fetches.push(FetchServeRun::new( + serve.clone(), + req, + self.timescale, + self.serving.provenance(), + )); } else { req.reject(Error::Version); } @@ -3681,6 +3795,18 @@ impl ServeLoop { // signal, so a spliced reader waiting on a skipped group // fails over instead of stalling on a live route. lite::SubscribeResponse::Start(start) => { + // The reply names whose content this subscription + // carries. A copy has one origin, so another one is a + // different track: hand it back. The held group + // streams go ahead, and still wait on the front + // admitting that origin. + if serve.subscriber.version.has_origin() { + if let Err(err) = self.serving.provenance().name(start.origin) { + tracing::debug!(track = %serve.name, origin = start.origin.id(), "subscription changed origin"); + return Poll::Ready(ServeEnd::GiveBack(err)); + } + *active.started.lock() = true; + } // A START describes the demand the SUBSCRIBE carried. // It applies only while the current start still matches // that demand (updates get no fresh START, so an update @@ -3772,9 +3898,14 @@ struct FetchServeRun { session: S, timescale: Option, group: u64, + /// The copy's origin, which FETCH_OK names on lite-07. + provenance: track::Provenance, state: FetchRunState, } +// A state machine's enum is its storage: one transient instance per stream, so the +// big variant is the working state, not padding held in bulk. +#[allow(clippy::large_enum_variant)] enum FetchRunState { Open { request: Option, @@ -3790,6 +3921,12 @@ enum FetchRunState { stream: Stream, frame_start: u64, }, + /// FETCH_OK named the origin: hold the frames until the front admits it. + Admit { + request: Option, + stream: Stream, + frame_start: u64, + }, Ingest { stream: Stream, producer: group::Producer, @@ -3799,7 +3936,12 @@ enum FetchRunState { } impl FetchServeRun { - fn new(serve: TrackServe, request: group::Request, timescale: Option) -> Self { + fn new( + serve: TrackServe, + request: group::Request, + timescale: Option, + provenance: track::Provenance, + ) -> Self { let session = serve.subscriber.session.clone(); let group = request.sequence(); Self { @@ -3807,6 +3949,7 @@ impl FetchServeRun { session, timescale, group, + provenance, state: FetchRunState::Open { request: Some(request) }, } } @@ -3907,11 +4050,16 @@ impl kio::Task for FetchServeRun { }; } FetchRunState::Answer { stream, .. } => { - // Lite has no FETCH_OK: a publisher without the group resets the - // stream instead. Accepting before the first byte (or a FIN, for an - // empty group) would resolve every joined `fetch_group` to a group - // that only fails on its first read, so wait for the answer. - let answered = ready!(stream.reader.poll_has_more(&mut cx)); + // A publisher without the group resets the stream. Accepting before + // its answer (FETCH_OK on lite-07, the first byte or a FIN before) + // would resolve every joined `fetch_group` to a group that only fails + // on its first read, so wait for it. FETCH_OK names whose content + // follows, which a copy only takes from one origin. + let answered = match self.serve.subscriber.version.has_origin() { + true => ready!(stream.reader.poll_decode::(&mut cx)) + .and_then(|ok| self.provenance.name(ok.origin)), + false => ready!(stream.reader.poll_has_more(&mut cx)).map(|_| ()), + }; let FetchRunState::Answer { request, stream, @@ -3920,10 +4068,35 @@ impl kio::Task for FetchServeRun { else { unreachable!() }; - let request = request.expect("request pending"); if let Err(err) = answered { tracing::debug!(track = %self.serve.name, group = self.group, %err, "fetch refused"); stream.writer.abort(&err); + request.expect("request pending").reject(err); + return Poll::Ready(()); + } + self.state = FetchRunState::Admit { + request, + stream, + frame_start, + }; + } + FetchRunState::Admit { .. } => { + let admitted = match self.serve.subscriber.version.has_origin() { + true => ready!(self.provenance.poll_admitted(waiter)), + false => Ok(()), + }; + let FetchRunState::Admit { + request, + stream, + frame_start, + } = std::mem::replace(&mut self.state, FetchRunState::Done) + else { + unreachable!() + }; + let request = request.expect("request pending"); + if let Err(err) = admitted { + tracing::debug!(track = %self.serve.name, group = self.group, %err, "fetch from another origin"); + stream.writer.abort(&err); request.reject(err); return Poll::Ready(()); } diff --git a/rs/moq-net/src/lite/version.rs b/rs/moq-net/src/lite/version.rs index 33a1d52bd0..acf43e92e5 100644 --- a/rs/moq-net/src/lite/version.rs +++ b/rs/moq-net/src/lite/version.rs @@ -206,6 +206,18 @@ impl Version { } } + /// Whether SUBSCRIBE_OK and FETCH_OK name the origin serving the request, which is + /// what a relay stitches a failover on. Added in lite-07 (with FETCH_OK itself); + /// older versions leave a relay the route's first hop. + #[allow(clippy::match_like_matches_macro)] + pub(crate) fn has_origin(self) -> bool { + // Match form so future versions default forward (CLAUDE.md convention). + match self { + Self::Lite01 | Self::Lite02 | Self::Lite03 | Self::Lite04 | Self::Lite05 | Self::Lite06 => false, + _ => true, + } + } + /// Whether announcements carry the route cost: the marginal cost of pulling /// the broadcast via this route, accumulated per link. Added in lite-06. /// Older versions carry nothing, so a received route stays at zero and ranks diff --git a/rs/moq-net/src/model/broadcast.rs b/rs/moq-net/src/model/broadcast.rs index 3b7f90306e..659984a922 100644 --- a/rs/moq-net/src/model/broadcast.rs +++ b/rs/moq-net/src/model/broadcast.rs @@ -116,6 +116,10 @@ struct SplicedState { // Names awaiting assignment to a route, in request order. pending: VecDeque>, + + // The origin the front serves, as its replies name it. Set by the front as its + // identity settles; `None` until then. + origin: Option, } impl BroadcastState { @@ -404,6 +408,17 @@ impl Producer { Poll::Ready((name, producer)) } + /// Record the origin the front serves; see [`Consumer::origin`]. Route-fed + /// broadcasts only. + pub(crate) fn set_origin(&self, origin: crate::Hop) { + if self.state.read().spliced.as_ref().and_then(|spliced| spliced.origin) == Some(origin) { + return; + } + let mut state = self.state.lock(); + let spliced = state.spliced.as_mut().expect("origin of a route-fed broadcast"); + spliced.origin = Some(origin); + } + /// Let go of every spliced track, aborting with `err` the ones never handed /// out by [`Self::poll_spliced_assigned`]. Called when the broadcast ends: /// whoever took the others decides how they end. @@ -911,6 +926,12 @@ impl Consumer { self.state.read().closing } + /// The origin serving this broadcast, as a reply for it names it: `None` for a + /// broadcast that is not route-fed, whose content originates on this origin. + pub(crate) fn origin(&self) -> Option { + self.state.read().spliced.as_ref().and_then(|spliced| spliced.origin) + } + #[doc(hidden)] #[deprecated(note = "a broadcast end carries no cause")] pub fn is_finished(&self) -> bool { diff --git a/rs/moq-net/src/model/front.rs b/rs/moq-net/src/model/front.rs index b56136b02a..f069c74f8d 100644 --- a/rs/moq-net/src/model/front.rs +++ b/rs/moq-net/src/model/front.rs @@ -21,7 +21,8 @@ use std::{ use crate::{Error, Hop, runtime::Instant, track}; /// A route the table selected for the front: the entry id, the endpoint that -/// originated it, and whether it is a broadcast published on this origin. +/// originated it (`None` for an empty chain: announced on this origin), and +/// whether it is a broadcast published on this origin. #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub(super) struct Candidate { pub route: u64, @@ -74,6 +75,9 @@ pub(super) enum Event { result: Result<(), Error>, delivered: bool, }, + /// A spliced copy's reply named the origin serving it. Its content is held + /// until the front [admits](Front::admit) that origin. + Origin { track: Arc, source: u64, origin: Hop }, /// A reader arrived on the track. Used { track: Arc }, /// The last reader left the track. @@ -131,13 +135,20 @@ pub(super) enum Identity { /// closing the front ends instead, so a newcomer gets a fresh broadcast /// rather than being spliced into one that is over. Local, - /// The serving route's first hop was absent or [`Hop::UNKNOWN`], which - /// identifies nobody: the front cannot resume through any other route, so - /// its source ending, or `route` leaving the table, ends it. - Anonymous { route: u64 }, - /// The first hop of the serving route. Routes sharing it are the same - /// origin reached another way and safe to resume through. + /// The serving route names no other origin: its chain is empty (a handler on + /// this origin, `here`) or its first hop is [`Hop::UNKNOWN`], which + /// identifies nobody. The front cannot resume through any other route, so its + /// source ending, or `route` leaving the table, ends it. + Anonymous { route: u64, here: bool }, + /// The first hop of the serving route, for sources whose replies name no + /// origin (older wires). Routes sharing it are the same origin reached + /// another way and safe to resume through. Publisher(Hop), + /// The origin the first source's replies named. Any route may lead back to it, + /// since an advertiser serving a prefix from a pool advertises one route for many + /// origins, so a replacement source is verified by its own replies instead of by + /// its route. + Origin(Hop), } /// Which routes qualify for a front's (re)selection. @@ -151,15 +162,19 @@ pub(super) enum Pin { Publisher(Hop), /// Only this route: the front never fails over. Route(u64), + /// Any served route, but this one while it stands: a replacement's content + /// is only known once it replies, so a live source is not traded for it. + Stay(u64), } impl Identity { - fn pin(self) -> Pin { + /// The origin the front's content comes from, as its own replies name it: + /// `None` for this origin, [`Hop::UNKNOWN`] for nobody identifiable. + fn origin(self) -> Option { match self { - Self::Undetermined => Pin::Any, - Self::Local => Pin::Local, - Self::Anonymous { route } => Pin::Route(route), - Self::Publisher(hop) => Pin::Publisher(hop), + Self::Local | Self::Anonymous { here: true, .. } => None, + Self::Undetermined | Self::Anonymous { here: false, .. } => Some(Hop::UNKNOWN), + Self::Publisher(hop) | Self::Origin(hop) => Some(hop), } } } @@ -195,6 +210,11 @@ struct Track { #[derive(Clone, Debug)] pub(super) struct Front { identity: Identity, + /// The origin the replies established, and the only one whose content is + /// admitted. `None` until a copy whose wire names origins replies. + named: Option, + /// The first source attached: its replies may name what the route only labelled. + first: Option, /// The attached source and the route that produced it. serving: Option<(u64, u64)>, /// Whether the serving source has begun closing, as of the last event that @@ -202,9 +222,8 @@ pub(super) struct Front { serving_closing: bool, /// The route an upstream request is in flight through. upstream: Option, - /// Routes excluded from selection: they refused the path while another - /// source was serving, or their source ended while still advertised. - refused: HashSet, + /// Routes excluded from selection: their source ended while still advertised. + excluded: HashSet, /// Why the last candidate fell through, reported if the front ends unresolved. last_err: Option, /// Whether the parked requesters were resolved (the first source attached). @@ -226,10 +245,12 @@ impl Front { pub(super) fn new(linger: Duration) -> Self { Self { identity: Identity::Undetermined, + named: None, + first: None, serving: None, serving_closing: false, upstream: None, - refused: HashSet::new(), + excluded: HashSet::new(), last_err: None, resolved: false, tracks: BTreeMap::new(), @@ -242,17 +263,37 @@ impl Front { /// Which routes qualify for the next selection. pub(super) fn pin(&self) -> Pin { - self.identity.pin() + match self.identity { + Identity::Undetermined => Pin::Any, + Identity::Local => Pin::Local, + Identity::Anonymous { route, .. } => Pin::Route(route), + Identity::Publisher(hop) => Pin::Publisher(hop), + Identity::Origin(_) => match self.serving { + Some((_, route)) => Pin::Stay(route), + None => Pin::Any, + }, + } + } + + /// The origin this front's replies name: `None` for this origin, see + /// [`Identity::origin`]. + pub(super) fn origin(&self) -> Option { + self.identity.origin() } /// The routes excluded from selection; the driver skips them. - pub(super) fn refused_routes(&self) -> &HashSet { - &self.refused + pub(super) fn excluded_routes(&self) -> &HashSet { + &self.excluded } - /// Forget refused routes that left the table (a reconnect is a fresh entry). + /// Forget excluded routes that left the table (a reconnect is a fresh entry). pub(super) fn retain_routes(&mut self, standing: impl Fn(u64) -> bool) { - self.refused.retain(|route| standing(*route)); + self.excluded.retain(|route| standing(*route)); + } + + /// The origin whose copies may deliver content, once replies established one. + pub(super) fn admit(&self) -> Option { + self.named } /// The attached source, if any. @@ -303,6 +344,7 @@ impl Front { result, delivered, } => self.track_ended(track, source, closing, result, delivered, &mut actions), + Event::Origin { track, source, origin } => self.origin_named(track, source, origin, &mut actions), Event::Used { track } => self.used(track, &mut actions), Event::Unused { track, now } => self.unused(track, now, &mut actions), Event::Deadline { now } => self.deadline(now, &mut actions), @@ -370,18 +412,11 @@ impl Front { self.last_err = Some(Error::Unroutable); actions.push(Action::Reselect); } - // An authoritative refusal of the path. It ends a front with no - // other source; a serving front merely skips the refuser. + // An authoritative refusal of the path ends the front, serving or + // not: a refusal never moves to another route. Err(Refusal { err, standing: true }) => { self.upstream = None; - match self.serving { - Some(_) => { - self.refused.insert(route); - self.last_err = Some(err); - actions.push(Action::Reselect); - } - None => self.end(err, actions), - } + self.end(err, actions); } } } @@ -399,6 +434,7 @@ impl Front { } } self.serving = Some((source, route)); + self.first.get_or_insert(source); self.serving_closing = false; if !self.resolved { self.resolved = true; @@ -425,7 +461,10 @@ impl Front { self.identity = match (candidate.local, candidate.first) { (true, _) => Identity::Local, (false, Some(hop)) if hop != Hop::UNKNOWN => Identity::Publisher(hop), - (false, _) => Identity::Anonymous { route: candidate.route }, + (false, first) => Identity::Anonymous { + route: candidate.route, + here: first.is_none(), + }, }; } @@ -439,7 +478,7 @@ impl Front { // A standing route can outlive the source it produced. Asking it again // would re-request the broadcast that just ended; another route to the // same publisher may still resume it. - self.refused.insert(route); + self.excluded.insert(route); self.serving = None; self.serving_closing = false; actions.push(Action::Detach { source }); @@ -454,7 +493,7 @@ impl Front { // A local publisher ending ends its broadcast; a newcomer at the // path gets a fresh one. An anonymous source can never be resumed. Identity::Local | Identity::Anonymous { .. } | Identity::Undetermined => self.end(Error::Dropped, actions), - Identity::Publisher(_) => actions.push(Action::Reselect), + Identity::Publisher(_) | Identity::Origin(_) => actions.push(Action::Reselect), } } @@ -466,10 +505,17 @@ impl Front { result: Result, actions: &mut Vec, ) { - let Some(track) = self.tracks.get_mut(&name) else { + if self.tracks.get(&name).map(|track| &track.state) != Some(&TrackState::Querying { source }) { return; - }; - if track.state != (TrackState::Querying { source }) { + } + // Once replies named the front's origin, a copy whose replies cannot is + // content nobody can vouch for: let it go, so the end cannot splice it + // either, and end the front; a re-request gets a fresh one. + if let Ok(info) = &result + && self.named.is_some() + && !info.names_origin + { + self.reject(source, actions); return; } let verdict = match result { @@ -493,6 +539,7 @@ impl Front { }; match verdict { Ok(()) => { + let track = self.tracks.get_mut(&name).expect("querying a known track"); track.state = TrackState::Spliced { source }; actions.push(Action::Splice { track: name.clone(), @@ -503,6 +550,71 @@ impl Front { } } + /// A spliced copy's reply named its origin: admit it, or end the front if it is + /// another origin's content. The copy held everything so far, so nothing of it + /// was delivered. + fn origin_named(&mut self, name: Arc, source: u64, origin: Hop, actions: &mut Vec) { + if self.serving.map(|(serving, _)| serving) != Some(source) || !self.tracks.contains_key(&name) { + return; + } + if !self.vouch(source, origin) { + self.reject(source, actions); + } + } + + /// End the front over a source carrying another origin's content. Its copies + /// delivered nothing, so a track spliced from one is aborted rather than left + /// to end with it, and the source is let go. + fn reject(&mut self, source: u64, actions: &mut Vec) { + let rejected: Vec> = self + .tracks + .iter() + .filter(|(_, track)| { + matches!(track.state, TrackState::Querying { source: s } | TrackState::Spliced { source: s } if s == source) + }) + .map(|(name, _)| name.clone()) + .collect(); + for name in rejected { + self.tracks.remove(&name); + actions.push(Action::Abort { + track: name, + err: Error::Dropped, + }); + } + actions.push(Action::Detach { source }); + self.end(Error::Dropped, actions); + } + + /// Whether `origin`, named by a reply from `source`, is the front's origin, + /// establishing it from the first source's first reply. + fn vouch(&mut self, source: u64, origin: Hop) -> bool { + if let Some(named) = self.named { + return named == origin; + } + match self.identity { + // The first source's reply names what its route could only label, since + // an advertiser serving a pool advertises one route for all of it. + Identity::Publisher(_) | Identity::Anonymous { .. } if self.first == Some(source) => { + self.identity = match origin { + Hop::UNKNOWN => Identity::Anonymous { + route: self + .serving + .map(|(_, route)| route) + .expect("a reply comes from a serving source"), + here: false, + }, + origin => Identity::Origin(origin), + }; + } + // Content already flowed under the route's label, from a wire that names no + // origin: only that origin matches. + Identity::Publisher(hop) if hop == origin => {} + _ => return false, + } + self.named = Some(origin); + true + } + fn track_ended( &mut self, name: Arc, @@ -677,6 +789,23 @@ mod tests { track::Info::default() } + /// Track metadata from a wire whose replies name the origin. + fn vouching() -> track::Info { + track::Info { + names_origin: true, + ..info() + } + } + + /// `source`'s copy of `video` named `origin`. + fn named(source: u64, origin: u64) -> Event { + Event::Origin { + track: name("video"), + source, + origin: Hop::from_wire(origin).unwrap(), + } + } + fn name(s: &str) -> Arc { Arc::from(s) } @@ -689,6 +818,11 @@ mod tests { /// A front serving `source` through `candidate`, with `video` spliced and read. fn serving(candidate: Candidate, source: u64) -> Front { + serving_with(candidate, source, None) + } + + /// [`serving`], with the copy of `video` replying that `origin` serves it. + fn serving_with(candidate: Candidate, source: u64, origin: Option) -> Front { let mut front = Front::new(LINGER); assert_actions( front.step(Event::Selected { @@ -718,13 +852,19 @@ mod tests { track: name("video"), source, closing: false, - result: Ok(info()), + result: Ok(match origin { + Some(_) => vouching(), + None => info(), + }), }), &[Action::Splice { track: name("video"), source, }], ); + if let Some(origin) = origin { + assert_actions(front.step(named(source, origin)), &[]); + } front } @@ -735,6 +875,177 @@ mod tests { assert_eq!(front.pin(), Pin::Publisher(hop(10))); } + /// An advertiser serving a pool advertises one route for all of it, so the + /// reply, not the route, names what the front serves. + #[test] + fn the_first_reply_names_the_origin() { + let front = serving_with(remote(1, 10), 100, Some(20)); + assert_eq!(front.identity, Identity::Origin(hop(20))); + assert_eq!(front.origin(), Some(hop(20))); + assert_eq!(front.admit(), Some(hop(20))); + // Any route may lead back to that origin, but the serving one stays. + assert_eq!(front.pin(), Pin::Stay(1)); + + // A reply naming nobody pins the front to its route. + let front = serving_with(remote(1, 10), 100, Some(0)); + assert_eq!(front.identity, Identity::Anonymous { route: 1, here: false }); + assert_eq!(front.origin(), Some(Hop::UNKNOWN)); + + // Nothing is admitted until a reply names something. + assert_eq!(serving(remote(1, 10), 100).admit(), None); + } + + /// A replacement source up to its TRACK_INFO: re-requested, resolved, and its + /// copy of `video` spliced (holding its content until its origin is admitted). + fn fail_over(front: &mut Front, candidate: Candidate, source: u64, info: track::Info) -> Vec { + front.step(Event::SourceClosed { + source: front.serving().unwrap(), + }); + front.step(Event::Selected { + best: Some(candidate), + serving_closing: false, + }); + front.step(Event::Resolved { + route: candidate.route, + result: Ok(source), + }); + front.step(Event::TrackInfo { + track: name("video"), + source, + closing: false, + result: Ok(info), + }) + } + + /// The serving source dies and the best remaining route has another first + /// hop, but its reply names the same origin: the front resumes there. + #[test] + fn a_replacement_naming_the_same_origin_resumes() { + let mut front = serving_with(remote(1, 10), 100, Some(20)); + assert_actions( + front.step(Event::SourceClosed { source: 100 }), + &[Action::Detach { source: 100 }, Action::Reselect], + ); + assert_eq!(front.pin(), Pin::Any); + assert_actions( + front.step(Event::Selected { + best: Some(remote(2, 11)), + serving_closing: false, + }), + &[Action::Request { route: 2 }], + ); + assert_actions( + front.step(Event::Resolved { + route: 2, + result: Ok(200), + }), + &[Action::Query { + track: name("video"), + source: 200, + }], + ); + assert_actions( + front.step(Event::TrackInfo { + track: name("video"), + source: 200, + closing: false, + result: Ok(vouching()), + }), + &[Action::Splice { + track: name("video"), + source: 200, + }], + ); + assert_actions(front.step(named(200, 20)), &[]); + assert_eq!(front.pin(), Pin::Stay(2)); + } + + /// A replacement whose reply names another origin is different content: its + /// copy, which held everything, is let go and the front ends. + #[test] + fn a_replacement_naming_another_origin_ends_the_front() { + for reply in [21, 0] { + let mut front = serving_with(remote(1, 10), 100, Some(20)); + fail_over(&mut front, remote(2, 10), 200, vouching()); + assert_actions( + front.step(named(200, reply)), + &[ + Action::Abort { + track: name("video"), + err: Error::Dropped, + }, + Action::Detach { source: 200 }, + Action::End { err: Error::Dropped }, + ], + ); + } + } + + /// A replacement whose wire names no origin cannot vouch for the front's: it + /// is let go before it is spliced. + #[test] + fn a_replacement_that_cannot_vouch_ends_the_front() { + let mut front = serving_with(remote(1, 10), 100, Some(20)); + assert_actions( + fail_over(&mut front, remote(2, 10), 200, info()), + &[ + Action::Abort { + track: name("video"), + err: Error::Dropped, + }, + Action::Detach { source: 200 }, + Action::End { err: Error::Dropped }, + ], + ); + } + + /// Once content flowed under a route's label from a wire naming no origin, a + /// later reply can only confirm the label, never replace it. + #[test] + fn a_reply_after_unnamed_content_must_match_the_route_label() { + let mut front = serving(remote(1, 10), 100); + fail_over(&mut front, remote(2, 10), 200, vouching()); + assert_actions( + front.step(named(200, 20)), + &[ + Action::Abort { + track: name("video"), + err: Error::Dropped, + }, + Action::Detach { source: 200 }, + Action::End { err: Error::Dropped }, + ], + ); + + let mut front = serving(remote(1, 10), 100); + fail_over(&mut front, remote(2, 10), 200, vouching()); + assert_actions(front.step(named(200, 10)), &[]); + assert_eq!(front.admit(), Some(hop(10))); + } + + /// A reply from a source the front already let go changes nothing. + #[test] + fn a_stale_reply_is_ignored() { + let mut front = serving_with(remote(1, 10), 100, Some(20)); + fail_over(&mut front, remote(2, 10), 200, vouching()); + assert_actions(front.step(named(100, 21)), &[]); + } + + /// A handler on this origin has an empty chain: its content is named by this + /// origin, though the front still never leaves its route. + #[test] + fn an_empty_chain_originates_here() { + let candidate = Candidate { + route: 1, + first: None, + local: false, + }; + let front = serving(candidate, 100); + assert_eq!(front.identity, Identity::Anonymous { route: 1, here: true }); + assert_eq!(front.origin(), None); + assert_eq!(front.pin(), Pin::Route(1)); + } + #[test] fn nothing_routable_ends_an_unresolved_front() { let mut front = Front::new(LINGER); @@ -818,7 +1129,7 @@ mod tests { front.step(Event::SourceClosed { source: 100 }), &[Action::Detach { source: 100 }, Action::Reselect], ); - assert!(front.refused_routes().contains(&1)); + assert!(front.excluded_routes().contains(&1)); assert_actions( front.step(Event::Selected { best: Some(remote(3, 10)), @@ -926,7 +1237,7 @@ mod tests { } #[test] - fn standing_refusal_while_serving_skips_the_refuser() { + fn standing_refusal_while_serving_ends_the_front() { let mut front = serving(remote(1, 10), 100); front.step(Event::Selected { best: Some(remote(2, 10)), @@ -940,10 +1251,9 @@ mod tests { standing: true, }), }), - &[Action::Reselect], + &[Action::End { err: Error::NotFound }], ); - assert!(front.refused_routes().contains(&2)); - assert_eq!(front.serving, Some((100, 1))); + assert!(front.ended()); } #[test] @@ -963,7 +1273,6 @@ mod tests { }), &[Action::Reselect], ); - assert!(front.refused_routes().is_empty()); assert!(!front.ended()); } @@ -1255,6 +1564,22 @@ mod tests { closing: false, result: Err(Error::NotFound), }, + Event::TrackInfo { + track: name("v"), + source: 200, + closing: false, + result: Ok(vouching()), + }, + Event::Origin { + track: name("v"), + source: 100, + origin: hop(20), + }, + Event::Origin { + track: name("v"), + source: 200, + origin: hop(21), + }, Event::TrackEnded { track: name("v"), source: 100, @@ -1303,6 +1628,16 @@ mod tests { .filter(|action| matches!(action, Action::Splice { .. })) .count(); assert!(splices <= 1); + // A copy's content is only admitted when the front serves the origin + // its reply named. + if let Event::Origin { source, origin, .. } = event + && !next.ended() + && front.serving() == Some(*source) + && front.tracks.contains_key("v") + { + assert_eq!(next.admit(), Some(*origin)); + assert_eq!(next.origin(), Some(*origin)); + } // A Detach names a source the front no longer serves from. for action in &actions { if let Action::Detach { source } = action { diff --git a/rs/moq-net/src/model/origin.rs b/rs/moq-net/src/model/origin.rs index a93764e92e..36f03255a1 100644 --- a/rs/moq-net/src/model/origin.rs +++ b/rs/moq-net/src/model/origin.rs @@ -583,20 +583,25 @@ fn fnv_key(name: &str, origins: impl IntoIterator) -> u64 { hash } -/// Ordering key for a route entry covering one prefix. Lower wins: an identified +/// Ordering key for a route entry resolving `path`. Lower wins: an identified /// chain (no 0) outranks an anonymous one regardless of cost, then the cheapest /// cost, then a broadcast published on this origin (it serves what is here, not a /// claim that has to ask), then the shortest hop chain, then a deterministic hash -/// of the prefix and chain so every node converges on the same winner, and finally +/// of `path` and the chain so every node converges on the same winner, and finally /// the newest announcement, so a reconnect under an otherwise identical route wins /// the moment it lands instead of after the transport retires the old session. -fn route_order(prefix: &Path, entry: &RouteEntry) -> (bool, Cost, bool, usize, u64, Reverse) { +/// +/// `path` is what is being resolved: the requested path for a request, the prefix +/// itself for an advertisement. Keying the hash on the requested path is what +/// spreads equal-cost advertisers of one prefix: each path picks its own winner +/// from the pool, rather than every path under the prefix hashing alike. +fn route_order(path: &Path, entry: &RouteEntry) -> (bool, Cost, bool, usize, u64, Reverse) { ( entry.is_anonymous(), entry.cost, !entry.local, entry.hops.len(), - fnv_key(prefix.as_str(), entry.hops.iter().copied()), + fnv_key(path.as_str(), entry.hops.iter().copied()), Reverse(entry.id), ) } @@ -768,7 +773,7 @@ impl RouteEntry { /// Whether `pin` admits this entry for a front's selection. fn qualifies(&self, pin: Pin) -> bool { match pin { - Pin::Any => true, + Pin::Any | Pin::Stay(_) => true, Pin::Local => self.local, Pin::Publisher(first) => self.hops.iter().next() == Some(&first), Pin::Route(id) => self.id == id, @@ -2175,6 +2180,8 @@ struct FrontTask { request: kio::Producer, /// Published for requesters once the first source fixes it; see [`RemoteFront::pin`]. pin: kio::Lock, + /// This origin's hop, which names content originating here. + hop: Hop, timers: Clock, } @@ -2201,6 +2208,18 @@ struct TrackIo { head: Option, /// Whether the track had a reader as of the last demand edge. used: bool, + /// Whether the spliced copy's named origin was handed to the machine. + reported: bool, +} + +impl TrackIo { + /// Every copy the driver holds for the track, with its source. + fn copies(&self) -> impl Iterator { + let query = self.query.as_ref().map(|(source, copy, _)| (*source, copy)); + let staged = self.staged.as_ref().map(|(source, copy)| (*source, copy)); + let copy = self.copy.as_ref().map(|(source, copy)| (*source, copy)); + query.into_iter().chain(staged).chain(copy) + } } /// Drives one front: feeds the world's events to a [`Front`] and performs the @@ -2215,6 +2234,7 @@ async fn run_front(task: FrontTask) { watch, request, pin, + hop, timers, } = task; @@ -2224,6 +2244,7 @@ async fn run_front(task: FrontTask) { Resolved(u64, Result), SourceClosed(u64), Info(Arc, u64, Result), + Origin(Arc, u64, Hop), Ended(Arc, u64, Result<(), Error>), Demand(Arc), Deadline, @@ -2231,6 +2252,16 @@ async fn run_front(task: FrontTask) { } let mut front = Front::new(TRACK_IDLE_LINGER); + // The origin the front's replies name. Content nobody identifies (an anonymous + // source, or an origin without a hop) gets a random one for this front's life, + // rather than 0: a downstream relay can still resume within it, and nothing + // else ever names it. + let anonymous = Hop::random(); + let named = |front: &Front| match front.origin() { + None if hop != Hop::UNKNOWN => hop, + Some(origin) if origin != Hop::UNKNOWN => origin, + _ => anonymous, + }; let mut sources: HashMap = HashMap::new(); let mut next_source = 0u64; // The in-flight upstream request: the route and its pending channel. @@ -2252,7 +2283,7 @@ async fn run_front(task: FrontTask) { *seen = watch.seen(); front.retain_routes(|route| table.routes.covers(&path.as_path(), route)); let best = table - .best_route(&path.as_path(), horizon, front.pin(), front.refused_routes()) + .best_route(&path.as_path(), horizon, front.pin(), front.excluded_routes()) .map(|entry| Candidate { route: entry.id, first: entry.hops.iter().next().copied(), @@ -2269,7 +2300,21 @@ async fn run_front(task: FrontTask) { loop { while let Some(event) = events.pop_front() { - for action in front.step(event) { + let actions = front.step(event); + // Published before any copy is admitted below: once a copy's content + // flows, a reply for it names this origin. + *pin.lock() = front.pin(); + broadcast.set_origin(named(&front)); + if let Some(origin) = front.admit() { + for io in tracks.values() { + for (_, copy) in io.copies() { + if let Some(provenance) = copy.provenance() { + provenance.admit(origin); + } + } + } + } + for action in actions { match action { Action::Reselect => events.push_back(select(&mut front, &sources, &mut seen)), Action::Request { route } => { @@ -2304,6 +2349,7 @@ async fn run_front(task: FrontTask) { }; front.identify(candidate); *pin.lock() = front.pin(); + broadcast.set_origin(named(&front)); if let Some(source) = source { let id = next_source; next_source += 1; @@ -2373,8 +2419,14 @@ async fn run_front(task: FrontTask) { Action::Detach { source } => { sources.remove(&source); // Its copies go with it; the segments they delivered stay - // spliced until a replacement resumes past them. + // spliced until a replacement resumes past them. Whatever they + // still hold for admission never will be. for io in tracks.values_mut() { + for (_, copy) in io.copies().filter(|(s, _)| *s == source) { + if let Some(provenance) = copy.provenance() { + provenance.refuse(); + } + } if io.copy.as_ref().is_some_and(|(s, _)| *s == source) { io.copy = None; } @@ -2431,6 +2483,7 @@ async fn run_front(task: FrontTask) { // edge the copy is asked to advance. io.edge = io.resume.resume_position(); io.copy = Some((source, copy)); + io.reported = false; } Action::Park { track: name } => { let Some(io) = tracks.get_mut(&name) else { continue }; @@ -2469,6 +2522,12 @@ async fn run_front(task: FrontTask) { Action::Abort { track: name, err } => { if let Some(mut io) = tracks.remove(&name) { tracing::debug!(name = %name, %err, "aborting track"); + // Nothing a copy still holds for admission will be read. + for (_, copy) in io.copies() { + if let Some(provenance) = copy.provenance() { + provenance.refuse(); + } + } let _ = io.resume.abort(err); } } @@ -2491,6 +2550,16 @@ async fn run_front(task: FrontTask) { // ends as that copy does. let waiting = io.staged.take().map(|(_, copy)| copy); let waiting = waiting.or_else(|| io.query.take().map(|(_, copy, _)| copy)); + // Whatever a copy still holds is only delivered if it is the + // front's origin; with none established, nobody vouches for it. + for copy in waiting.iter().chain(io.copy.as_ref().map(|(_, copy)| copy)) { + if let Some(provenance) = copy.provenance() { + match front.admit() { + Some(origin) => provenance.admit(origin), + None => provenance.refuse(), + } + } + } if let Some(copy) = waiting && io.resume.is_used() { @@ -2546,6 +2615,12 @@ async fn run_front(task: FrontTask) { { return Poll::Ready(Step::Ended(name.clone(), *source, result)); } + if let Some((source, copy)) = &io.copy + && !io.reported && let Some(provenance) = copy.provenance() + && let Poll::Ready(origin) = provenance.poll_named(waiter) + { + return Poll::Ready(Step::Origin(name.clone(), *source, origin)); + } // Watch the demand edge in whichever direction is unmet. let edge = match io.used { true => io.resume.poll_unused(waiter), @@ -2575,6 +2650,7 @@ async fn run_front(task: FrontTask) { warm: None, head: None, used: false, + reported: false, }, ); Event::TrackAssigned { track: name } @@ -2602,6 +2678,15 @@ async fn run_front(task: FrontTask) { } } Step::SourceClosed(source) => Event::SourceClosed { source }, + Step::Origin(name, source, origin) => { + let Some(io) = tracks.get_mut(&name) else { continue }; + io.reported = true; + Event::Origin { + track: name, + source, + origin, + } + } Step::Info(name, source, result) => { let closing = sources.get(&source).is_some_and(|s| s.is_closing()); let Some(io) = tracks.get_mut(&name) else { continue }; @@ -3131,7 +3216,7 @@ impl OriginState { } /// The best served route covering `path` (absolute) for a requester seeing - /// `horizon`, skipping the `refused` entry ids. + /// `horizon`, skipping the `excluded` entry ids. /// /// The most specific covering prefix wins outright, so a narrow advertise-only /// announcement shadows a broad served one: requests under it resolve @@ -3139,12 +3224,26 @@ impl OriginState { /// prefix, the cheapest served one is picked by [`route_order`]. /// /// Only announced routes are candidates: an unannounced broadcast serves - /// nobody, and does not shadow anything either. `pin` is the front's + /// nobody, and does not shadow anything either. The hash tie-break is keyed + /// on `path`, so equal-cost advertisers of one prefix share its paths. `pin` is the front's /// identity: only routes it admits are candidates, since a route from anyone /// else is different content rather than an alternate path (see [`Front`]). /// A broadcast published on this origin competes on cost like any other /// route and wins a tie. - fn best_route(&self, path: &Path, horizon: Horizon, pin: Pin, refused: &HashSet) -> Option<&RouteEntry> { + fn best_route(&self, path: &Path, horizon: Horizon, pin: Pin, excluded: &HashSet) -> Option<&RouteEntry> { + // A route a front stays on wins while it can still serve, whatever else + // appeared since: see [`Pin::Stay`]. + if let Pin::Stay(route) = pin + && let Some(entry) = self.routes.covering(path).find(|entry| { + entry.id == route + && entry.advertised + && entry.scope.matches(path.as_str()) + && horizon.admits(entry) + && entry.serves(path) + }) { + return Some(entry); + } + // Covering prefixes of one path form a chain, so the deepest node with a // candidate holds the unique longest prefix; walking down, the last such // node decides. @@ -3158,12 +3257,12 @@ impl OriginState { .filter(|entry| entry.scope.matches(path.as_str())) .filter(|entry| horizon.admits(entry)) .filter(|entry| entry.qualifies(pin)) - .filter(|entry| !refused.contains(&entry.id)) + .filter(|entry| !excluded.contains(&entry.id)) .peekable(); if candidates.peek().is_some() { best = candidates .filter(|entry| entry.serves(path)) - .min_by_key(|entry| route_order(&entry.prefix, entry)); + .min_by_key(|entry| route_order(path, entry)); } } best @@ -3985,6 +4084,7 @@ impl Consumer { watch, request, pin, + hop: self.hop, timers: self.timers.clone(), })); kio::Pending::new(Requesting::queued(consumer).with_path(requested).with_stats(scope)) @@ -4847,6 +4947,56 @@ mod tests { announced.assert_next_ended("room"); } + /// Equal-cost advertisers of one prefix share its paths: a set of requested + /// paths spreads across the pool, and one path always resolves to the same + /// advertiser, whatever order the routes arrived in. + /// Pinned so `spreadHash` in `js/net` picks the same pool member for a path. + #[test] + fn spread_hash_matches_js() { + assert_eq!(fnv_key("pool/job-0", [origin(10)]), 0xefb5e20a66101c32); + assert_eq!(fnv_key("pool/job-0", [origin(11)]), 0x0eb0a91370ff6653); + } + + #[tokio::test] + async fn equal_cost_pool_spreads_paths() { + const WORKERS: [u64; 4] = [10, 11, 12, 13]; + const PATHS: usize = 64; + + // The first hop of the route each path resolves to on a node whose pool + // arrived in `order`. + fn winners(order: impl Iterator) -> Vec { + let producer = origin(1).produce(); + let _pool: Vec = order + .map(|id| { + producer + .dynamic("pool", Route::default().with_hops(hops(&[id])).with_cost(3)) + .unwrap() + }) + .collect(); + let table = producer.shared.read(); + (0..PATHS) + .map(|i| { + let path = Path::new(&format!("pool/job-{i}")).to_owned(); + let entry = table + .best_route(&path.as_path(), Horizon::default(), Pin::Any, &HashSet::new()) + .expect("the pool serves every path"); + entry.hops.iter().next().copied().unwrap() + }) + .collect() + } + + let forward = winners(WORKERS.into_iter()); + let reverse = winners(WORKERS.into_iter().rev()); + assert_eq!(forward, reverse, "a path must resolve the same way on every node"); + + // Not an assertion about any two paths, which a correct hash may put on + // one worker: only that the set does not pile onto a few. + for worker in WORKERS { + let share = forward.iter().filter(|hop| **hop == origin(worker)).count(); + assert!(share >= PATHS / 16, "worker {worker} took {share} of {PATHS} paths"); + } + } + #[tokio::test] async fn identical_reannounce_is_invisible() { let producer = origin(1).produce(); @@ -6140,6 +6290,61 @@ mod tests { } } + #[tokio::test] + async fn refusal_while_serving_does_not_fall_through_to_a_broader_route() { + let producer = origin(1).produce(); + let consumer = producer.consume(); + let broad = producer.dynamic("", Route::default().with_hops(hops(&[10]))).unwrap(); + + let pending = consumer.request_broadcast("jobs/a"); + let served = broadcast::Info::new().produce(); + queued(&broad).await.accept(&served); + let resolved = pending.await.expect("the broad route serves it"); + + // The same publisher claims a narrower prefix, which wins selection and refuses. + let narrow = producer + .dynamic("jobs", Route::default().with_hops(hops(&[10]))) + .unwrap(); + queued(&narrow).await.reject(Error::NotFound); + + tokio::time::timeout(Duration::from_secs(5), resolved.closed()) + .await + .expect("the narrower refusal must end the front, not keep the broader route"); + assert!( + broad.poll_requested_broadcast(&kio::Waiter::noop()).is_pending(), + "the refusal fell through to the broader route" + ); + } + + #[tokio::test] + async fn refusal_while_serving_does_not_move_to_a_sibling() { + let producer = origin(1).produce(); + let consumer = producer.consume(); + let broad = producer.dynamic("", Route::default().with_hops(hops(&[10]))).unwrap(); + + let pending = consumer.request_broadcast("jobs/a"); + let served = broadcast::Info::new().produce(); + queued(&broad).await.accept(&served); + let resolved = pending.await.expect("the broad route serves it"); + + // Two paths from the same publisher claim a narrower prefix; the cheaper refuses. + let cheap = producer + .dynamic("jobs", Route::default().with_hops(hops(&[10])).with_cost(1)) + .unwrap(); + let sibling = producer + .dynamic("jobs", Route::default().with_hops(hops(&[10])).with_cost(2)) + .unwrap(); + queued(&cheap).await.reject(Error::NotFound); + + tokio::time::timeout(Duration::from_secs(5), resolved.closed()) + .await + .expect("the refusal must end the front"); + assert!( + sibling.poll_requested_broadcast(&kio::Waiter::noop()).is_pending(), + "the refusal moved to a sibling in the tier" + ); + } + #[tokio::test] async fn most_specific_prefix_shadows() { let producer = origin(1).produce(); @@ -6494,6 +6699,176 @@ mod tests { pending.await.expect("re-request resolves through the rival"); } + /// A copy of "video" whose reply named `member`, as a Lite07 session's is. + fn vouching(source: &broadcast::Producer, member: u64) -> track::Producer { + let info = track::Info { + names_origin: true, + ..Default::default() + }; + let track = source.create_track("video", info).unwrap(); + track.provenance().name(origin(member)).unwrap(); + track + } + + /// Wait for the front to admit or refuse a vouching copy, which a Lite07 session + /// waits on before delivering anything. + async fn admission(track: &track::Producer) -> Result<(), Error> { + let provenance = track.provenance(); + let mut admitted = None; + settle(|| match provenance.poll_admitted(&kio::Waiter::noop()) { + Poll::Ready(result) => { + admitted = Some(result); + true + } + Poll::Pending => false, + }) + .await; + admitted.unwrap() + } + + /// A pool front: "room/alice" served through a route labelled `first` (at + /// cost 5) by a copy whose reply named `member`, one "before" group delivered. + async fn pool_rig(first: u64, member: u64) -> (ResumeRig, Dynamic, broadcast::Producer) { + let producer = origin(1).produce(); + let consumer = producer.consume(); + let server = producer + .dynamic("room", Route::default().with_hops(hops(&[first])).with_cost(5)) + .unwrap(); + + let pending = consumer.request_broadcast("room/alice"); + let request = queued(&server).await; + let source = broadcast::Info::new().produce(); + let track = vouching(&source, member); + request.accept(&source); + + let resolved = pending.await.expect("resolves"); + let mut subscription = resolved + .track("video") + .unwrap() + .subscribe(None) + .await + .expect("subscribe"); + admission(&track) + .await + .expect("the first reply names the front's origin"); + let mut group = track.append_group().unwrap(); + group.write_frame(crate::Timestamp::ZERO, b"before".as_ref()).unwrap(); + group.finish().unwrap(); + let mut group = subscription + .recv_group() + .await + .expect("recv group") + .expect("track ended early"); + let frame = group.read_frame().await.expect("read frame").expect("frame"); + assert_eq!(&frame.payload[..], b"before"); + + let rig = ResumeRig { + producer, + resolved, + subscription, + incumbent_track: track, + }; + (rig, server, source) + } + + /// A downstream relay reaches a pool through advertisements whose first hop + /// labels the pool, not the member serving the path. The reply names the + /// member, so a failover through a route with another label still resumes + /// when the replacement names the same member, and the relay names that + /// member in its own replies. + #[tokio::test] + async fn failover_resumes_on_the_origin_the_reply_names() { + let (mut rig, incumbent, source) = pool_rig(10, 20).await; + assert_eq!(rig.resolved.origin(), Some(origin(20))); + + let standby_server = rig.standby(&[11]); + drop(incumbent); + drop(source); + + let request = queued(&standby_server).await; + let replacement = broadcast::Info::new().produce(); + let track = vouching(&replacement, 20); + request.accept(&replacement); + admission(&track).await.expect("the same origin is admitted"); + let mut group = track.append_group().unwrap(); + group.write_frame(crate::Timestamp::ZERO, b"before".as_ref()).unwrap(); + group.finish().unwrap(); + let mut group = track.append_group().unwrap(); + group.write_frame(crate::Timestamp::ZERO, b"resumed".as_ref()).unwrap(); + group.finish().unwrap(); + + let mut group = rig + .subscription + .recv_group() + .await + .expect("subscription survives the failover") + .expect("track ended early"); + let frame = group.read_frame().await.expect("read frame").expect("frame"); + assert_eq!(&frame.payload[..], b"resumed"); + assert_eq!(rig.resolved.origin(), Some(origin(20))); + } + + /// Two pool members behind the same advertised label are different content: + /// a failover never splices one member's frames onto another's. + #[tokio::test] + async fn failover_never_splices_another_pool_member() { + let (mut rig, incumbent, source) = pool_rig(10, 20).await; + let standby_server = rig.standby(&[10]); + drop(incumbent); + drop(source); + rig.incumbent_track.abort(Error::Dropped).unwrap(); + + let request = queued(&standby_server).await; + let rival = broadcast::Info::new().produce(); + let track = vouching(&rival, 21); + request.accept(&rival); + assert!( + admission(&track).await.is_err(), + "another member's content is never admitted" + ); + + let err = rig.subscription.recv_group().await.err().expect("subscription ends"); + assert!(matches!(err, Error::Dropped), "unexpected end: {err}"); + } + + /// A front serving a named origin does not trade its live source for a + /// better route: which member the new route leads to is unknown until it + /// replies. Newcomers join it rather than starting a second copy. + #[tokio::test] + async fn a_named_origin_stays_on_its_live_route() { + let (rig, _incumbent, _source) = pool_rig(10, 20).await; + let cheaper = rig + .producer + .dynamic("room", Route::default().with_hops(hops(&[11])).with_cost(1)) + .unwrap(); + + let consumer = rig.producer.consume(); + let joined = consumer.request_broadcast("room/alice").await.expect("joins"); + assert!(joined.is_clone(&rig.resolved), "a newcomer joins the live front"); + for _ in 0..100 { + tokio::task::yield_now().await; + } + assert!( + cheaper.poll_requested_broadcast(&kio::Waiter::noop()).is_pending(), + "the front must not re-request through the new route" + ); + } + + /// A front names the origin of what it serves in its replies: this origin's hop + /// for content originating here, and a random hop of its own for content nobody + /// identifies, never 0. + #[tokio::test] + async fn a_front_names_an_origin_for_its_content() { + let (rig, _server, _source) = ResumeRig::new(&[0]).await; + let anonymous = rig.resolved.origin().expect("a front names its origin"); + assert_ne!(anonymous, Hop::UNKNOWN); + assert_ne!(anonymous, origin(1)); + + // Content originating here is named by this origin. + let (rig, _server, _source) = ResumeRig::new(&[]).await; + assert_eq!(rig.resolved.origin(), Some(origin(1))); + } + #[tokio::test] async fn anonymous_routes_never_resume() { // An empty hop chain identifies nobody, so two of them must not pass for diff --git a/rs/moq-net/src/model/track.rs b/rs/moq-net/src/model/track.rs index e15fd8786a..41bedc971f 100644 --- a/rs/moq-net/src/model/track.rs +++ b/rs/moq-net/src/model/track.rs @@ -106,6 +106,10 @@ pub struct Info { /// The publisher's priority for this track, used only to break ties between /// subscriptions of equal subscriber priority. Reported in TRACK_INFO (Lite05+). pub priority: u8, + /// Whether a copy's replies name the origin serving it (SUBSCRIBE_OK and FETCH_OK + /// on Lite07+), recorded in its [`Provenance`]. The origin's failover only + /// splices a copy that can vouch for the front's origin this way. + pub(crate) names_origin: bool, } impl Default for Info { @@ -114,6 +118,7 @@ impl Default for Info { timescale: Timescale::default(), max_age: DEFAULT_MAX_AGE, priority: 0, + names_origin: false, } } } @@ -141,6 +146,89 @@ impl Info { } } +/// Which origin a copy's content comes from: the one its replies named, and the one +/// its consumer admits. +/// +/// A relay's front splices copies from several sources into one logical track, and +/// only copies of one origin may join. A session whose replies name the serving +/// origin records each reply here and holds the copy's content until the front +/// admits that origin, so a copy of another origin never delivers anything. +#[derive(Clone, Default)] +pub(crate) struct Provenance(kio::Shared); + +#[derive(Default)] +struct ProvenanceState { + named: Option, + admit: Option, + refused: bool, +} + +impl ProvenanceState { + fn admitted(&self) -> bool { + self.named.is_some() && self.named == self.admit + } +} + +impl Provenance { + /// Record the origin a reply named. A copy has one origin for its whole life, so + /// a reply naming another means the upstream is serving a different track. + pub(crate) fn name(&self, origin: crate::Hop) -> Result<()> { + let mut state = self.0.lock(); + match state.named { + Some(named) if named != origin => Err(Error::Dropped), + Some(_) => Ok(()), + None => { + state.named = Some(origin); + Ok(()) + } + } + } + + /// Wait for a reply to name the origin. + pub(crate) fn poll_named(&self, waiter: &kio::Waiter) -> Poll { + self.0 + .poll(waiter, |state| match state.named { + Some(_) => Poll::Ready(()), + None => Poll::Pending, + }) + .map(|state| state.named.expect("predicate guaranteed a name")) + } + + /// The origin whose content may be delivered. Replaces an earlier one: a front + /// learns its origin from the first reply, after it had only the route's label. + pub(crate) fn admit(&self, origin: crate::Hop) { + if self.0.read().admit != Some(origin) { + self.0.lock().admit = Some(origin); + } + } + + /// Nothing more of this copy will be admitted: its consumer let it go. + pub(crate) fn refuse(&self) { + if !self.0.read().refused { + self.0.lock().refused = true; + } + } + + /// Whether the named origin is admitted now, for content that cannot wait for it. + pub(crate) fn is_admitted(&self) -> bool { + self.0.read().admitted() + } + + /// Wait until the named origin is admitted, or fail once the copy is refused. + /// Admission wins: a refusal only stops what was still waiting. + pub(crate) fn poll_admitted(&self, waiter: &kio::Waiter) -> Poll> { + self.0 + .poll(waiter, |state| match state.admitted() || state.refused { + true => Poll::Ready(()), + false => Poll::Pending, + }) + .map(|state| match state.admitted() { + true => Ok(()), + false => Err(Error::Dropped), + }) + } +} + #[derive(Default)] pub(crate) struct TrackState { // The publisher's properties, once known; always Some for Subscriber/Producer. @@ -259,6 +347,10 @@ pub(crate) struct TrackState { // The reverse fetch queue (see [`FetchState`]), same reasoning: cache-miss // `fetch_group` calls enqueue here and a `Dynamic` drains. fetch: kio::Shared, + + // Which origin this copy's content comes from, same reasoning: the session + // records replies, the consumer admits. + provenance: Provenance, } /// A cached group plus its bookkeeping in the track's `lookup` map. @@ -1790,6 +1882,11 @@ impl Producer { }) } + /// Which origin this copy's content comes from; see [`Provenance`]. + pub(crate) fn provenance(&self) -> Provenance { + self.state.read().provenance.clone() + } + /// Create a [`Dynamic`] handle that serves on-demand fetches of uncached /// (old) groups. Most producers never need this; a relay creates one to fetch /// past groups from upstream. @@ -2505,6 +2602,15 @@ impl Consumer { } } + /// Which origin this copy's content comes from; see [`Provenance`]. `None` for a + /// spliced logical track, which is made of copies. + pub(crate) fn provenance(&self) -> Option { + match &self.inner { + ConsumerKind::Plain(state) => Some(state.read().provenance.clone()), + ConsumerKind::Spliced(_) => None, + } + } + /// The newest group, when it is already cached: resolved synchronously, without /// counting as a fetch or a delivery. The IETF publisher snapshots its frame count to /// resolve Largest Object; a group that is not immediately available reads as no edge. diff --git a/rs/moq-net/tests/datagram.rs b/rs/moq-net/tests/datagram.rs index 4cc78b1297..3cd596cf25 100644 --- a/rs/moq-net/tests/datagram.rs +++ b/rs/moq-net/tests/datagram.rs @@ -35,6 +35,10 @@ struct Fixture { /// A publisher and a subscriber joined over the mock transport, sharing one /// datagram-carrying track. async fn connect_datagram_track() -> Fixture { + connect_datagram_track_on("moq-lite-05").await +} + +async fn connect_datagram_track_on(version: &str) -> Fixture { let publisher = produce_origin(1); let consumer_origin = produce_origin(2); @@ -42,7 +46,7 @@ async fn connect_datagram_track() -> Fixture { let producer = broadcast.create_track("datagrams", None).unwrap(); broadcast.announce(Default::default()).unwrap(); - let mut options = MockConnectOptions::new("moq-lite-05".parse::().unwrap()); + let mut options = MockConnectOptions::new(version.parse::().unwrap()); options.server_publish = Some(publisher.consume()); options.client_subscribe = Some(consumer_origin.clone()); let pair = connect_mock(options).await; @@ -110,6 +114,34 @@ async fn datagrams_reach_the_subscriber_in_order() { .expect("timed out"); } +/// Lite-07 holds a subscription's content until SUBSCRIBE_OK names its origin, so a +/// track that only ever sends datagrams must still send SUBSCRIBE_OK for any to arrive. +/// Datagrams racing ahead of it are dropped, so this keeps sending until one lands. +#[tokio::test] +async fn datagram_only_track_names_its_origin_on_lite07() { + tokio::time::timeout(TEST_TIMEOUT, async { + let mut fixture = connect_datagram_track_on("moq-lite-07-wip").await; + + for sequence in 0.. { + fixture + .producer + .insert_datagram( + sequence, + Timestamp::from_millis(sequence).unwrap(), + bytes::Bytes::from_static(PAYLOAD), + ) + .unwrap(); + let arrived = tokio::time::timeout(Duration::from_millis(10), fixture.subscriber.recv_datagram()).await; + if let Ok(datagram) = arrived { + assert_eq!(&datagram.unwrap().unwrap().payload[..], PAYLOAD); + break; + } + } + }) + .await + .expect("no datagram arrived: SUBSCRIBE_OK never named the origin"); +} + /// MoQ Transport has no datagram mapping: groups still flow, inserted datagrams do not. #[tokio::test] async fn ietf_does_not_deliver_datagrams() { diff --git a/rs/moq-net/tests/pool_failover.rs b/rs/moq-net/tests/pool_failover.rs new file mode 100644 index 0000000000..d70083ba65 --- /dev/null +++ b/rs/moq-net/tests/pool_failover.rs @@ -0,0 +1,117 @@ +//! A relay spreading one prefix across a pool of workers, seen from downstream. +//! +//! Workers claim the same prefix, so the pool relay advertises one route for all +//! of them, labelled by whichever member ranks first for the prefix. Which member +//! serves a path is the relay's per-path choice, so the label cannot say what a +//! downstream relay is receiving; the SUBSCRIBE_OK and FETCH_OK replies do. When +//! the serving worker dies, the pool relay re-serves the path from another worker, +//! and the downstream relay must end its subscription rather than splice that +//! worker's frames onto the first one's. + +mod support; + +use std::time::Duration; + +use moq_net::{Hop, Timestamp, Version, origin}; +use support::harness::{MockConnectOptions, connect_mock}; + +const TIMEOUT: Duration = Duration::from_secs(10); + +fn produce_origin(hop: u64) -> origin::Producer { + let (producer, driver) = origin::Producer::new(origin::Config::new(Hop::new(hop).unwrap())); + tokio::spawn(support::harness::run(driver)); + producer +} + +/// A worker claiming the "pool" prefix: every request is answered with a fresh +/// broadcast whose "video" track keeps producing groups carrying the worker's name. +fn worker(hop: u64) -> origin::Producer { + let producer = produce_origin(hop); + let dynamic = producer.dynamic("pool", origin::Route::default()).unwrap(); + let name = format!("w{hop}").into_bytes(); + tokio::spawn(async move { + while let Ok(request) = dynamic.requested_broadcast().await { + let broadcast = moq_net::broadcast::Info::new().produce(); + let track = broadcast.create_track("video", None).unwrap(); + request.accept(&broadcast); + let name = name.clone(); + tokio::spawn(async move { + let _broadcast = broadcast; + loop { + let Ok(mut group) = track.append_group() else { return }; + group.write_frame(Timestamp::ZERO, name.clone()).unwrap(); + group.finish().unwrap(); + tokio::time::sleep(Duration::from_millis(5)).await; + } + }); + } + }); + producer +} + +#[tokio::test] +async fn downstream_failover_never_splices_another_pool_member() { + tokio::time::timeout(TIMEOUT, async { + let version: Version = "moq-lite-07-wip".parse().unwrap(); + let pool = produce_origin(10); + let downstream = produce_origin(30); + + // Each worker publishes its claim to the pool relay. + let mut members = Vec::new(); + for hop in [20, 21] { + let producer = worker(hop); + let mut options = MockConnectOptions::new(version); + options.client_publish = Some(producer.consume()); + options.server_subscribe = Some(pool.clone()); + members.push((hop, producer, connect_mock(options).await)); + } + + // The downstream relay reaches the pool through one session. + let mut options = MockConnectOptions::new(version); + options.server_publish = Some(pool.consume()); + options.client_subscribe = Some(downstream.clone()); + let _link = connect_mock(options).await; + + let consumer = downstream.consume(); + consumer.routed("pool/job").await.unwrap(); + let remote = consumer.request_broadcast("pool/job").await.unwrap(); + let mut subscription = remote.track("video").unwrap().subscribe(None).await.unwrap(); + + let mut group = subscription.recv_group().await.unwrap().expect("a first group"); + let first = group.read_frame().await.unwrap().expect("a frame").payload.to_vec(); + + // Kill whichever worker the pool relay picked for this path. + let serving = members + .iter() + .position(|(hop, ..)| first == format!("w{hop}").into_bytes()) + .expect("a pool member served the path"); + let (_, _, pair) = members.remove(serving); + pair.client.abort(moq_net::Error::Cancel); + + // Every group this subscription delivers is the first worker's; it ends + // instead of carrying on with the survivor's. + while let Ok(Some(mut group)) = subscription.recv_group().await { + while let Some(frame) = group.read_frame().await.unwrap_or(None) { + assert_eq!(frame.payload.to_vec(), first, "spliced another pool member's frames"); + } + } + + // A fresh request is served by the survivor. + let (hop, ..) = &members[0]; + let survivor = format!("w{hop}").into_bytes(); + let remote = consumer.request_broadcast("pool/job").await.unwrap(); + let track = remote.track("video").unwrap(); + let mut subscription = track.subscribe(None).await.unwrap(); + let mut group = subscription.recv_group().await.unwrap().expect("a group"); + let frame = group.read_frame().await.unwrap().expect("a frame"); + assert_eq!(frame.payload.to_vec(), survivor); + + // A group from before the subscription is fetched through both relays, each + // FETCH_OK naming the survivor. + let mut fetched = track.fetch_group(0, None).await.unwrap(); + let frame = fetched.read_frame().await.unwrap().expect("a frame"); + assert_eq!(frame.payload.to_vec(), survivor); + }) + .await + .expect("timed out"); +}