diff --git a/quest/m0/remove-gossip.md b/quest/m0/remove-gossip.md index ecf38db9b7..80b0635a49 100644 --- a/quest/m0/remove-gossip.md +++ b/quest/m0/remove-gossip.md @@ -40,4 +40,4 @@ announcing `.internal/origins`. ## Related -- [Cluster routing](/quest/m1/cluster-routing/README.md) - takes its topology from configured links only +- [Cluster topology](/quest/m1/cluster-routing/topology.md) - takes its topology from configured links only, so it waits on this diff --git a/quest/m0/wildcard/README.md b/quest/m0/wildcard/README.md index 0b106cd870..a41beec582 100644 --- a/quest/m0/wildcard/README.md +++ b/quest/m0/wildcard/README.md @@ -24,7 +24,7 @@ is the catch-all `**`, a claim about every path at once. The cost of enumerating is real even though its last measurement is stale. One announcement measured 8.8 KB per relay plus 4.3 KB per additional route before prefix routes made a standby route a table entry; -[Cluster routing](/quest/m1/cluster-routing/README.md) owns remeasuring it. Whatever the current +[Cluster routing's memory benchmark](/quest/m1/cluster-routing/memory.md) remeasures it. Whatever the current number, every relay that hears an announcement pays it whether or not anything there subscribes, so "workers times broadcasts" is that number multiplied across the fleet in resident memory. diff --git a/quest/m1/README.md b/quest/m1/README.md index b560c4d583..d4df9c9fdb 100644 --- a/quest/m1/README.md +++ b/quest/m1/README.md @@ -22,7 +22,8 @@ QUIC studies there on that rule. ## Required -- [Cluster routing](/quest/m1/cluster-routing/README.md) - an announcement says where a broadcast originates, not how to reach it, and a relay hears only the prefixes its clients asked for +- [Cluster idle timeout](/quest/m1/cluster-idle-timeout.md) - a relay notices a silent peer relay within seconds, not after the shared 30 s QUIC idle timeout +- [Cluster routing](/quest/m1/cluster-routing/README.md) - a relay learns the relay graph once and each announcement once, not once per neighbour, and redundant publishers share an epoch instead of `--hop` - [Track tail interop](/quest/m1/track-tail-interop.md) - a Rust publisher ending a track with a group in flight is read to its end by the JS subscriber, and the reverse, in `just test interop` - [Merge queue](/quest/m1/merge-queue.md) - the required checks run on `merge_group`, so a stale green check can no longer break main - [Binding audio delay](/quest/m1/binding-surface.md) - moq-ffi and every wrapper configure and observe audio playout delay @@ -49,7 +50,7 @@ QUIC studies there on that rule. - [TS PSI reassembly](/quest/m1/ts-psi-reassembly.md) - `import ts` reads a PAT or PMT that spans packets or follows a nonzero pointer_field instead of aborting, and one corrupted section costs a repetition and a counted `CRC_error`, not the import - [RTMP interleaving](/quest/m1/rtmp-interleaving.md) - isolate partial messages before optimizing assembly copies - [TS stats module](/quest/m1/ts-stats-module.md) - on dev, the TS stats types move under `ts::stats` as `Snapshot` and `Stream`, with an owned `track` -- [Same-hop importers](/quest/m1/hop-aligned-import.md) - importers sharing a `--hop` and fed one stream publish identical groups and timestamps, so failover survives +- [Same-hop importers](/quest/m1/hop-aligned-import.md) - importers fed one stream publish identical groups and timestamps, so failover between a redundant pair survives - [Audio capture without ALSA link](/quest/m1/capture-alsa-link.md) - moq-audio capture and playback build on Linux without linking libasound - [Capture by default](/quest/m1/capture-default.md) - moq-video and moq-audio build `capture` by default, so pre-merge checks test it and the capture gate goes away - [Ship capture and playback](/quest/m1/cli-packaging.md) - a released moq binary can capture and play, which no distribution currently enables diff --git a/quest/m1/broadcast-epoch/README.md b/quest/m1/broadcast-epoch/README.md index 96a0172d3e..0b315118b6 100644 --- a/quest/m1/broadcast-epoch/README.md +++ b/quest/m1/broadcast-epoch/README.md @@ -14,9 +14,10 @@ The epoch rides in the path, so it survives any moq-transport relay, and no wire message changes. At an epoch-aware relay, a request for a bare name resolves to its newest live epoch on every protocol version. -Non-goals: redundant publishers sharing one epoch, and failing over between -them faster than the keep-alive (a question -[Cluster routing](/quest/m1/cluster-routing/README.md) owns). Also out of scope: trusting the publisher's clock (a far-future epoch +Non-goals: pooling redundant publishers that share one epoch, which +[Cluster routing](/quest/m1/cluster-routing/README.md)'s selection and +`--hop` removal own; this line only keeps an epoch a caller supplies. Also out +of scope: trusting the publisher's clock (a far-future epoch wins until its route goes away). ## Plan diff --git a/quest/m1/cluster-idle-timeout.md b/quest/m1/cluster-idle-timeout.md new file mode 100644 index 0000000000..5bbf7d739b --- /dev/null +++ b/quest/m1/cluster-idle-timeout.md @@ -0,0 +1,29 @@ +# [S] Cluster idle timeout + +## Goal + +A relay notices a silent peer relay within seconds, not after the 30 s QUIC +idle timeout it shares with every viewer. moq.pro's routing simulator found +that failure detection, not routing, sets every outage window: a silent link +or relay loss goes unnoticed for the idle timeout in every routing design, +and subscribes through it go nowhere meanwhile. + +## Plan + +- Cluster sessions get their own idle timeout and keep-alive, shorter than + the `--quic-*` defaults (`rs/moq-tokio/src/quic.rs`, 30 s idle and 5 s + keep-alive). A keep-alive alone only keeps a quiet session open; the + timeout is what detects loss. +- QUIC's effective idle timeout is the smaller of the two endpoints' + (RFC 9000 section 10.1), so the dialing relay's setting bounds both sides of + a cluster link. Say so where the setting is documented. +- Pick the default in the PR with its reasoning: loss detection against + spurious drops on a lossy intercontinental link. +- Docs: `doc/bin/relay/cluster.md` and `doc/bin/relay/config.md`. + +Public API: adds a relay config field; lands on `main`. + +## Related + +- [Cluster routing](/quest/m1/cluster-routing/README.md) - its liveness and failover depend on this detection +- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - viewers wait out the same timeout for an old route diff --git a/quest/m1/cluster-routing/README.md b/quest/m1/cluster-routing/README.md index 821bd5f080..ee1ec30ac9 100644 --- a/quest/m1/cluster-routing/README.md +++ b/quest/m1/cluster-routing/README.md @@ -1,29 +1,63 @@ -# [XL] Cluster routing +# Cluster routing ## Goal -A broadcast event reaches each relay at most once, and a relay learns only the -prefixes its own clients asked for. An announcement says a path exists at an -origin relay, at a cost; how to reach that origin comes from a shared relay -topology, so no announcement inside a cluster carries a hop list. A relay's -memory scales with what it serves, not with what the mesh knows. This -questline records the design; its wire and implementation quests are planned -from the simulator's report. +Cut cluster gossip. A relay learns how to reach each other relay once, from +a topology shared across the cluster rather than repeated in every route, +and learns each announcement once rather than once per neighbour. Every relay +still holds the ledger of live announcements that ANNOUNCE_REQUEST and a cold +SUBSCRIBE need, and a broadcast under overlapping prefixes routes to one +origin deterministically. Redundancy an operator configures may deliver more +than one copy. Non-goals: warm re-origination (a warm relay would be one more origin with a cost, so leave room for it), and a permanently mixed-version cluster. ## Plan -This is a questline with no children yet. Its children (wire and -implementation) are planned next in a separate `/quest-plan` session, from -the decisions and open questions below. +### Decisions + +Settled in the 2026-09-30 `/quest-plan`: + +- The goal stays wide. The design below is the current candidate, not a + decision; [Propagation](/quest/m1/cluster-routing/propagation.md) settles + how announcements move and writes the implementation children that follow. +- Topology is a cluster message kept apart from routes, built from configured + links only (`--cluster-connect`, the connect API, LAN mDNS). Gossip + discovery goes first, but only the topology child waits on it. +- Every relay holds a record for every live announcement, since + ANNOUNCE_REQUEST and a SUBSCRIBE with no prior announce both need one. + Scoping that knowledge by demand needs registries, which Propagation may + defer to m2. +- `--hop` goes. An epoch-qualified concrete path (`foo/@`) is the + identity of a source: origins that announce the same one are + interchangeable, which is how a redundant pair is expressed. It is strictly + better than a Hop ID, which is per session, so a connection could not + publish several broadcasts with different identities. A path a claim + produces keeps Wildcard's per-origin identity, even once announced + concretely, since each worker's output is its own. +- The line lands on `dev`: deleting `--hop` and the publisher's Hop setup + parameter breaks a published CLI and wire. Wire changes go in the current + wip version (`moq-lite-07-wip` today, dropping `Hop Base` and `Hop Keep` + before they publish). If lite-07 is finalized first for the Wildcard + rollout, the remaining children move to the next wip version; finalizing + never waits on this line. +- The before/after memory figure is a committed benchmark, measured first. +- A short idle timeout for cluster sessions is its own m1 quest, + [Cluster idle timeout](/quest/m1/cluster-idle-timeout.md), since it helps + today's path vector too. +- Reduced flooding ([RFC 9667](https://www.rfc-editor.org/rfc/rfc9667)) and + registries are m2 unless Propagation pulls registries in; Propagation writes + whichever it defers. +- Between clusters, announcements stay path vector with cluster ids as hops, + in the last child. ### Why not path vector or Babel Today every relay advertises its best route to every peer not already in the -hop chain. One publish costs about R·(d-1) announces for R relays of mesh -degree d, every relay learns every broadcast (`.stats` and `.internal` +hop chain, and pulls every peer's full table with an empty-prefix +ANNOUNCE_REQUEST. One publish costs about R·(d-1) announces for R relays of +mesh degree d, every relay learns every broadcast (`.stats` and `.internal` included), and a link change rewrites every route crossing it. On moq.pro's live fleet (26 PoPs, average degree about 5) each relay receives every event about five times. Babel ([RFC 8966](https://www.rfc-editor.org/rfc/rfc8966)) @@ -42,7 +76,12 @@ the loops. Hiding the routes through a peer that withdrew the path ([#4399](https://github.com/moq-dev/moq/pull/4399)) reduces that path hunting but does not end it; see the findings below. -### Decisions +### Candidate design + +The children own their parts: topology in +[Topology](/quest/m1/cluster-routing/topology.md), existence and registries +in [Propagation](/quest/m1/cluster-routing/propagation.md), origin choice in +[Selection](/quest/m1/cluster-routing/selection.md). - Existence is split from reachability. An announcement carries the path, its origin relay, and the origin's cost, and nothing about the path to it. @@ -59,76 +98,28 @@ but does not end it; see the findings below. dropped. Purging on a liveness report instead drops a restarted origin's broadcasts until their re-announce lands, which the simulator showed as live broadcasts reported offline. -- The topology is configured: `--cluster-connect` or the connect API gives the - relay graph and link costs. Relays flood per-link liveness among themselves - with a per-link seqno. The seqno is scoped to the relay's incarnation, so a - restarted relay's links supersede its stale ones instead of looking older. - Configured links are the only source: gossip discovery is removed by - [Remove gossip](/quest/m0/remove-gossip.md), and LAN mDNS dials peers - that then count as configured links. - - A relay batches the liveness reports it sends, its own and those it - forwards, for a short hold-down (50 ms in the simulator), and recomputes - its trees after a matching delay, as OSPF's SPF delay does. Unbatched, one - relay restart at 340 relays sent half a million messages; batched, 27k. - - On session up, relays exchange a digest (each reporter's incarnation and - its seqno per link) and send only what the other lacks. A reporter's - reports flood separately, so its newest seqno alone would hide a missing - older one. The digest already counts the fresh report of the link that - just came up; sending the database first replays the relay's own report - from when the link went down, and the peer drops the link it is using. +- Relays flood per-link liveness among themselves with a per-link seqno scoped + to the relay's incarnation, batched for a short hold-down, with a digest + exchange on session up. - A relay picks the origin with the lowest shortest-path distance plus origin cost, ties broken by rendezvous hashing (HRW) of the requested path and the - origin id, and forwards along its shortest path. Distance compares cost, - then hop count, so every hop strictly shortens it even across `?cost=0` - links. That is a shortest path to a virtual node linked to every origin, so - it is loop-free whenever relays agree on the topology. The longest covering + origin id, and forwards along its shortest path. The longest covering prefix still ranks first, per [Wildcard](/quest/m0/wildcard/README.md). The - simulator saw no loop while views agreed, and HRW split an equal-cost pool - 63/49 where today's hash of the announced prefix sends all of it to one - sibling. -- The first relay's choice rides the SUBSCRIBE, and transit relays forward - toward that origin by topology alone, never re-selecting. Re-selection - against another existence view loops: a relay that lost a specific claim - falls back to a broader one through a relay still routing to the specific - one ([RFC 8966 section 3.5.4](https://www.rfc-editor.org/rfc/rfc8966#section-3.5.4)). - If the origin no longer serves the path, it refuses, and the first relay - selects again. -- SUBSCRIBE and FETCH carry a visited-relay list end to end. It catches loops - while liveness views disagree and names the path for stats. Narrowing it to - cluster hops is later work. The serving origin's identity rides the reply, - per Wildcard's Spread quest. On live's graph no disagreement looped across - 20 seeds of link, cost, and relay churn; the list is a safety net. -- Announcements are on demand. A relay forwards only the union of its clients' - ANNOUNCE_REQUEST prefixes, never the empty prefix. A wide prefix that many - edges' viewers request is that customer's cost. `.stats` becomes ordinary - demand. + first relay's choice rides the SUBSCRIBE, and transit relays forward toward + that origin by topology alone. SUBSCRIBE and FETCH carry a visited-relay + list as a loop safety net. +- Announcements to clients are on demand: a relay forwards only the union of + its clients' ANNOUNCE_REQUEST prefixes. `.stats` becomes ordinary demand. - Registries are an optional, configured tier: moq-relay in a registry mode, - one or more per region. - - An ingest relay registers its broadcasts with its nearest registry, and an - edge sends its ANNOUNCE_REQUEST there. - - Registries form a small full mesh and flood existence among themselves, so - an event crosses an ocean once per remote registry, not once per relay. - Announce latency is about one round trip to the nearest registry, - whatever the path length. - - A relay fails over to the next-nearest registry and reconciles its view - instead of treating the lost session as ends, so a registry failure never - reports a live broadcast offline. - - The reconcile is the relay's full live set at its current seqno, which - ends whatever it leaves out. The registry's snapshot carries each - relevant origin's last reconcile seqno, and the relay names the origins it - already holds, so a view kept through a freeze ends what ended meanwhile. - - With no registry reachable, a relay freezes: it keeps its view, learns - nothing new, and alerts. Falling back to flooding would cascade the - failure. -- Without registries (self-hosting), existence floods along the shortest-path - tree, one copy per relay. A relay forwards an event only when it changes its - view, so a duplicate copy, from trees built on disagreeing liveness, stops - there. A relay that gains a child in its tree (a view change or a new - session) pushes that origin's reset and records to it; without the push a - relay misses events while views disagree. -- Between clusters, announcements stay path vector with cluster ids as the - hops, like BGP between autonomous systems. A customer's on-prem cluster is - one hop, and an announcement naming the receiving cluster is dropped. + one or more per region. Registries form a small full mesh and flood + existence among themselves, so an event crosses an ocean once per remote + registry. A relay fails over to the next-nearest registry and reconciles; + with none reachable it freezes its view and alerts rather than falling back + to flooding. +- Without registries, existence floods along the shortest-path tree, one copy + per relay. A relay forwards an event only when it changes its view, and a + relay that gains a child in its tree pushes that origin's reset and records + to it. - The cluster switches versions as a whole; older lite and IETF sessions stay at its edges. @@ -138,7 +129,7 @@ The report is moq.pro's `just rs sim` (every scenario on live's graph) and `just rs sim sweep` (synthetic regional graphs of 34, 340, and 1020 relays). It carries messages and bytes by kind, per relay and cross-region, convergence, loop, stall, and failover windows, and state per relay and per -registry. What decides the wire: +registry. - Existence costs about one message per relay per event flooded, and one per interested relay plus one per remote registry with registries. On live, a @@ -150,13 +141,16 @@ registry. What decides the wire: loss and restore plus a relay loss and restart cost about 290 MiB of liveness flooding and 54 MiB of session-up digests against under 1 MiB of existence for twenty publishes, while a publish costs exactly one message - per relay. A reduced flooding topology - ([RFC 9667](https://www.rfc-editor.org/rfc/rfc9667)) is the known fix for - the flooding. + per relay. +- Unbatched, one relay restart at 340 relays sent half a million liveness + messages; batched for 50 ms, 27k. - Failure detection, not routing, sets every outage window: a silent link or relay loss is noticed after the 30 s QUIC idle timeout in every candidate, - and subscribes through it go nowhere until then. Cluster sessions need a - short idle timeout; a keepalive only keeps a quiet session open. + and subscribes through it go nowhere until then. +- The simulator saw no loop while views agreed, and HRW split an equal-cost + pool 63/49 where today's hash of the announced prefix sends all of it to one + sibling. On live's graph no disagreement looped across 20 seeds of link, + cost, and relay churn. - One registry per region and two cost about the same; two halves the busiest registry's load. A registration sent on a dead registry session that nobody has noticed yet waits for the failover. @@ -172,76 +166,31 @@ registry. What decides the wire: the same scenario costs 18.6M messages instead of 189k. Per-origin seqnos, above, end both. -### Memory before and after - -On-demand announcements bound the route table by demand; measure that saving -before and after the implementation lands. Keep two costs apart: the route -table and per-announcement state (`RouteEntry` plus `ServeState`) scale with -announcements and routes, while the served-content cache a `ServeState` -materializes scales with demand. - -Every published figure is stale. The old baseline was 8.8 KB per announced -broadcast plus 4.3 KB per extra route on `adad52b`, measured with two -throwaway `moq-net` examples driving an origin under a counting allocator and -reading `/proc/self/statm`. Since then -[moq#2989](https://github.com/moq-dev/moq/pull/2989) cut `kio`'s inline waiter -slots from 32 to 4 (a `kio::State<()>` went from 896 B to about 200 B, and -`kio`'s `tests/waiter_allocs.rs` pins that lever as spent), and -[moq#3225](https://github.com/moq-dev/moq/pull/3225) made a standby route a -table entry rather than an object graph. Neither example is committed, since -they need `#[doc(hidden)]` size probes on private types. Rebuild them and -restate the per-broadcast and per-route cost, the per-peer session -bookkeeping (`announce_ids`, `held`, `watched`), and the shed threshold on a -degree-5, 1 GB node. Chat-shaped traffic (one broadcast per channel or per -chatter) depends on the answer. - -### Open questions - -- Sharding registries by HRW over a prefix key once one registry cannot hold - everything, and what that key is. -- A mixed-version bridge, if a fleet cannot switch at once. -- How long a relay keeps an ended path's seqno. A new origin incarnation - clears it; within one, it must outlive every delayed copy of the start. -- How a push to a new tree child ends what the child missed within the same - incarnation. Its reset only marks a new incarnation, so the push may need a - watermark like the registry reconcile's: it ends only what the child holds - at or below it, and a delayed push cannot erase newer records. -- How long a relay keeps the records of an origin that never comes back. They - stop being reported once it is unreachable, but nothing removes them. -- How an edge routes a SUBSCRIBE for a path none of its clients asked to - announce. It holds no route for it, and asking a registry first adds a round - trip before the first byte. -- What a cold ANNOUNCE_REQUEST reports as live. Answering from the local view - keeps a relay from waiting on peers but reports an empty set until the - registry's replay lands, on every new prefix rather than in a rare race. -- What remains of announce compression's hop-tail half (`Hop Base` and - `Hop Keep` in the lite draft) once only cluster boundaries carry hops. -- What replaces `--hop` first-hop failover. Today two publishers sharing a Hop - ID are one source that relays fail over between at a group boundary - (`doc/bin/cli.md` "Redundant publishers", - `doc/concept/use-case/contribution.md`). Inside a cluster no announcement - carries a hop list, so two encoders on different ingest relays become two - origins. Keep the documented behavior or change the docs in the same PR. - Say whether same-hop semantics survive, or - [Same-hop importers](/quest/m1/hop-aligned-import.md)' work is thrown - away. The redundant ingest study folded in here: whether the pair claims one - `@` epoch, what enforces the aligned groups and matching catalog the - docs only ask for, and who declares the incumbent dead early (a failover - service that retracts it, or active-active delivery to the relay). Weigh it - against the moq-transport rule that each publisher of a namespace must be - asked (#3697). The answer may be a no-go. -- Whether equal-cost next hops should spread by a hash of the path. A fixed - tie-break sends every path through the same neighbour and its failure takes - them all. +### Remaining work + +Once every child has landed: + +- An end-to-end test: a multi-relay cluster in `rs/moq-relay/tests` where a + publish, an end, a relay restart, and a redundant-pair failover each reach + a subscriber on a far relay, with time mocked. +- Rewrite `doc/bin/relay/cluster.md` into the operator's view (topology, + link costs, idle timeout, redundant pairs), and add a routing page under + `doc/concept`. ## Required -- [Wildcard](/quest/m0/wildcard/README.md) - the longest-prefix rule, pool spread, and reply identity this selection builds on +- [Memory benchmark](/quest/m1/cluster-routing/memory.md) - a committed benchmark states per-announcement, per-route, and per-peer relay memory, before anything changes +- [Topology](/quest/m1/cluster-routing/topology.md) - relays learn the relay graph once from a cluster message, apart from routes +- [Propagation](/quest/m1/cluster-routing/propagation.md) - decides how each announcement reaches every relay once, and writes the implementation children +- [Selection](/quest/m1/cluster-routing/selection.md) - a broadcast under overlapping prefixes routes to one origin deterministically, and same-epoch origins are one source +- [Remove `--hop`](/quest/m1/cluster-routing/hop-removal.md) - redundant publishers share an explicit `@`, and `--hop` and the publisher's Hop ID are gone +- [Between clusters](/quest/m1/cluster-routing/inter-cluster.md) - announcements crossing a cluster boundary stay path vector with cluster ids as hops ## Related -- [Same-hop importers](/quest/m1/hop-aligned-import.md) - builds on `--hop` failover, which this line must keep or replace -- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - a redundant pair would share one epoch - +- [Wildcard](/quest/m0/wildcard/README.md) - the longest-prefix rule, pool spread, and reply identity Selection builds on +- [Cluster idle timeout](/quest/m1/cluster-idle-timeout.md) - failure detection sets every outage window, today and after this line +- [Same-hop importers](/quest/m1/hop-aligned-import.md) - the importer half of a redundant pair; `--hop` removal re-keys it to a shared epoch +- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - a redundant pair shares one epoch - [Cross-relay delivery under bursts](/quest/m1/cross-relay-bursts.md) - its #4349 report also shows closed broadcasts announced for up to 229 s and flapping between Retracted and Announced across nodes, evidence for per-incarnation seqnos - [Routing cost domains](/quest/m3/routing-cost-domains.md) - cost across the cluster boundaries this keeps path vector diff --git a/quest/m1/cluster-routing/hop-removal.md b/quest/m1/cluster-routing/hop-removal.md new file mode 100644 index 0000000000..2dcdf8221f --- /dev/null +++ b/quest/m1/cluster-routing/hop-removal.md @@ -0,0 +1,45 @@ +# [M] Redundant publishers share an epoch, not a hop + +## Goal + +Redundant publishers share an explicit `@` instead of a `--hop`, and +the publisher-facing Hop ID is gone: `moq`'s `--hop` / `MOQ_HOP` and the +deprecated `--origin` spelling, and the Hop setup parameter a publisher +sends. An epoch is strictly better, since a Hop ID is per session: one +connection cannot publish several broadcasts with different identities. + +## Plan + +- The pool semantics already live in + [Selection](/quest/m1/cluster-routing/selection.md): origins announcing the + same epoch-qualified concrete path are one source. This quest deletes what + that makes dead, including `Pin::Publisher(Hop)` in + `rs/moq-net/src/model/front.rs` and `RouteEntry::qualifies`. +- A redundant pair passes one explicit epoch; the broadcast-publish path keeps + an epoch a caller supplies ([Broadcast epochs](/quest/m1/broadcast-epoch/README.md)). + Give `moq` a flag for it if Broadcast epochs has not. +- Relays keep `cluster.id` as their identity; `--hop` fills it today + (`rs/moq-cli/src/args.rs`), so split that coupling. Decide whether relays + still need the Hop setup parameter or learn identity from topology. +- Re-key [Same-hop importers](/quest/m1/hop-aligned-import.md)'s docs and 1+1 + tests to a shared epoch. +- Docs: rewrite "Redundant publishers" in `doc/bin/cli.md` and + `doc/concept/use-case/contribution.md`, then search the repo for `--hop`, + `MOQ_HOP`, and `--cluster-id` examples, including demo recipes. + +Open: whether anything declares an incumbent dead faster than the cluster +idle timeout (a failover service that retracts it, or active-active delivery +to the relay). Neither is required to land this. + +Public API and wire: removes a CLI flag and the publisher's Hop setup +parameter; lands with the line on `dev`. + +## Required + +- [Selection](/quest/m1/cluster-routing/selection.md) - owns the same-epoch pool this relies on +- [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - a restart is a new path, not a splice +- [Same-hop importers](/quest/m1/hop-aligned-import.md) - the importer fixes this re-keys + +## Related + +- [Cluster idle timeout](/quest/m1/cluster-idle-timeout.md) - bounds how long a dead incumbent holds its pair diff --git a/quest/m1/cluster-routing/inter-cluster.md b/quest/m1/cluster-routing/inter-cluster.md new file mode 100644 index 0000000000..43386d65ef --- /dev/null +++ b/quest/m1/cluster-routing/inter-cluster.md @@ -0,0 +1,40 @@ +# [M] Routing between clusters + +## Goal + +Announcements crossing a cluster boundary stay path vector with cluster ids +as the hops, like BGP between autonomous systems. A customer's on-prem +cluster is one hop, and an announcement naming the receiving cluster is +dropped. Inside a cluster nothing carries a list of relay hops. + +## Plan + +- A boundary is a configured link to a relay of another cluster; decide how + a relay knows its own cluster id and its peer's. +- What crosses a boundary is what the chosen + [Propagation](/quest/m1/cluster-routing/propagation.md) design holds, plus + its cluster-id list. An imported record keeps that list as provenance, so + whichever boundary relay exports it appends its own cluster id to the full + list; stripping it on import would let A → B → C → A re-enter A. +- The remote origin is outside the receiving cluster's topology, so the + importing boundary relay announces the record inside its cluster as the + origin, with the boundary link's cost added. The serving origin's identity + still rides the reply, so re-originating does not merge two sources. +- Cluster-id lists are short, so they need no `Hop Base`/`Hop Keep` + compression. +- Cost across the boundary is plain configured link cost; business policy + and incomparable costs are + [Routing cost domains](/quest/m3/routing-cost-domains.md)'s. + +Wire: the boundary announcement in the current wip lite version, with the +draft updated in the same PR. Tests: two clusters with two boundary links +and an origin that is not a boundary relay, and three clusters in a cycle with different ingress and egress relays, where +an announcement neither loops nor re-enters its origin cluster. + +## Required + +- [Propagation](/quest/m1/cluster-routing/propagation.md) - the record shape a boundary exports + +## Related + +- [Routing cost domains](/quest/m3/routing-cost-domains.md) - designs policy on this path vector diff --git a/quest/m1/cluster-routing/memory.md b/quest/m1/cluster-routing/memory.md new file mode 100644 index 0000000000..a97d2e7b38 --- /dev/null +++ b/quest/m1/cluster-routing/memory.md @@ -0,0 +1,46 @@ +# [S] Relay memory benchmark + +## Goal + +A committed benchmark states what a relay's memory costs per announced +broadcast, per extra route, and per peer session, so the rest of +[Cluster routing](/quest/m1/cluster-routing/README.md) has a before figure and +reports its after figure against the same code. Chat-shaped traffic (one +broadcast per channel or per chatter) and Wildcard's "workers times +broadcasts" argument both depend on the answer. + +## Plan + +Every published figure is stale. The old baseline was 8.8 KB per announced +broadcast plus 4.3 KB per extra route on `adad52b`, measured with two +throwaway `moq-net` examples driving an origin under a counting allocator and +reading `/proc/self/statm`. Since then +[moq#2989](https://github.com/moq-dev/moq/pull/2989) cut `kio`'s inline waiter +slots from 32 to 4, and [moq#3225](https://github.com/moq-dev/moq/pull/3225) +made a standby route a table entry rather than an object graph. Neither +example was committed, since they needed `#[doc(hidden)]` size probes on +private types. + +- Measure through the public API with a counting allocator (as + `rs/moq-net/benches/session.rs` already does for groups), not private size + probes, so the benchmark survives the redesign. +- Sweep both axes: broadcasts announced, and peers each one is routed through, + so a cost that grows with the table rather than the touched path shows as a + slope. +- Keep two costs apart: the route table and per-announcement state + (`RouteEntry`, `ServeState` in `rs/moq-net/src/model/origin.rs`) scale with + announcements and routes, while the served-content cache a `ServeState` + materializes scales with demand. Per-peer bookkeeping is the lite + publisher's `live` map and `AnnounceEncoder` entries, the lite subscriber's + `Announced.routes`, and the IETF `watched` and `held`. +- Run it in the nightly benchmark job + ([Benchmark regressions in CI](/quest/m1/bench-ci.md)) or its own nightly + step if that has not landed. +- Record the before figures in the PR and in the questline README, and derive + the shed threshold on a degree-5, 1 GB node from them. The implementation + child that replaces per-peer routes reports the after figures. + +## Related + +- [Wildcard](/quest/m0/wildcard/README.md) - cites the stale figure +- [Benchmark regressions in CI](/quest/m1/bench-ci.md) - where the benchmark runs nightly diff --git a/quest/m1/cluster-routing/propagation.md b/quest/m1/cluster-routing/propagation.md new file mode 100644 index 0000000000..1614b590ab --- /dev/null +++ b/quest/m1/cluster-routing/propagation.md @@ -0,0 +1,73 @@ +# [M] Plan announcement propagation + +## Goal + +Decide how an announcement reaches every relay in a cluster once, rather than +once per neighbour, and how every relay keeps the ledger of live announcements +that ANNOUNCE_REQUEST and a cold SUBSCRIBE need. The output is a decision +recorded in the [questline README](/quest/m1/cluster-routing/README.md), +worked counterexamples, and rewritten implementation quests, not production +code. + +## Plan + +Decided: every relay holds a record for every live announcement, and an +announcement names its origin relay and cost, not a path to it +([Topology](/quest/m1/cluster-routing/topology.md) supplies reachability). + +Candidates to choose between or combine: + +- **Tree flooding with per-origin seqnos.** The README's candidate design: + existence floods along the origin's shortest-path tree, a relay forwards an + event only when it changes its view, and per-origin seqnos with + incarnation resets stop stale revivals. This is the same as "skip a + neighbour already closer to the origin": a relay forwards only to its + children in that origin's tree. +- **A central node owns gossip.** Registries: an optional tier, one or more + per region, full-meshed among themselves. An ingest relay registers with + its nearest, an edge sends its ANNOUNCE_REQUEST there, and a relay fails + over to the next-nearest and reconciles (its full live set at its current + seqno, ending whatever it leaves out; the registry's snapshot carries each + origin's last reconcile seqno). With none reachable a relay freezes its + view and alerts. The simulator found registries buy cross-ocean bytes and + one-round-trip announce latency, not message count, at live's 34 relays. +- **A compressed ledger.** The records are a replicated ledger of every live + announcement; sync it with digests (anti-entropy) and compress it, such as + prefix-shared paths, instead of per-event messages. + +Measure the choice with the simulator's scenarios (messages and bytes per +event, convergence, stale revival, restart, failover) and the +[Memory benchmark](/quest/m1/cluster-routing/memory.md)'s figures. + +Open questions the chosen design must answer: + +- How long a relay keeps an ended path's seqno. A new origin incarnation + clears it; within one, it must outlive every delayed copy of the start. +- How a relay that just became another's tree child catches up on what it + missed within the same incarnation. Its reset only marks a new incarnation, + so the push may need a watermark like the registry reconcile's: it ends only + what the child holds at or below it, and a delayed push cannot erase newer + records. +- How long a relay keeps the records of an origin that never comes back. They + stop being reported once it is unreachable, but nothing removes them. +- Whether registries belong in this line. If deferred, write an m2 quest for + them with sharding (HRW over a prefix key) and a mixed-version bridge as its + open questions, and one for reduced flooding + ([RFC 9667](https://www.rfc-editor.org/rfc/rfc9667)). + +The implementation quests it writes replace the empty-prefix ANNOUNCE_REQUEST +between peers and per-hop path-vector announcements, drop `Hop Base` and +`Hop Keep` from the current wip lite version, and report the memory +benchmark's after figures. Add each one to the Required of +[Selection](/quest/m1/cluster-routing/selection.md) and +[Routing between clusters](/quest/m1/cluster-routing/inter-cluster.md), so +neither starts on a record shape that is not implemented yet. + +## Required + +- [Memory benchmark](/quest/m1/cluster-routing/memory.md) - the before figures the choice is weighed against + +## Related + +- [Cross-relay delivery under bursts](/quest/m1/cross-relay-bursts.md) - its report of closed broadcasts announced for minutes is evidence for per-origin seqnos +- [Announcement shapes](/quest/m2/announce-shapes.md) - announcements between relays must keep their shape diff --git a/quest/m1/cluster-routing/selection.md b/quest/m1/cluster-routing/selection.md new file mode 100644 index 0000000000..c75f2b5b81 --- /dev/null +++ b/quest/m1/cluster-routing/selection.md @@ -0,0 +1,72 @@ +# [L] Deterministic origin selection + +## Goal + +A broadcast under overlapping announcements routes to one origin chosen +deterministically, and every relay forwards toward that origin by topology. +Origins that announce the same epoch-qualified concrete path are one source, +so a subscriber moves between them at a group boundary when the incumbent +ends or becomes unreachable. + +## Plan + +Decided: + +- The longest covering prefix ranks first, per + [Wildcard](/quest/m0/wildcard/README.md), and origin choice is + per broadcast, not per announced prefix. +- An epoch-qualified concrete path (`foo/@`) is a source's identity: + every origin announcing it is interchangeable, which is how a redundant + pair works without `--hop`. The first relay fails over between them at a + group boundary. A path with no epoch, or one a claim produces, keeps + Wildcard's per-origin identity (the origin SUBSCRIBE_OK names), since two + workers' groups differ. That includes a claim worker's derived output once + it is announced concretely: it mirrors the input's epoch + (`.pro/transcode//foo.hang/@e`), so after a double claim two workers + announce the same epoch-qualified path and must not pool. This extends + Wildcard's resume rule to concrete same-epoch origins. + +Candidate mechanics: + +- A relay picks the origin with the lowest shortest-path distance plus origin + cost, ties broken by rendezvous hashing (HRW) of the requested path and the + origin id, and forwards along its shortest path. That is a shortest path to + a virtual node linked to every origin, so it is loop-free whenever relays + agree on the topology. +- The first relay's choice rides the SUBSCRIBE and FETCH, and transit relays + forward toward that origin by topology alone, never re-selecting. + Re-selection against another existence view loops: a relay that lost a + specific claim falls back to a broader one through a relay still routing to + the specific one + ([RFC 8966 section 3.5.4](https://www.rfc-editor.org/rfc/rfc8966#section-3.5.4)). + A refusal follows Wildcard's refusal rule: only a capacity refusal lets the + first relay select once more within the same longest-prefix tier, excluding + the refusing origin, and any other refusal is terminal. +- SUBSCRIBE and FETCH carry a visited-relay list end to end. It catches loops + while liveness views disagree and names the path for stats. The serving + origin's identity rides the reply, per Wildcard's first-hop resume rule. + +Open: + +- Whether equal-cost next hops should spread by a hash of the path. A fixed + tie-break sends every path through the same neighbour, and its failure + takes them all. +- How a relay tells claim output from a redundant pair. One candidate: a + concrete path an origin announces under its own claim keeps per-origin + identity, since a redundant publisher claims nothing. + +Wire: SUBSCRIBE and FETCH fields in the current wip lite version, with the +draft updated in the same PR. Tests cover an HRW split across an equal-cost +pool, refusal and reselection, a same-epoch pair failing over mid-track +with no timestamp rewind, and a concrete double claim whose loser's +subscribers end and resubscribe rather than splice. + +## Required + +- [Topology](/quest/m1/cluster-routing/topology.md) - distance to each origin +- [Propagation](/quest/m1/cluster-routing/propagation.md) - the record shape that names each origin +- [Wildcard](/quest/m0/wildcard/README.md) - the longest-prefix rule, pool spread, and reply identity this builds on + +## Related + +- [Epoch primitive](/quest/m1/epoch.md) - parses the `@` segment that makes a path a source identity diff --git a/quest/m1/cluster-routing/topology.md b/quest/m1/cluster-routing/topology.md new file mode 100644 index 0000000000..35adf5f6a4 --- /dev/null +++ b/quest/m1/cluster-routing/topology.md @@ -0,0 +1,61 @@ +# [L] Cluster topology + +## Goal + +Relays learn the relay graph once, from a cluster message kept apart from +routes, instead of every route repeating how to reach its origin. Each relay +knows every configured link, whether it is up, and its cost, and computes its +shortest path to every other relay. No announcement is needed to learn that +relay K exists or how to reach it. + +## Plan + +Decided: the topology is configured. `--cluster-connect` or the connect API +gives the relay graph and link costs, and LAN mDNS dials peers that then count +as configured links. Gossip discovery is gone first, by +[Remove gossip](/quest/m0/remove-gossip.md). + +Candidate mechanics, from the simulator (see the questline README's +findings): + +- Relays flood per-link liveness among themselves with a per-link seqno. The + seqno is scoped to the relay's incarnation, so a restarted relay's links + supersede its stale ones instead of looking older. +- A relay batches the liveness reports it sends, its own and those it + forwards, for a short hold-down (50 ms in the simulator), and recomputes its + shortest paths after a matching delay, as OSPF's SPF delay does. Unbatched, + one relay restart at 340 relays sent half a million messages; batched, 27k. +- On session up, relays exchange a digest (each reporter's incarnation and + its seqno per link) and send only what the other lacks. A reporter's + reports flood separately, so its newest seqno alone would hide a missing + older one. The digest already counts the fresh report of the link that just + came up; sending the database first replays the relay's own report from + when the link went down, and the peer drops the link it is using. +- Distance compares cost, then hop count, so every hop strictly shortens it + even across `?cost=0` links. + +Reduced flooding ([RFC 9667](https://www.rfc-editor.org/rfc/rfc9667)) is +out of scope; [Propagation](/quest/m1/cluster-routing/propagation.md) writes +it as an m2 quest if it stays deferred. + +Open: moq-transport cluster peers. Cluster links accept moq-transport 17+ +today via the cluster extension (`drafts/draft-lcurley-moq-cluster.md`), +which these messages do not reach. +Either extend that draft, or refuse a cluster session that does not negotiate +the wip lite version, per the README's "the cluster switches versions as a +whole". Never accept a peer that cannot carry the topology. + +Wire: new cluster-session messages in the current wip lite version, with the +draft updated in the same PR. Tests drive topologies in process with mocked +time, including restart, a link flapping, and a digest racing a link that +just came up. Expose the computed graph where operators already look +(`/nodes` in `rs/moq-relay/src/internal.rs`). + +## Required + +- [Remove gossip](/quest/m0/remove-gossip.md) - configured links are the only topology source + +## Related + +- [Cluster idle timeout](/quest/m1/cluster-idle-timeout.md) - liveness only reports what the session layer detects +- [Drain](/quest/m1/drain/README.md) - a second relay per PoP joins this topology diff --git a/quest/m1/hop-aligned-import.md b/quest/m1/hop-aligned-import.md index 9b34dd42b0..250076067b 100644 --- a/quest/m1/hop-aligned-import.md +++ b/quest/m1/hop-aligned-import.md @@ -13,7 +13,11 @@ exits (#4354). ## Plan Decided: same-hop publishers MUST publish the same broadcasts and tracks. -Keep `--hop`, and make every container importer (ts, fmp4, flv, mkv, and the +Keep `--hop` for now: the users' bugs are on one relay today. Cluster +routing replaces it with a shared explicit epoch on `dev` (decided +2026-09-30), and its [`--hop` removal](/quest/m1/cluster-routing/hop-removal.md) +re-keys this quest's docs and tests. The importer work holds under either +key. Make every container importer (ts, fmp4, flv, mkv, and the SRT, RTMP, and HLS gateways that reuse them) meet the contract when fed one encoded stream. Capture is out: two encoders never align. @@ -40,5 +44,5 @@ survives the standby joining and the incumbent stopping. ## Related -- [Cluster routing](/quest/m1/cluster-routing/README.md) - drops hop lists inside a cluster and must say whether same-hop failover survives; it also owns splicing across first hops and two encoders, which this does not attempt +- [Remove `--hop`](/quest/m1/cluster-routing/hop-removal.md) - on dev, a redundant pair shares an explicit epoch instead, and this quest's docs and tests are re-keyed to it - [Broadcast epochs](/quest/m1/broadcast-epoch/README.md) - a redundant pair shares one epoch diff --git a/quest/m3/routing-cost-domains.md b/quest/m3/routing-cost-domains.md index a21210a687..f32615b6ae 100644 --- a/quest/m3/routing-cost-domains.md +++ b/quest/m3/routing-cost-domains.md @@ -56,7 +56,7 @@ quests. Open wire/API choices belong to this design exercise. ## Required -- [Cluster routing](/quest/m1/cluster-routing/README.md) - this designs on its inter-cluster path vector +- [Routing between clusters](/quest/m1/cluster-routing/inter-cluster.md) - this designs on its inter-cluster path vector ## Related