chore: backport #2425 and #2393 to v1 - #2439
Conversation
### 💡 Overview The SFU's Twirp requests carry no client identity today - the SDK name and version only appear in the `JoinRequest` and `SendStats` payloads. This adds an `X-Stream-Client` header to every SFU RPC, mirroring what we already send to the coordinator. This gap surfaced during the Android 1.31.0 incident, where `startNoiseCancellation` raced ahead of the WS join, the SFU replied `ERROR_CODE_PARTICIPANT_NOT_FOUND`, and the SDK treated it as terminal and looped. The SFU couldn't gate a server-side mitigation on the client version because the RPC didn't say who was calling. The SFU already makes decisions based on the client string (codec selection, for one), so having it on every RPC lets the backend diagnose and mitigate client-specific regressions without waiting for an SDK release. ### 📝 Implementation notes The header carries the same client id shape the coordinator already receives through the user agent: | SDK | `X-Stream-Client` | | --- | --- | | React | `stream-video-react-v1.43.1` | | React Native | `stream-video-react-native-v1.44.2` | | everything else | `stream-video-js-v<version>` | `getStreamClientId` lives next to the existing `getSdkName` / `getSdkVersion` helpers so the `SdkType` mapping stays in one place, and the version comes from the same source that feeds `JoinRequest` and the stats payloads. `setSdkInfo` runs at module init in both `react-sdk/index.ts` and `react-native-sdk/src/index.ts`, well before any SFU client is constructed at join time; with SDK info unset the value degrades to `stream-video-js-v0.0.0-development`. The WS `/ws` endpoint already carries `cid`, `user_id` and `api_key` as query params, so nothing changes there. [Format agreed with Marcelo in the thread](https://getstream.slack.com/archives/C040262MY9K/p1788870872729689). **Backend prerequisite:** `X-Stream-Client` is a non-safelisted header on a cross-origin Twirp POST, so the SFU's CORS `Access-Control-Allow-Headers` must include it before this ships - otherwise the preflight blocks every SFU RPC from the web SDK. **Known inconsistency:** the coordinator derives its own name from `SdkType[type].toLowerCase()`, so it sees `react_native` / `plain_javascript` where this header says `react-native` / `js`. Aligning the coordinator side is out of scope here. The other platforms (Android, iOS, Flutter, Unity) need the same header to close the gap fleet-wide; that's tracked separately per SDK team. 🎫 Ticket: https://linear.app/stream/issue/REACT-1169/identify-the-sdk-and-version-on-sfu-twirp-requests-x-stream-client 📑 Docs: n/a - internal transport header, no public API change. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * SFU requests now include a standardized SDK client identifier in their headers, including the platform and version (for example, React or React Native). * SDK identifiers now consistently fall back to a JavaScript identifier when platform information is unavailable or unsupported. * Added coverage to verify SDK identifier formatting and SFU request headers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> (cherry picked from commit 2312002)
) `call.accepted` / `call.rejected` / `call.missed` reach clients only over the coordinator WebSocket, best-effort and with no store-and-forward. The WS pings every 25s while the ring window is 30s, so a silently dead socket may never even be *detected* inside the window. The caller sits on a ringing screen while the callee is already in the call alone. This is the client half of D8: the caller reconciles by pull. After a quiet period with no ring event it reads `GET /call/{type}/{id}/ring_state` every 5s until the ring settles or the window closes — join when somebody accepted, cancel when everyone rejected, leave when the call already ended. "Nobody answered" stays with the local auto-drop. A new `src/ringing/` holds the four pieces, one file each: - **`RingStatePoller`** — the loop. Caller-only, on by default; disable or tune with `StreamClientOptions.ringStatePolling` (`false | { startAfterMs, intervalMs }`, defaults 15s / 5s). An incoming ring event restarts the quiet period rather than killing the poller, since in a group ring one rejection does not settle it. Bounded by `auto_cancel_timeout_ms` so it always resolves before the local auto-drop. - **`reconcileRingState`** — the single place a ring outcome is decided, shared by the WebSocket handlers and the poller. Takes no event and no payload: it reads `call.state` and branches on caller vs callee. `call_ended_at` / `session_ended_at` is checked *before* `accepted_by`, so an already-ended session is never joined. - **`RingTimeout`** — the ringing auto-drop, moved out of `Call` because it is the same shape as the poller. It also refuses to arm without ring settings; the old code returned early on a missing settings object but then read `settings.ring` unguarded. - **`resolveOwnRingOutcome`** — the current user's *own* accept or reject, which may have landed on another device. `reconcileAsCallee` used to defer to an inline effect in `Call.registerEffects`, so the callee's rules were split across two files. `Call.getRingState(callSessionId?)` wraps the endpoint; the session id defaults to the current session, and is passed explicitly to read one that has already ended. `CallState.updateFromRingState` merges the polled maps into the matching session so `session$` subscribers see what the dropped event carried. The response type comes from `src/gen/coordinator` — only `GetCallRingStateResponse` was taken from the regenerated output. The outcome used to be decided twice: `watchCallAccepted` / `watchCallRejected` read it off the event payload, and the poller had its own copy. They could drift, and the poller's copy only covered the caller. `CallState.updateFromEvent` is an `all` listener and `dispatchEvent` drains those before the typed ones, so by the time a ring handler runs the event's session is already on the call state. That makes the payload redundant, and both paths now run the same state-driven reconciler — the handlers rely on `updateFromEvent`, the poller applies the polled state itself. Two consequences worth review: - **A failed join keeps the ring open.** `doJoin` restores the ringing state when a join fails, so `RINGING -> JOINING -> RINGING`. Treating that first transition as the end of the ring left the call ringing with no auto-drop and no poller — the exact failure this PR exists to fix. `JOINING` is now transient for the poller, `RingTimeout` checks the state when its timer fires, and a failed join is reported non-terminal so the next poll retries. - **The callee branch is explicit**, and leave failures are caught rather than escaping the listener as an unhandled rejection. Acting twice on one outcome is prevented by the `RINGING` check in `reconcileRingState`, `singleFlight` on `Call.join`, and `leave()`'s own `LEFT` check. No event-identity dedup needed. **Drive-by fix:** `RingCallEvents` extracted from `AllClientCallEvents` — an object type, not a string union — so it resolved to `never`, the mapped registry type collapsed to `{}`, and the handler registry was unchecked. `call.missed` was registered nowhere as a result. It stays unregistered here, matching `main`: nothing in the reconciler acts on `missed_by`, so `RingTimeout` owns the "nobody answered" case until the server-owned ring timeout (VID-1427) lands. The polled `missed_by` map still reaches `session$` for consumers, and a `call.missed` event still restarts the poller's quiet period, since it proves the socket is alive. The watchdogs used to be cancelled outright before `leave()`'s first await, so a leave that failed left the call ringing with neither of them running. `RingTimeout` gained a `pause()` that keeps its deadline, and `leave()` resumes both watchdogs from its rejection handler when the call is still `RINGING` and nothing else is queued on the join/leave tag - so a transient `reject()` failure does not silently strand the ring, and the resumed timeout fires at its original deadline rather than extending the window. Now that two paths can pull the caller into a call, the tracing has to say which one did — a poll-driven join is itself a dropped-event signal, and without the distinction there is no way to measure how often the poller rescues a ring the socket lost. `CoordinatorJoin` client events gain a `source` field alongside `join_reason`: `ring-ws` from the event handlers, `ring-poll-api` from the poller. - It travels on a new local-only `joinSource` option on `Call.join()`, destructured out so it never reaches the coordinator with the rest of `JoinCallData`. `join` is wrapped in `singleFlight`, so a second concurrent trigger has its argument discarded along with the rest — the source always names the angle that *first* reached `join()`, which is the intended reading. - The reporter mirrors `join_reason` throughout: a per-cid map, a snapshot on the stage pair, and the same three `CoordinatorJoin` emissions, always read off the snapshot so all three agree. It is scoped to one join lifecycle by clearing it both on entry when absent and in a `finally` — the first stops a *nested* lifecycle inheriting it (a reconnect can open one during `join()`'s inter-attempt sleep, when `doJoin` has restored the pre-join state), the second stops a *later* `CoordinatorJoin` reported outside any lifecycle, which is the shape of a fast reconnect. - `source` is not in the coordinator OpenAPI schema yet, so it is typed by a local extension of the generated `ClientEvent` rather than by hand-editing generated code. The backend ignores unknown fields, so reporting it early is dropped rather than rejected — **but if the server lands the field as `join_source` instead, the client keeps sending `source` and the data silently never arrives.** Worth settling the name before the server half ships. **Merge note:** `main` moved `callingX.endCall(call, 'remote')` out of `watchCallEnded` and into `leave()`, fired off the new `reason: 'ended'`. The reconciler's two explicit `endCall` calls are gone in favour of that, otherwise every ended ring would have fired it twice. **Dogfood:** the web dialer gets a Pronto-only pane showing the live ring state next to an on-demand read of the endpoint, and learns `?coordinator_url` / `?use_local_coordinator` (like the home and join pages) plus `?call_id` / `?call_type` to pin a ring to one call. The RN app gets the same pane over the ringing call, call type / call id inputs to pin the ring, and a coordinator-URL override plus a polling on/off switch in the env switcher — both persisted, so the client built for a background push uses them too. The endpoint is live on the default edge; exercised by hand against it, the response matches the generated type field for field and the pane populates on ring and resets on cancel. The `source` field is covered at the unit level only — reproducing a *dropped* WS event end-to-end needs the coordinator socket suppressed while HTTP stays alive, so the assertions are on the outgoing `POST /call_client_event` body rather than on a dashboard. The server discards the field until its half ships anyway. 🎫 Ticket: https://linear.app/stream/issue/VID-1444/pollable-ring-state-the-caller-reconciles-a-dropped-ring-outcome-by 🔗 Backend: GetStream/chat#16110 📑 Docs: pending — the customer-facing doc must carry the re-ring staleness caveat (until VID-1322) and the double-join guard recipe. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **New Features** - Added automatic ring-outcome polling for outgoing calls. - Calls now respond appropriately when participants accept, reject, miss, or end a call. - Added configurable polling timing, including an option to disable polling. - Added a development-only panel for inspecting ring state and manually refreshing results. - **Bug Fixes** - Improved synchronization between ring outcomes and call state, including missed and ended calls. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Backport note: adapted to `release-v1`, which does not have the `ClientState` refactor (#2422). `clientState` -> `clientStore`, `ClientState` -> `StreamVideoWriteableStateStore`, and the new `CallState.updateFromRingState` uses the `this.setCurrentValue` class field instead of the module-level import. No behaviour change. (cherry picked from commit 73bfc98)
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Bundle sizeBuilt package output. Sizes in KB; delta vs
|
`rtc/e2ee/workerMessages.ts` was swept into the #2393 squash on `main` and rode along into the backport. It is unrelated to ring-state polling and nothing imports it on either branch, so it is dead code. Removing it here keeps the unused `WORKER_TIMEOUT_MS` runtime constant out of the v1 build. The file stays on `main`, where the in-flight E2EE work is expected to pick it up.
💡 Overview
Backports the two merged
backport-v1PRs torelease-v1, in merge order:feat(client): identify the SDK and version on SFU RPCs- cherry-picked clean.feat(client): poll ring state to reconcile a dropped ring outcome- adapted, see below.rtc/e2ee/workerMessages.ts, see below.Both commits keep their original author and carry a
(cherry picked from commit ...)trailer pointing at the commit onmain.Important
Merge this with Rebase and merge, not squash. Squashing collapses both backports into a single commit and replaces the body with this PR description, which drops the cherry-pick trailers and the original authorship.
📝 Implementation notes
release-v1does not have #2422 (feat(client): replace the two client state stores with a single ClientState), which is not flagged for backport. Every conflict in #2393 traced back to that refactor, so the backport adapts to the old API instead of pulling the refactor in:clientState->clientStoreandClientState->StreamVideoWriteableStateStoreinCall.tsand the affected tests.CallState.updateFromRingStatecalls thethis.setCurrentValueclass field rather than feat(client): replace the two client state stores with a single ClientState #2422's module-level import. This hunk merged without a conflict and produced code that does not compile (Cannot find name 'setCurrentValue'); onlytsccaught it.The adaptation is naming only, no behaviour change.
packages/client/src/ringing/*is byte-identical tomain, andCall.tsdiffers frommainonly by the rename plus one pre-existing JSDoc line that belongs torelease-v1.#2393 also added
packages/client/src/rtc/e2ee/workerMessages.ts, which is unrelated to ring-state polling and has no importers on either branch. It appears to have been swept into that PR's squash from the in-flight E2EE work. A third commit removes it from the backport so the unusedWORKER_TIMEOUT_MSruntime constant does not reach the v1 build; the file is untouched onmain, where the E2EE work should pick it up or clean it up.Apart from that removal, #2393 is carried whole, including its react-dogfood and RN dogfood changes, so the commit matches the original. Happy to strip it to
packages/**if we would rather keep sample apps off the release branch.✅ Verification
tsc --noEmitonpackages/client: clean.StreamVideoClient*integration test files needSTREAM_SECRETand fail to load locally without it; they run in CI.prettier --checkandeslint --max-warnings 0on all 39 changed files: clean.🎫 Ticket: n/a (backport)