fix(web): server voice plays again (media-src blob:) - #52
NetDevAutomate wants to merge 5 commits into
Conversation
Reported: the header badge says "Kokoro (server)", nothing is ever heard, and OpenVox appears to see nothing. Reproduced in Chromium against the real local OpenVox by clicking the voice toggle: /api/tts/health, /warm and /speak all answered 200 (speak with audio/wav), then the browser refused the audio: Loading media from 'blob:http://127.0.0.1:.../...' violates the following Content Security Policy directive: "default-src 'self'". Note that 'media-src' was not explicitly set, so 'default-src' is used as a fallback. play() rejected with NotSupportedError and the engine swallowed it. The same page with the policy bypassed played to the end. The server tier plays the host's WAV through URL.createObjectURL (since adf7493, 23 Aug); the policy (`default-src 'self'`, no media-src) arrived on 2 Sep in 1f10f05, and 'self' does not match a blob: URL. Why nothing caught it: the only browser test of this path needs a running OpenVox, so it skips in CI (which has none), and its "played" check counted play() calls rather than playback. RED, each for the stated reason: - the CSP has no media-src allowing blob:, in both modes (2 unit tests); - with the host's answers faked at the network edge, so CI runs it, the "Voice enabled" audio is refused by the page's own policy; - a play() the page refuses produces no notice, so the learner cannot tell a refused audio from a dead server; - a host that answers health in 4s is treated as absent. Measured here: OpenVox's /v1/models takes 1.2-2.7s, so 5 of 8 /api/tts/health calls ran past the page's 2.5s probe budget, and 1 page load in 5 fell back to system voices with a healthy host; - the real-OpenVox browser test now asserts the audio played (it fails on the same policy refusal). Guard that passes now and must keep passing: blob: appears in no directive but media-src (2 unit tests).
The server tier plays the host's WAV through `new Audio(URL.createObjectURL(blob))`. Since 2 Sep (1f10f05) the page's policy has been `default-src 'self'` with no media-src, and 'self' does not match a blob: URL, so the browser refused every utterance: /api/tts/speak answered 200 audio/wav, play() rejected with NotSupportedError, and the learner heard nothing while the badge said "Kokoro (server)". media-src now allows 'self' and blob:, in both modes. The exception is media-only (a unit guard fails if blob: appears in any other directive): a blob: URL can only be minted by this origin's own script, and script-src stays 'self' 'unsafe-eval', so no new script source is admitted. Chosen over a data: URL (broader: any markup can carry one) and over streaming the audio from a GET URL (the element could no longer read a 503 and fall back to system voices). Verified in Chromium: the faked-host playback test and the four CSP unit tests pass; the whole header suite (test_web_app.py, 32) passes. The real-OpenVox browser test resolved the system-voice tier on this run -- the 2.5s health budget, fixed separately below, not the policy.
_speakServer resolved a refused play() and an element error in silence, so the policy regression above showed as a dead speech server: the badge stayed "Kokoro (server)", OpenVox had answered every request, and the only trace was one console line. The learner, reasonably, went looking at OpenVox. A refused or undecodable utterance now raises the settings store's existing `tts:engine-notice` toast -- a listener nothing in the engine had emitted since the in-browser neural tier was removed. One notice per utterance (the error event and the rejected play() arrive together), and none for audio that stop() or a newer utterance cut short, which rejects play() with AbortError. NotAllowedError and NotSupportedError get a plain-language hint; anything else is named as it is. Engine notices now stay up 8s rather than the toast's 2s default, which is gone before the sentence is read. Unchanged: no tier switch on a refusal (falling back to system voices is a separate decision, and they need the same user activation), and the fetch-failure path still only logs. Verified: the refusal test passes (undecodable bytes, faked host); the engine e2e file (13) and the JS unit suite (167) pass.
The page gave /api/tts/health 2.5s before treating the host as absent. OpenVox's own /v1/models takes 1.2-2.7s on this Mac, and the health route waits on it. Measured on 29 Sep against the real OpenVox: 5 of 8 health calls ran past 2.5s, and 1 page load in 5 resolved to system voices with a healthy host -- a badge that changes between reloads for no visible reason. The real-OpenVox browser test landed on system voices in one run for the same reason. The budget is now 8s. It still ends a hung host's probe (the point of the bound: an init that never resolves looks broken), and leaves room for a tablet's LAN hop. A speech server that is not running on this machine refuses the connection at once (health against a closed local port: 13ms), so the budget changes nothing in the common "no server" case. Verified: after the change, 10 of 10 page loads resolved the server tier against the same OpenVox (health 2.4-2.6s on that run, 6 of 8 over the old budget); the server-TTS e2e file passed 11/11 twice with OpenVox running, and the engine e2e file 13/13.
CHANGELOG [Unreleased] Fixed: the media-src policy regression (2 Sep to now), the silent playback refusal, and the 2.5s health budget. The voice guide gains a Playback notice bullet under Controls: the badge names the tier, not whether this device could play a given utterance. Verified: mkdocs build --strict (docs extra) and the docs tests (524 passed, 3 skipped).
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The implementation addresses the regression with narrowly scoped policy changes and comprehensive deterministic coverage.
Review effort: Balanced
Findings: None
What changed in this PR
Restores server-side web voice playback and improves failure visibility and health-probe reliability.
Changes:
- Allows
blob:audio through the CSP’smedia-src. - Surfaces playback failures via longer-lived toast notifications.
- Adds deterministic browser coverage and extends the server health timeout.
| File | Description |
|---|---|
CHANGELOG.md |
Documents voice playback fixes. |
docs/voice-output.md |
Documents playback failure notices. |
packages/studyloop/src/studyloop/web/app.py |
Permits blob media in CSP. |
packages/studyloop/src/studyloop/web/static/components.js |
Extends TTS notice visibility. |
packages/studyloop/src/studyloop/web/static/tts-engine.js |
Reports playback failures and extends probe timeout. |
packages/studyloop/tests/e2e/test_server_tts.py |
Adds browser-level playback regressions. |
packages/studyloop/tests/test_web_app.py |
Verifies CSP scope in both modes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
Full unit suite (default markers) on e659f8c + docs, this Mac: 7401 passed, 31 failed, 14 errors, 16 skipped, in 15m54s. The 45 failing ids against the recorded machine-specific set ( |
Web voice has been silent since 2 Sep. The badge said Kokoro (server), OpenVox answered every request, and nothing was heard.
Root cause
The server voice tier plays the host's WAV through
new Audio(URL.createObjectURL(blob))(since adf7493, 23 Aug). On 2 Sep, 1f10f05 set the page's policy todefault-src 'self'with nomedia-src.'self'does not match ablob:URL, so the browser refused every utterance:play()rejected withNotSupportedError, and_speakServerswallowed it. Reproduced in Chromium against the real local OpenVox by clicking the header voice toggle:/api/tts/health,/warmand/speakall returned 200 (speak withaudio/wav, 66 KB), and then the audio was refused. The same page with the policy bypassed played to the end.Changes (test-first: RED d7e0821)
media-src 'self' blob:in both modes. It applies to media only; a unit guard fails ifblob:appears in any other directive. Ablob:URL can only be created by this origin's own script, andscript-srcstays'self' 'unsafe-eval'. I chose this over adata:URL (broader, because any markup can carry one) and over a GET audio URL (the element could no longer read a 503 and fall back to system voices).tts:engine-noticetoast, which nothing in the engine had emitted since the neural tier was removed. It fires once per utterance, and not for audio thatstop()or a newer utterance interrupted. Engine notices now stay up 8s instead of 2s./v1/modelstakes 1.2 to 2.7s. Measured: 5 of 8 health calls went past 2.5s, and 1 page load in 5 fell back to system voices while the host was healthy. After the change, 10 of 10 loads got the server tier. A speech server that isn't running still fails at once (13ms against a closed local port).Why nothing caught it
The only browser test of this path needs a running OpenVox, so it skips in CI, which has none. Its "played" check also counted
play()calls rather than playback. With OpenVox up it did fail, but only on its console-error check. The new tests fake the host's answers at the network edge withpage.route, so the page's real CSP header and the realtts-engine.jsare what gets tested, and CI runs them:The real-OpenVox browser test now asserts that the audio played, and it drives the header toggle the way a learner does.
Tested
tests/e2e/test_server_tts.py: 11/11, twice, with OpenVox running. On RED: 3 new tests failed plus the real-OpenVox test, each for the stated reason.tests/e2e/test_web_tts_engine.py: 13/13. JS unit suite: 167/167.test_web_app.py: 32/32.mkdocs build --strictand the docs tests (524 passed, 3 skipped).Not verified: tablets. The iPad and Android devices listed under Verified devices were checked in August, before the regression. This fix was re-verified in desktop Chromium only.
Noticed, not changed
stop()pauses the element, but a pause fires neitherendednorerror, so_speakServer's promise never settles after a stop and its blob URL is only released when the page unloads./api/tts/speakfetch (network error) still only logs.