Skip to content

fix(web): server voice plays again (media-src blob:) - #52

Open
NetDevAutomate wants to merge 5 commits into
mainfrom
fix/voice-csp-media-src
Open

NetDevAutomate wants to merge 5 commits into
mainfrom
fix/voice-csp-media-src

Conversation

@NetDevAutomate

Copy link
Copy Markdown
Owner

Web voice has been silent since 2 Sep. The badge said Kokoro (server), OpenVox answered every request, and nothing was heard.

Root cause

The server voice tier plays the host's WAV through new Audio(URL.createObjectURL(blob)) (since adf7493, 23 Aug). On 2 Sep, 1f10f05 set the page's policy to default-src 'self' with no media-src. 'self' does not match a blob: URL, so the browser refused every utterance:

Loading media from 'blob:http://127.0.0.1:.../...' violates the following Content
Security Policy directive: "default-src 'self'". Note that 'media-src' was not
explicitly set, so 'default-src' is used as a fallback.

play() rejected with NotSupportedError, and _speakServer swallowed it. Reproduced in Chromium against the real local OpenVox by clicking the header voice toggle: /api/tts/health, /warm and /speak all returned 200 (speak with audio/wav, 66 KB), and then the audio was refused. The same page with the policy bypassed played to the end.

Changes (test-first: RED d7e0821)

  1. 7483538: media-src 'self' blob: in both modes. It applies to media only; a unit guard fails if blob: appears in any other directive. A blob: URL can only be created by this origin's own script, and script-src stays 'self' 'unsafe-eval'. I chose this over a data: URL (broader, because any markup can carry one) and over a GET audio URL (the element could no longer read a 503 and fall back to system voices).
  2. 878e826: a refused or undecodable utterance now raises the store's existing tts:engine-notice toast, which nothing in the engine had emitted since the neural tier was removed. It fires once per utterance, and not for audio that stop() or a newer utterance interrupted. Engine notices now stay up 8s instead of 2s.
  3. e659f8c: the health probe budget goes from 2.5s to 8s. OpenVox's own /v1/models takes 1.2 to 2.7s. Measured: 5 of 8 health calls went past 2.5s, and 1 page load in 5 fell back to system voices while the host was healthy. After the change, 10 of 10 loads got the server tier. A speech server that isn't running still fails at once (13ms against a closed local port).
  4. b9cc072: CHANGELOG, plus a Playback notice bullet in the voice guide.

Why nothing caught it

The only browser test of this path needs a running OpenVox, so it skips in CI, which has none. Its "played" check also counted play() calls rather than playback. With OpenVox up it did fail, but only on its console-error check. The new tests fake the host's answers at the network edge with page.route, so the page's real CSP header and the real tts-engine.js are what gets tested, and CI runs them:

  • the host's audio plays under the page's own policy;
  • audio the page can't play produces a notice that is still visible after 3s;
  • a host that answers health in 4s keeps the server tier.

The real-OpenVox browser test now asserts that the audio played, and it drives the header toggle the way a learner does.

Tested

  • tests/e2e/test_server_tts.py: 11/11, twice, with OpenVox running. On RED: 3 new tests failed plus the real-OpenVox test, each for the stated reason.
  • tests/e2e/test_web_tts_engine.py: 13/13. JS unit suite: 167/167. test_web_app.py: 32/32.
  • Docs: mkdocs build --strict and the docs tests (524 passed, 3 skipped).
  • Full unit suite: see the comment below.

Not verified: tablets. The iPad and Android devices listed under Verified devices were checked in August, before the regression. This fix was re-verified in desktop Chromium only.

Noticed, not changed

  • stop() pauses the element, but a pause fires neither ended nor error, so _speakServer's promise never settles after a stop and its blob URL is only released when the page unloads.
  • A failed /api/tts/speak fetch (network error) still only logs.

Reported: the header badge says "Kokoro (server)", nothing is ever heard,
and OpenVox appears to see nothing. Reproduced in Chromium against the real
local OpenVox by clicking the voice toggle: /api/tts/health, /warm and
/speak all answered 200 (speak with audio/wav), then the browser refused
the audio:

  Loading media from 'blob:http://127.0.0.1:.../...' violates the following
  Content Security Policy directive: "default-src 'self'". Note that
  'media-src' was not explicitly set, so 'default-src' is used as a fallback.

play() rejected with NotSupportedError and the engine swallowed it. The same
page with the policy bypassed played to the end. The server tier plays the
host's WAV through URL.createObjectURL (since adf7493, 23 Aug); the policy
(`default-src 'self'`, no media-src) arrived on 2 Sep in 1f10f05, and
'self' does not match a blob: URL.

Why nothing caught it: the only browser test of this path needs a running
OpenVox, so it skips in CI (which has none), and its "played" check counted
play() calls rather than playback.

RED, each for the stated reason:
- the CSP has no media-src allowing blob:, in both modes (2 unit tests);
- with the host's answers faked at the network edge, so CI runs it, the
  "Voice enabled" audio is refused by the page's own policy;
- a play() the page refuses produces no notice, so the learner cannot tell
  a refused audio from a dead server;
- a host that answers health in 4s is treated as absent. Measured here:
  OpenVox's /v1/models takes 1.2-2.7s, so 5 of 8 /api/tts/health calls ran
  past the page's 2.5s probe budget, and 1 page load in 5 fell back to
  system voices with a healthy host;
- the real-OpenVox browser test now asserts the audio played (it fails on
  the same policy refusal).

Guard that passes now and must keep passing: blob: appears in no directive
but media-src (2 unit tests).
The server tier plays the host's WAV through
`new Audio(URL.createObjectURL(blob))`. Since 2 Sep (1f10f05) the page's
policy has been `default-src 'self'` with no media-src, and 'self' does not
match a blob: URL, so the browser refused every utterance: /api/tts/speak
answered 200 audio/wav, play() rejected with NotSupportedError, and the
learner heard nothing while the badge said "Kokoro (server)".

media-src now allows 'self' and blob:, in both modes. The exception is
media-only (a unit guard fails if blob: appears in any other directive): a
blob: URL can only be minted by this origin's own script, and script-src
stays 'self' 'unsafe-eval', so no new script source is admitted. Chosen
over a data: URL (broader: any markup can carry one) and over streaming
the audio from a GET URL (the element could no longer read a 503 and fall
back to system voices).

Verified in Chromium: the faked-host playback test and the four CSP unit
tests pass; the whole header suite (test_web_app.py, 32) passes. The
real-OpenVox browser test resolved the system-voice tier on this run -- the
2.5s health budget, fixed separately below, not the policy.
_speakServer resolved a refused play() and an element error in silence, so
the policy regression above showed as a dead speech server: the badge
stayed "Kokoro (server)", OpenVox had answered every request, and the only
trace was one console line. The learner, reasonably, went looking at
OpenVox.

A refused or undecodable utterance now raises the settings store's existing
`tts:engine-notice` toast -- a listener nothing in the engine had emitted
since the in-browser neural tier was removed. One notice per utterance
(the error event and the rejected play() arrive together), and none for
audio that stop() or a newer utterance cut short, which rejects play() with
AbortError. NotAllowedError and NotSupportedError get a plain-language hint;
anything else is named as it is. Engine notices now stay up 8s rather than
the toast's 2s default, which is gone before the sentence is read.

Unchanged: no tier switch on a refusal (falling back to system voices is a
separate decision, and they need the same user activation), and the
fetch-failure path still only logs.

Verified: the refusal test passes (undecodable bytes, faked host); the
engine e2e file (13) and the JS unit suite (167) pass.
The page gave /api/tts/health 2.5s before treating the host as absent.
OpenVox's own /v1/models takes 1.2-2.7s on this Mac, and the health route
waits on it. Measured on 29 Sep against the real OpenVox: 5 of 8 health
calls ran past 2.5s, and 1 page load in 5 resolved to system voices with a
healthy host -- a badge that changes between reloads for no visible reason.
The real-OpenVox browser test landed on system voices in one run for the
same reason.

The budget is now 8s. It still ends a hung host's probe (the point of the
bound: an init that never resolves looks broken), and leaves room for a
tablet's LAN hop. A speech server that is not running on this machine
refuses the connection at once (health against a closed local port: 13ms),
so the budget changes nothing in the common "no server" case.

Verified: after the change, 10 of 10 page loads resolved the server tier
against the same OpenVox (health 2.4-2.6s on that run, 6 of 8 over the old
budget); the server-TTS e2e file passed 11/11 twice with OpenVox running,
and the engine e2e file 13/13.
CHANGELOG [Unreleased] Fixed: the media-src policy regression (2 Sep to
now), the silent playback refusal, and the 2.5s health budget. The voice
guide gains a Playback notice bullet under Controls: the badge names the
tier, not whether this device could play a given utterance.

Verified: mkdocs build --strict (docs extra) and the docs tests (524
passed, 3 skipped).
Copilot AI balanced review requested due to automatic review settings September 29, 2026 09:01

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation addresses the regression with narrowly scoped policy changes and comprehensive deterministic coverage.

Review effort: Balanced
Findings: None

What changed in this PR

Restores server-side web voice playback and improves failure visibility and health-probe reliability.

Changes:

  • Allows blob: audio through the CSP’s media-src.
  • Surfaces playback failures via longer-lived toast notifications.
  • Adds deterministic browser coverage and extends the server health timeout.
File Description
CHANGELOG.md Documents voice playback fixes.
docs/​voice-output.md Documents playback failure notices.
packages/​studyloop/​src/​studyloop/​web/​app.py Permits blob media in CSP.
packages/​studyloop/​src/​studyloop/​web/​static/​components.js Extends TTS notice visibility.
packages/​studyloop/​src/​studyloop/​web/​static/​tts-engine.js Reports playback failures and extends probe timeout.
packages/​studyloop/​tests/​e2e/​test_server_tts.py Adds browser-level playback regressions.
packages/​studyloop/​tests/​test_web_app.py Verifies CSP scope in both modes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@NetDevAutomate

Copy link
Copy Markdown
Owner Author

Full unit suite (default markers) on e659f8c + docs, this Mac: 7401 passed, 31 failed, 14 errors, 16 skipped, in 15m54s.

The 45 failing ids against the recorded machine-specific set (docs/architecture/plan-integration/receipts/full-suite-control-item4-2026-09-18.md): 44 are in it. The 45th is agent-session-tools/tests/test_sync_conversation_integrity.py::test_concatenated_remote_dump_with_existing_archive, the known sqlite3 subprocess timeout, in a package this PR does not touch. No failing id comes from a file this PR changes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants