From 8751f833f33a0a71c79e5146fc1a91f2d59279e4 Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 20:30:01 -0500 Subject: [PATCH 1/8] Restore plan working document for Phase B (drops before PR) --- PLAN-mc-prompter.md | 300 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 300 insertions(+) create mode 100644 PLAN-mc-prompter.md diff --git a/PLAN-mc-prompter.md b/PLAN-mc-prompter.md new file mode 100644 index 0000000..ecfecc5 --- /dev/null +++ b/PLAN-mc-prompter.md @@ -0,0 +1,300 @@ +# mc-prompter: teleprompter and AI producer, implementation plan + +Working document for the feat-prompter branch. Not intended to merge; it guides the build and gets deleted (or distilled into skill references) before the PR. Written 2026-07-09 from a recon pass over the module and external research on real-time local ASR, teleprompter prior art, and live cueing UX. Revised the same day after an adversarial three-lens review (technical feasibility, module conventions, product); the review's fixes are folded in throughout and marked where they changed a decision. + +## Decisions record (2026-07-09, BMad) + +- LLM runtime for producer mode: Ollama is the default provider of a new `[llm]` config lane, with the standard provider-ladder pattern (local-first, other rungs opt-in later). +- Scope: all three tiers built on this one branch, in phases, with the classic teleprompter working end to end first. +- Skill shape: one service skill, `mc-prompter`, following the mc-audio pattern (no stage, no gate, no project.json state). The `record` stage stays creator-owned. +- Cue channels: visual-only in v1. The kokoro spoken tier is designed here but ships as a fast-follow behind a config flag. +- Landing (approved 2026-07-09): stacked PRs. PR 1 = Phase A plus minimal docs and the help row, PR 2 = Phase B, PR 3 = Phases C+D. All developed on this branch. + +## Product overview: three tiers + +Tier 1 is a classic teleprompter with the full standard feature set, no AI, no model downloads. Tier 2 is voice-follow: the scroll tracks the speaker through a known script using streaming local ASR plus a deterministic alignment algorithm, no LLM. Tier 3 is producer mode: a rundown file (duration, ordered talking points, intro, wrap), a rolling transcript, and a small local LLM that keeps the speaker on track with rate-limited visual cues. + +Each tier is progressive enhancement over the one below it. Tier 1 works on any machine with only the module installed. Tier 2 requires the prompter-lab workspace (ASR models, consent-gated download). Tier 3 additionally requires a running Ollama. + +Everything chosen is cross-platform (Windows, macOS, Linux): browser UI, browser mic capture, sherpa-onnx, silero VAD, kokoro-onnx, Ollama. This is the module's first fully cross-platform lane, and the sherpa-onnx dependency incidentally opens a path to the cross-platform transcription lane already on TODO.md. + +## Architecture + +### Skill shape + +New folder `skills/mc-prompter/`: + +``` +skills/mc-prompter/ + SKILL.md service skill: what it does, how to launch, tier gating + customize.toml [workflow] block + prompter defaults + references/ + cueing.md the cue design contract: escalation ladder, budget, vocabulary, + replan rules, coverage semantics, headphones/AEC dependency + rundown-spec.md the rundown.md file format specification (time math included) + scripts/ + ensure_workspace.py prompter-lab builder (mirrors mc-audio's, consent-gated) + run_prompter.py stdlib launcher: validates workspace, port probe, launches server + server/ the application package (both launch paths run python -m server.main) + __init__.py + main.py aiohttp app: HTTP + WebSocket + static UI; lazy tier-2/3 imports + asr.py sherpa-onnx streaming recognizer + silero VAD (workspace only) + align.py script-follow alignment engine (pure stdlib) + producer.py rundown state machine, replanner, cue engine, Ollama client + rundown.py rundown.md parser (pure stdlib) + script_ingest.py script.md / markdown / plain text ingestion (pure stdlib) + static/ vanilla HTML/JS/CSS, no build step + tests/ self-running test-*.py files (convention below) + fixtures/ +``` + +Conventions honored: SKILL.md frontmatter with exactly `name` and `description`; scripts invoked only via `uv run` with PEP 723 headers; scripts take explicit resolved arguments and do no config discovery; the skill reads only its own folder, `_bmad/scripts/`, and project files; nothing user-specific ships in the module; config keys kebab-case. + +### The prompter-lab workspace + +Heavy dependencies live in a persistent venv at `{engines-path}/prompter-lab`, exactly like audio-lab: + +``` +/.venv/ aiohttp (same pin as the PEP 723 header), numpy, + sherpa-onnx (pinned), soundfile +/models/ ASR models (below); later the kokoro pair +/out/ session artifacts (take logs, session transcripts, script backups) +``` + +`ensure_workspace.py --check` exits 0/4 like mc-audio's; the build asks consent before any download. Tier 1 does not require the workspace at all: `run_prompter.py` launches the server with ASR disabled when the workspace is absent, using `uv run` with aiohttp as a PEP 723 dependency. When the workspace exists, the launcher runs the venv interpreter instead, resolved portably (`.venv/bin/python` on POSIX, `.venv\Scripts\python.exe` on Windows). + +Two-launch-path discipline (review finding): + +- Both paths execute the server identically as `python -m server.main` from the scripts directory, so intra-package imports resolve the same way in both. A test executes main.py under a bare env with only aiohttp and asserts the tier-1 routes come up. +- main.py imports asr.py (and anything touching numpy/soundfile) lazily, only inside the workspace-present branch. align.py, rundown.py, and script_ingest.py stay pure stdlib. +- The aiohttp version is pinned identically in the PEP 723 header and the workspace venv, bumped together. + +Port handling (review finding): default port 8770 (mc-ograf's ephemeral verifier uses 8771). On launch, probe the port; if occupied, query `/health` (which reports session and script identity), then offer kill-and-replace or auto-increment to the next free port. The chosen port is printed and written to a session file the skill reads back. `--port` overrides. The server binds 127.0.0.1 by default; `--lan` opts into 0.0.0.0 for the phone remote and tablet displays, and the docs note the Windows Firewall consent dialog this triggers. The home page shows the LAN URL only when it is actually reachable. + +ASR models downloaded into `models/`: + +- Primary: sherpa-onnx export of nvidia nemotron-speech-streaming-en-0.6b, int8 (the streaming sibling of the parakeet family; NVIDIA Open Model License, commercial use permitted). Chunk setting 560 ms as the default latency point. +- Fallback for low-end hardware: a small streaming zipformer English export (Apache-2.0), selectable via config. +- VAD: silero VAD via sherpa-onnx's built-in VoiceActivityDetector (one dependency covers both). + +### Server and UI + +One aiohttp process serving four pages plus a WebSocket: + +- `/` home: pick a source (project script.md, rundown.md, pasted/loaded text), configure, preflight, launch. Shows the reconciled rundown plan before a show starts. +- `/prompt` the prompter display: fullscreen scroll surface, all tier-1 features, the producer rail when tier 3 is active. +- `/remote` phone-as-remote over LAN: play/pause, speed, jump to marker, next/prev section, and in producer mode the point-list override controls. Control only, no mic. The remote URL embeds a per-session token and is presented as a QR code on the home page, so a random LAN device cannot drive the prompter mid-show. +- `/overlay` OBS browser source: transparent background, renders only the ambient rail and cue cards for live shows. + +Concurrency model (review finding, binding): sherpa-onnx decode is CPU-bound and blocking, so it never runs on the event loop. A dedicated ASR thread consumes a bounded queue of PCM frames and publishes recognition events back via `call_soon_threadsafe`. Ollama ticks run as background tasks with hard timeouts and never hold shared state across an await. WebSocket fan-out never awaits decode or LLM work. The server smoke test asserts `/remote` command round-trip latency stays low while the replay harness saturates the ASR path. + +Audio capture and ownership (review findings, binding): + +- Mic capture happens on the server machine. getUserMedia requires a secure context, which http over LAN is not, so a tablet pointed at `/prompt` is display-only. The WS protocol separates the display role from the audio-producer role: the server grants a capture token to exactly one localhost connection; frames without the current token are rejected and the UI names the mic owner. Mic-on-remote-device would require shipping TLS and is explicitly deferred. +- getUserMedia constraints request `echoCancellation: false, noiseSuppression: false, autoGainControl: false, channelCount: 1`; browser speech processing measurably degrades ASR input. The preflight screen reads back `track.getSettings()` and surfaces what was actually applied, since browsers may ignore constraints. When the kokoro spoken tier ships, headphones-only output is what keeps AEC unnecessary; cueing.md records that dependency. +- The AudioContext is created with `{ sampleRate: 16000 }` so the browser resamples; if the browser refuses the rate, a small resampler in the worklet handles the conversion (naive decimation from 44.1 kHz aliases into the speech band). The preflight screen verifies `context.sampleRate`. The worklet ships ~120 ms PCM16 mono frames over the WebSocket. +- Backpressure has a policy at both ends: the browser checks `ws.bufferedAmount` and drops frames past a threshold; the server frame queue is bounded and drops oldest on overflow while raising an "ASR behind real-time" state on the rail. The replay harness asserts queue-depth behavior. + +State flows back to all connected pages over the same WebSocket (scroll position, VAD state, transcript tail, cue events), so the remote and overlay stay in sync with the prompter display. + +UI is vanilla HTML/JS/CSS with no build step. Display settings (mirror flips, font, colors, margins, eyeline position) persist in localStorage per device, with the `[prompter]` config values as defaults; a beam-splitter rig and the operator's browser keep independent settings without reconfiguration each launch. + +### Tier 1: the standard feature checklist + +Scroll and timing: + +- Smooth continuous scroll, speed as WPM with live +/- adjustment (keyboard, wheel, remote) +- Timed mode: give total duration, speed is continuously re-derived (remaining words over remaining time, recomputed on resume and after any jump), with the timer display showing drift from plan +- Pause/resume (spacebar), jump forward/back, jump to marker, restart +- Countdown before scroll starts +- Elapsed and remaining time, estimated read time from word count at the creator's measured wpm (`[owner] wpm` from the studio config when available) + +Display: + +- Mirror flip horizontal, vertical, and both (beam-splitter rigs) +- Font family/size, text and background colors, margins, line height +- Adjustable eyeline/cue marker (position, style) +- Fullscreen; works on a second monitor or a tablet pointed at the same URL (display-only on remote devices, see audio ownership above) +- Per-device settings persistence (localStorage), config defaults underneath + +Script handling: + +- Markdown and plain text; project `script.md` ingestion (below) +- Inline bracket notes render dimmed and are never matched by voice-follow +- Named markers/sections for jumping +- Edit-in-place from the home page between takes; edits write back to the source file with a timestamped backup copied to the workspace `out/` first, so the prompted text and the pipeline artifact never silently diverge + +Remote: + +- Keyboard shortcuts throughout; bluetooth presenters and USB foot pedals work as keyboard emulators for free +- `/remote` phone page over LAN, session-token URL via QR code + +### Script ingestion (pipeline tie-in) + +`script.md` from a project is directly consumable: plain spoken prose. Ingestion handles the two inline marker types: + +- `[INVENTED]` flags render as a subtle badge, toggleable off +- `[TAKE s-s]` lines render dimmed with a "have it already" badge, since those lines were already spoken well in the interview footage and may not need re-recording; a toggle hides them entirely + +The prompter takes a path argument; the skill resolves it from the project when launched inside the pipeline flow ("record with the teleprompter") or accepts any file standalone. + +### Tier 2: voice-follow alignment engine + +Deterministic, no LLM. The prior art (bounded-window Levenshtein prefix matching, the PromptSmart hold-and-re-anchor behavior) consumed Web-Speech-style utterance partials; a streaming transducer behaves differently, so the contract is adapted for sherpa-onnx output (review finding, binding): + +- Input contract: the engine consumes token deltas since the last partial, not whole hypotheses. The last K tokens (K around 3 to 5) of the hypothesis are held provisional because beam search can revise the tail between partials; the anchor commit lags the hypothesis head by K tokens and absorbs revisions. On endpoint detection (`is_final`), the anchor hard-commits and tail-tracking state resets for the next segment. BPE pieces merge to words during normalization before matching. +- Normalize both script tokens and ASR tokens: lowercase, strip punctuation, expand common number/abbreviation forms at index-build time (a small normalization table; "2026" also indexes as "twenty twenty six") +- Maintain a monotonic anchor (last committed script token). On each delta batch, take a lookahead window from the anchor (window size proportional to utterance length plus a constant, on the order of 2x + 10 tokens) and find the window prefix minimizing Levenshtein distance to the pending recognized words; the best prefix end becomes the provisional anchor +- Silence (VAD) produces no partials, so the scroll holds; ad-libs fail to match and the anchor holds until speech re-matches within the window +- Escape hatches: click/tap any word to re-anchor, arrow keys nudge the anchor, and a paragraph-skip gesture jumps the window when the creator deliberately skips content +- The scroll controller eases toward the anchor position rather than jumping; the anchor leads the eyeline by the measured end-to-end latency times the current speaking rate, so the eyeline sits where the speaker actually is, not where ASR last confirmed +- Match state is visible: matched text subtly tinted behind the eyeline, so trust in the tracker is inspectable + +Latency: the honest budget is end-to-end and includes terms the naive sum misses: capture framing (~120 ms) + ASR chunk emission (560 ms configured) + decode compute (hardware-dependent, grows under OBS load) + alignment (<10 ms) + scroll easing (a deliberate time constant). Realistic eyeline-follows-voice latency is 1 to 2 s depending on hardware. The replay harness measures capture-timestamp-to-anchor-update wall time on each target platform, and the measured number feeds the eyeline lead default. No fixed latency claim ships in docs. + +Preflight (review finding, this is the try-once-never-again defense): before any take, the preflight screen enumerates input devices with a picker persisted per machine, shows a live level meter, reports the applied audio constraints and sample rate, and runs a 10-second "read this sentence" tracking test that demonstrates the match-tint following before a real take starts. If recognition confidence is garbage, it fails loudly and names the device in use. + +The engine is a pure-stdlib module. Its fixture suite is built from recorded real partial sequences (the actual partial/final event stream captured from the model over the replay WAV), not hand-written final transcripts, covering: verbatim read, ad-lib excursion and return, skipped paragraph, number/abbreviation mismatch, repeated-phrase script traps, and tail-revision events. + +### Tier 3: producer mode + +Inputs: a rundown file, the rolling transcript from the same ASR stream, and the show clock. + +Show clock semantics (review finding, binding): the plan's arithmetic never keys off server or page start. An explicit GO LIVE control (on `/prompt` and `/remote`) starts the show clock after any pre-roll, and a plan-hold control freezes elapsed time and the state machine during BRB or technical trouble while VAD and the transcript keep running so context is not lost. Both are fixture-tested scenarios. + +The producer is two cooperating parts: + +- A deterministic state machine (code, not LLM). It tracks elapsed time against per-segment budgets, and it re-plans rather than merely flagging lateness: on every tick, remaining show time is redistributed across uncovered segments proportionally to their original budgets, with wrap-minutes protected as a hard reserve. Green/yellow/red state is always computed against the current re-plan, never the original rundown. When redistribution would push any segment below a feasibility floor, the state machine emits a card-tier CUT suggestion ("DROP: point 4, or 90s each"). This replan behavior is the core producer value (a countdown that only turns red is a nag, not a producer) and is specified in cueing.md as part of the binding contract. +- An LLM tick (Ollama): scheduled adaptively, and at VAD pause events, it receives a compact state block (rundown with per-point coverage, the current re-plan, elapsed vs plan, the last ~60 seconds of transcript) and returns structured JSON: proposed coverage transitions, current-topic guess, an optional suggested cue with tier and text, and a one-line reason. Cheap keyword/fuzzy matching runs continuously between ticks as a first-pass coverage signal the LLM confirms or overrides. + +Coverage semantics (review finding, binding): coverage is sticky and monotonic in the state machine; the LLM may only propose uncovered-to-covered transitions, never reversions, so the rail cannot flicker. "Next" is defined as the first uncovered point in rundown order, which stays well-defined when the creator covers points out of order. The human is the final authority: `/remote` gains producer controls, a tappable point list with mark-covered, skip, and make-current, so one tap mid-show rescues any model misjudgment. + +LLM tick budget (review finding, binding): the tick must never starve the ASR thread. Concretely: + +- Requests use Ollama structured outputs (`format` with a JSON schema), `think: false`, temperature 0, a hard `num_predict` cap, and `keep_alive` so the model stays resident +- The prompt keeps a stable prefix (system + rundown first, rolling transcript last) so Ollama prefix caching skips reprocessing +- Cadence is adaptive: the next tick is scheduled at `max(15 s, 3x last tick wall time)`, and ticks are skipped entirely while the ASR queue depth signals CPU pressure +- Default model is `qwen3:4b` where Ollama reports GPU/Metal offload; on CPU-only machines the producer startup check recommends and falls back to a sub-2B tag (`qwen3:1.7b`). Every request carries a hard timeout; a timed-out tick is dropped, not queued + +The cue engine (code) is the final authority on delivery: it applies the density setting, the one-active-cue rule, per-interval budget, and tier gating. The LLM proposes; the state machine disposes. A wrong suggestion costs nothing because rate limiting, tiering, and coverage stickiness are deterministic. + +Visual cue surface (v1, from the cueing research): + +- Ambient tier, always on: a rail showing current point, next point, and a green/yellow/red segment-time state computed against the re-plan (Toastmasters vocabulary), plus overall show progress. No motion, no reading required +- Card tier: a single quiet card ("NEXT: pricing demo", "STRETCH: 4 min left, 1 point to go", "DROP: point 4, or 90s each"), released at pauses, auto-expiring, never stacked +- Attention tier: the card flashes/enlarges for time-critical states ("WRAP", "2:00 OVER"), the one tier allowed to appear mid-sentence +- Vocabulary is the broadcast lexicon (standby, wrap, stretch, hard wrap, time remaining); `references/cueing.md` is the binding contract + +Free-talk support: a rundown with no script body per point is exactly the "5 ideas in this order plus intro and wrap" show. The intro and wrap can carry full scripted text (prompted via tier 1/2) while the middle segments run producer-only. The segment handoff is explicit UI behavior, not hand-waved: when a scripted segment ends (anchor reaches section end, or manual next-section), `/prompt` switches to a large-type rail view showing the current bullet set; entering the next scripted segment switches back to the scroll surface. + +Spoken tier (designed now, shipped later behind `spoken-cues = false`): short formulaic kokoro phrases only ("thirty seconds", "wrap"), synthesized by a persistent kokoro instance (~150 to 300 ms for a short cue on CPU), released only at pauses, hard requirement that output routes to headphones (this is also what keeps browser AEC unnecessary). Never speech-over-speech except a true emergency tier. The mix-minus principle from IFB practice: the speaker must never hear their own voice back. + +### The rundown artifact + +New file format, specified in `references/rundown-spec.md`. The starter template ships through mc-setup's `assets/` into the studio like tokens and format profiles, so Manny and any skill can read it as a project file without crossing skill-folder boundaries; mc-prompter owns the spec and does any template-based drafting, and Manny routes to it (review finding). + +```markdown +--- +show: "Why local models win" +duration-minutes: 30 +cue-density: normal # hands-off | minimal | normal | chatty +wrap-minutes: 3 +--- + +## Intro (3 min) + +Full scripted intro text here, prompted normally. + +## Point 1: The cost argument (5 min) + +- cloud bills compound, local is capex +- the 4090 anecdote + +## Point 2: Latency (5 min) +... + +## Wrap (3 min) + +Scripted wrap text. +``` + +Time math (review finding, in the spec): per-segment minutes are optional; unbudgeted segments split the remaining time evenly. If explicit minutes exceed duration-minutes, duration-minutes wins and the parser warns at load; the home page shows the reconciled plan before the show starts. The parser accepts exactly `(N min)` and `(Nm)` heading suffixes and rejects anything else with a line-numbered error, because hand-written and Manny-drafted rundowns will produce creative variants on day one. Frontmatter `cue-density` overrides the config value (most specific wins). + +Segments with prose bodies prompt as script; segments with only bullets run producer-only. For pipeline projects the file lives at `{projects-path}//rundown.md`; standalone shows pass any path. This artifact also fills the episode-plan gap the livestream formats already reference (the 1.0.x per-episode stream-pack fast-follow needs the same file), and the planned mc-research skill becomes its natural upstream. + +### Config: new studio sub-tables + +Two new tables in `[modules.manticore]`, seeded from mc-setup's `[defaults]`: + +```toml +[defaults.prompter] +workspace = "prompter-lab" # resolved {engines-path}/{prompter.workspace} +asr-provider = "nemotron-streaming" # nemotron-streaming (default) | zipformer-small | none +cue-density = "normal" # hands-off | minimal | normal | chatty +spoken-cues = false # kokoro tier, fast-follow +port = 8770 + +[defaults.llm] +provider = "ollama" # the only implemented rung; others planned, opt-in +model = "qwen3:4b" # any ollama tag; producer falls back to a sub-2B tag on CPU-only +endpoint = "http://localhost:11434" +api-key-env = "" # stays empty for local lanes, pattern-consistent +``` + +mc-setup changes: + +- Add both tables to `[defaults]`, a short optional interview step (offer the teleprompter, ask about producer mode and Ollama only if wanted), and both table names to the step 8 write list +- Migration (review finding, important): do NOT add these tables to the step 1a 0.x classifier list; that list defines what makes a studio 0.x, and adding 1.1 tables to it would mislabel every current 1.0.x studio as 0.x and run the full migration flow on it. Instead add a separate backfill rule alongside it: a config that has the 1.0 tables but is missing `[prompter]` or `[llm]` is a current studio predating the teleprompter; backfill both tables surgically from `[defaults]` and offer the optional prompter interview +- check_deps.py gains an optional `ollama` row (producer mode only) with install pointers +- PIPELINE.md's conventions line mentions the new sub-tables + +The `[llm]` lane follows the enforcement pattern: the producer accepts only `provider = "ollama"` and exits 3 for anything else, so no planned lane ever pretends to work. + +## Module integration checklist + +- `skills/module-help.csv`: one new row with the full 13-column schema (module, skill, display-name, menu-code, description, action, args, phase, preceded-by, followed-by, required, output-location, outputs); phase `anytime`, preceded-by and followed-by left empty per the mc-audio/mc-ograf service-skill precedent +- `.claude-plugin/marketplace.json`: add `./skills/mc-prompter` to the skills array (version bump at release, not in this PR) +- `skills/mc-pipeline/PIPELINE.md`: record stage row gains a sentence noting mc-prompter as the optional tool for the creator-owned record stage; no stage table change, no gate change +- mc-script SKILL.md step 6: the "now record" handoff names the teleprompter option (naming only, no cross-skill read) +- mc-agent (Manny) capabilities: route "teleprompter", "prompt me", "run my show", "producer mode" to mc-prompter; for rundown drafting Manny routes to mc-prompter or works from the studio-installed template, never reads mc-prompter's folder +- docs/user-guide.md: new section; README skill table row; TODO.md: add the spoken-cue fast-follow and the sherpa-onnx transcription-lane opportunity + +## Phasing + +- Phase A, classic teleprompter: workspace-less launch path, aiohttp server with the concurrency skeleton, `/prompt` with the full tier-1 checklist, `/remote` with session token, script.md ingestion with marker handling, home page, settings persistence, port handling. Verifiable end to end with zero models +- Phase B, voice-follow: ensure_workspace.py, browser audio path with the constraint/ownership/backpressure rules, sherpa-onnx + VAD integration on the ASR thread, align.py with the transducer contract and its recorded-partials fixture suite, preflight screen with device picker and tracking test, tracking UX (hold, re-anchor, click-to-anchor, match tinting, eyeline lead) +- Phase C, producer: rundown parser + spec, producer state machine with replanner and cue engine plus time-warped fixtures (running long, running short, out-of-order coverage, GO LIVE / hold), Ollama tick with the budget rules, ambient rail + cards on `/prompt` and `/overlay`, `/remote` producer controls, `[prompter]`/`[llm]` config plumbing, mc-setup + check_deps integration +- Phase D, finish: docs, help catalog row, marketplace entry, Manny routing, take-log hook, cross-platform smoke instructions, changelog + +Landing (approved 2026-07-09): develop everything on this branch, land as stacked PRs: PR 1 = Phase A plus minimal docs and the help row (a complete, useful teleprompter on its own), PR 2 = Phase B, PR 3 = Phases C+D. Each is independently green and valuable, and review feedback on the foundations arrives before the producer is built on top of them. + +## Testing and verification + +Test convention (review finding, matches the quality gate CI): the CI discovers `skills/*/scripts/tests/test-*.py` (hyphenated) and runs each file directly via `uv run`; there is no pytest. Every test file carries its own PEP 723 header, uses stdlib unittest with `unittest.main()`, and runs with no models, no network, no downloads. + +- `test-align.py`: the transducer-contract fixtures (recorded partial sequences committed as small JSON fixtures; the five failure-mode cases plus tail-revision events) +- `test-producer.py`: replanner and cue engine under time-warped scenarios (running long with replan and CUT suggestion, running short with STRETCH, point skipped, out-of-order coverage, GO LIVE and hold semantics, budget/ladder enforcement, coverage stickiness) +- `test-rundown.py`: parser, time-math reconciliation, heading-suffix rejection with line numbers +- `test-script_ingest.py`: marker handling, bracket-note exclusion +- `test-server.py`: aiohttp smoke via its own PEP 723 aiohttp dependency; routes up under the bare tier-1 env; WS state fan-out; remote-latency-under-ASR-load assertion lives here but auto-skips (exit 0 with a message) when the workspace is absent, so CI never needs models +- One small recorded fixture WAV (a few seconds, 16 kHz mono) committed under `scripts/tests/fixtures/`; nothing generates audio in-test +- The ASR replay harness (feed the fixture WAV through the same code path the WebSocket uses, measure capture-to-anchor wall time) is a documented manual check per platform, not a CI assertion; CI has no workspace and wall-clock assertions are flaky by construction +- Manual test matrix documented in the skill: macOS (reference), Windows, Linux; beam-splitter mirror check; phone remote; OBS overlay; device-picker preflight + +## Risks and mitigations + +- Nemotron streaming model quality/latency on low-end CPUs: zipformer-small fallback behind `asr-provider`, and tier 1 works with `none` +- ASR token timestamps are start-only in sherpa-onnx: alignment keys on token text order, not timestamps, so this costs nothing +- CPU contention between ASR, the LLM tick, and OBS on one machine: the tick budget rules above (adaptive cadence, pressure-skip, small-model fallback, timeouts) plus the bounded-queue drop policy keep voice-follow latency from compounding; the rail surfaces "ASR behind real-time" instead of silently lagging +- Ollama absent or model not pulled: producer mode degrades to the deterministic rail (timing and replan cues still work, coverage judgments off); the UI says exactly what is missing +- Repeated phrases in scripts confusing alignment: bounded window plus monotonic anchor limits damage; fixture-tested +- Browser mic pitfalls (wrong device, processing constraints ignored, wrong sample rate): the preflight screen with device picker, level meter, applied-settings readback, and the 10-second tracking test catches all of these before the first real take +- Scope creep: the kokoro spoken tier and the mc-cut take-log consumption are explicitly out of v1 + +## Explicitly out of scope (future hooks) + +- Spoken kokoro cue tier (designed, config key ships false; fast-follow) +- mc-cut consuming the take log (`out/take-log.json`, script positions and timestamps per take) to pre-anchor cut plans +- Chat/vision inputs to the producer (reading live chat is a natural producer input later) +- Cloud LLM rungs for `[llm]`, paid ASR rungs +- TLS for mic capture on remote devices (tablet-as-mic) +- Cross-platform batch transcription lane via sherpa-onnx parakeet-tdt offline export (separate TODO item this branch makes cheaper) From d777f4155e44f6b0e75fe1c42cf65370221dd669 Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 21:32:46 -0500 Subject: [PATCH 2/8] Add voice-follow alignment engine with real captured ASR fixtures --- skills/mc-prompter/scripts/server/align.py | 517 +++++++++++++ .../scripts/server/script_ingest.py | 41 + .../scripts/tests/fixtures/align/adlib.json | 703 ++++++++++++++++++ .../scripts/tests/fixtures/align/digits.json | 343 +++++++++ .../scripts/tests/fixtures/align/skip.json | 288 +++++++ .../tests/fixtures/align/stopword.json | 302 ++++++++ .../scripts/tests/fixtures/align/tailrev.json | 304 ++++++++ .../tests/fixtures/align/takeskip.json | 376 ++++++++++ .../tests/fixtures/align/verbatim.json | 570 ++++++++++++++ .../mc-prompter/scripts/tests/test-align.py | 607 +++++++++++++++ .../scripts/tests/test-script_ingest.py | 109 ++- 11 files changed, 4159 insertions(+), 1 deletion(-) create mode 100644 skills/mc-prompter/scripts/server/align.py create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/adlib.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/digits.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/skip.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/stopword.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/tailrev.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/takeskip.json create mode 100644 skills/mc-prompter/scripts/tests/fixtures/align/verbatim.json create mode 100644 skills/mc-prompter/scripts/tests/test-align.py diff --git a/skills/mc-prompter/scripts/server/align.py b/skills/mc-prompter/scripts/server/align.py new file mode 100644 index 0000000..2db1293 --- /dev/null +++ b/skills/mc-prompter/scripts/server/align.py @@ -0,0 +1,517 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# /// +"""Voice-follow alignment engine for mc-prompter (Phase B). Pure stdlib. + +Consumes the streaming-transducer partial hypothesis stream (BPE token lists +per segment, as emitted by server/asr.py) and maintains a monotonic committed +anchor into the script's speakable words (the same global word indexing the +UI renders as data-i spans; see script_ingest.speakable_words()). + + aligner = Aligner(speakable_words(doc), + take_ranges=take_word_ranges(doc)) + result = aligner.feed(tokens, segment, final) + # result == {"anchor": int, "moved": bool, "held": bool} + aligner.set_anchor(word_index) # human override, backwards allowed + aligner.anchor # committed global word index, -1 initially + +take_ranges is the list of [start, end) global word-index pairs for the +script's take blocks (script_ingest.take_word_ranges, same indexing as +speakable_words). Take paragraphs are already-recorded footage: they are +dimmed or hidden in the UI and a presenter normally skips them aloud, so the +aligner treats their words as free to skip (see Matching below). + +Input contract (matches the observed nemotron-streaming behavior recorded in +the Phase B contract and the committed fixtures): + + feed() receives the FULL hypothesis token list for the current segment on + every partial. Hypotheses grow incrementally but the tail revises at the + word level (BPE pieces append into the previous word, e.g. "workshop" + becoming "workshopod"), so the last k_provisional merged words are treated + as provisional and never advance the committed anchor. On final=True the + whole hypothesis commits. A new segment id resets hypothesis tracking + (endpoint reset cleared the recognizer state) but never rewinds the + committed anchor. Timestamps are not used. + +Matching: + + BPE pieces merge to words (a piece starting with a space, or the first + piece, starts a new word). Both script words and hypothesis words are + normalized: casefold, punctuation stripped, hyphen/slash split, and a + small documented number/ordinal table expands digit forms to spoken words + on BOTH sides, so script "2026" matches spoken "twenty twenty six" and a + recognizer that emits digits still matches a written-out script. + + Newly committed hypothesis words (pending) are matched against a script + window starting just past the committed anchor, window = window_mult * + len(pending) + window_base expanded NON-TAKE tokens; take-block tokens + inside that span ride along for free, so a take paragraph of any length + never consumes window budget and the words after it stay reachable. A + word-level Levenshtein DP picks the window prefix end with minimum + distance. Per-word substitution cost: exact 0, prefix or difflib-fuzzy + 0.5, else 1. Insertion of an unmatched spoken word costs 1; deletion of + a script window word costs 0.25, except take-block words which delete + at cost 0 (the presenter is expected to skip them; if they DO read the + take aloud, substitution matching inside the take still works and the + anchor tracks through it). Deletions must be cheap: at cost 1, aligning + the pending words after N skipped script words (cost N) always loses to + matching nothing (cost len(pending)), so skips and recognizer word + drops would stall the anchor forever. At 0.25 a deliberate skip is + followed within the window while spurious far matches stay unprofitable + (a fuzzy match only pays for itself within about two words of + deletion). The anchor advances to the window position of the last + pending word that aligned at cost <= 0.5. The committed anchor word is + never a take-block word: if the last match lands inside a take (the + take was read aloud, or a word fuzzy-matched into it), the broadcast + anchor advances to the next non-take word so the UI always has a + visible span to scroll to (falling back to the previous non-take word + when the take runs to the end of the script). + + Hold rule (the gate): off-script speech must not move the anchor. A + match only counts toward the gate if the spoken token is informative: + length >= 4, or not in the small stopword set below. Conversational + ad-libs are stopword-dense and scripts are full of forward stopword + duplicates, so stopword coincidences ("and", "the", "we") must not be + evidence of progress. When the pending batch contains informative + tokens, moving requires (a) at least one informative token matched + EXACTLY (cost 0), and (b) informative-matches / (informative-pending + + interior-deletions) >= hold_threshold, where interior-deletions counts + unmatched non-take window words strictly between the first and last + matched position. Charging interior deletions makes scattered + coincidence matches fail (they straddle unmatched script) while a + genuine skip still passes (its matches are contiguous, the skipped + words sit BEFORE the first match). A batch of only stopwords (common in + small partial commits: "and the") falls back to the same ratio over all + pending words, so verbatim reading still advances promptly. When the + gate fails, the anchor does not move and held=True is returned. Held + pending words are consumed, so the window never fills with off-script + speech and the next on-script words re-match from the committed anchor. + + End of script: when the window is truncated by the document end and the + gate fails, the batch may be the script tail plus overflow speech + committed together in one final ("...thanks for watching" continuing + past the last script word). Overflow past the end cannot match + anything, so the gate is retried over only the pending words up to the + last matched one; if that passes, the tail anchors instead of holding + forever. + + feed() clamps each batch to the last 24 RAW words before normalization. + Pending is normally 1-5 words, but a mid-segment Aligner rebuild + (script edit while a segment is in flight) resets _consumed and the + next partial commits the whole hypothesis so far as one batch; _match + is O(len(pending) * window) with window proportional to pending, i.e. + quadratic, and a 300-word batch would block the asyncio event loop for + seconds. Older words are stale context anyway: the window matcher only + needs recent speech. The clamp is on raw words (before number expansion + multiplies tokens) so the DP input is bounded no matter what the words + expand to. + + set_anchor() jumps anywhere, backwards included (the human is the + authority), and consumes the current hypothesis so already-spoken words + are not re-matched against the new position. + +No I/O, no threads, no imports beyond stdlib. server/main.py owns the wiring +(ASR events in, anchor broadcasts out). +""" + +import difflib + +_ONES = ["zero", "one", "two", "three", "four", "five", "six", "seven", + "eight", "nine", "ten", "eleven", "twelve", "thirteen", "fourteen", + "fifteen", "sixteen", "seventeen", "eighteen", "nineteen"] +_TENS = {2: "twenty", 3: "thirty", 4: "forty", 5: "fifty", 6: "sixty", + 7: "seventy", 8: "eighty", 9: "ninety"} +# Small ordinal table: the common single-word ordinals. Larger ordinals fall +# back to cardinal expansion of the digits ("42nd" -> "forty two"), close +# enough for the fuzzy per-word cost. +_ORDINALS = {"1st": "first", "2nd": "second", "3rd": "third", "4th": "fourth", + "5th": "fifth", "6th": "sixth", "7th": "seventh", + "8th": "eighth", "9th": "ninth", "10th": "tenth", + "11th": "eleventh", "12th": "twelfth"} + +_FUZZY_RATIO = 0.75 +_FUZZY_COST = 0.5 +_DELETE_COST = 0.25 # script window word skipped; see the module docstring +_TAKE_DELETE_COST = 0.0 # take-block words are free to skip + +# Apostrophe-family characters stripped inside words so "don't", "don’t" and +# "donʼt" all normalize to "dont". U+02BC (modifier letter apostrophe) is +# Unicode category Lm and would otherwise survive isalnum() as part of the +# token; the curly quotes would split the word instead. +_APOSTROPHES = "'’ʼ‘‛" + +# Uninformative tokens for the hold gate: a match on one of these (or any +# token shorter than 4 characters that IS in this set) is a coincidence, not +# evidence the presenter is on script. Deliberately tiny and limited to +# common function words of length <= 3 (length >= 4 already counts as +# informative regardless of this set). +_STOPWORDS = frozenset( + "a an and are as at be but by do for he i in is it me my no not of on " + "or so the to us was we you".split()) + +# feed() batch clamp, in raw (pre-normalization) words; see the module +# docstring ("feed() clamps each batch..."). +_PENDING_CAP = 24 + + +def _informative(token): + """True when a match on this token counts toward the hold gate.""" + return len(token) >= 4 or token not in _STOPWORDS + + +def merge_bpe(tokens): + """Merge BPE pieces into words. + + A piece starting with a space, or the first piece, starts a new word; + every other piece (including bare punctuation pieces like "," and ".") + appends to the current word. Empty results are dropped. + """ + words = [] + for piece in tokens: + if piece.startswith(" ") or not words: + words.append(piece.strip()) + else: + words[-1] += piece.strip() + return [w for w in words if w] + + +def _expand_number(digits): + """Expand a digit string to spoken words (the documented small table). + + 0-19 and tens from the ones/tens tables; 100-999 as "N hundred [rest]"; + 1000-9999 read as digit pairs the way years are spoken ("2026" -> + "twenty twenty six", "1995" -> "nineteen ninety five"), with the special + cases "2000" -> "two thousand", "1900" -> "nineteen hundred", "2007" -> + "two thousand seven", "1907" -> "nineteen oh seven"; 10000-999999 as + "N thousand [rest]"; anything larger digit by digit. + """ + n = int(digits) + if n < 20: + return [_ONES[n]] + if n < 100: + tens, ones = divmod(n, 10) + return [_TENS[tens]] + ([_ONES[ones]] if ones else []) + if n < 1000: + hundreds, rest = divmod(n, 100) + return [_ONES[hundreds], "hundred"] + (_expand_number(str(rest)) if rest else []) + if n < 10000: + hi, lo = divmod(n, 100) + if lo == 0 and hi % 10 == 0: + return [_ONES[hi // 10], "thousand"] + if lo == 0: + return _expand_number(str(hi)) + ["hundred"] + if hi % 10 == 0 and lo < 10: + return [_ONES[hi // 10], "thousand"] + _expand_number(str(lo)) + if lo < 10: + return _expand_number(str(hi)) + ["oh"] + _expand_number(str(lo)) + return _expand_number(str(hi)) + _expand_number(str(lo)) + if n < 1_000_000: + thousands, rest = divmod(n, 1000) + return (_expand_number(str(thousands)) + ["thousand"] + + (_expand_number(str(rest)) if rest else [])) + return [_ONES[int(d)] for d in digits] + + +def normalize_word(word): + """Normalize one raw word to a list of match tokens (possibly empty). + + Casefold; "%" becomes the word "percent"; the apostrophe family + (_APOSTROPHES: ASCII, curly, and U+02BC modifier letter) is removed so + "don't", "don’t", "donʼt" and "dont" all match; commas inside digit + runs are removed; every + other non-alphanumeric character splits the word (hyphens, slashes, + trailing punctuation). Digit pieces expand through the number table; + digit+ordinal-suffix pieces go through the ordinal table; mixed pieces + split at digit/letter boundaries. + """ + w = word.casefold().replace("%", " percent ") + for ch in _APOSTROPHES: + w = w.replace(ch, "") + # Drop commas used as thousands separators before they split the number. + w = "".join(ch for i, ch in enumerate(w) + if not (ch == "," and 0 < i < len(w) - 1 + and w[i - 1].isdigit() and w[i + 1].isdigit())) + pieces = [] + buf = "" + for ch in w: + if ch.isalnum(): + buf += ch + elif buf: + pieces.append(buf) + buf = "" + if buf: + pieces.append(buf) + + out = [] + for piece in pieces: + if piece in _ORDINALS: + out.append(_ORDINALS[piece]) + continue + if piece.isdigit(): + out.extend(_expand_number(piece)) + continue + if piece.isalpha(): + out.append(piece) + continue + # Mixed digits and letters: split at boundaries ("90s", "42nd"). + if piece[:-2].isdigit() and piece[-2:] in ("st", "nd", "rd", "th"): + out.extend(_expand_number(piece[:-2])) + continue + run = "" + for ch in piece: + if run and ch.isdigit() != run[-1].isdigit(): + out.extend(_expand_number(run) if run.isdigit() else [run]) + run = "" + run += ch + if run: + out.extend(_expand_number(run) if run.isdigit() else [run]) + return out + + +def _word_cost(a, b): + """Per-word similarity cost: exact 0, prefix/fuzzy 0.5, else 1.""" + if a == b: + return 0.0 + if len(a) >= 3 and len(b) >= 3 and (a.startswith(b) or b.startswith(a)): + return _FUZZY_COST + if difflib.SequenceMatcher(None, a, b).ratio() >= _FUZZY_RATIO: + return _FUZZY_COST + return 1.0 + + +def _match(pending, window, take_flags=None): + """Word-level Levenshtein of pending against the best window prefix. + + take_flags marks window positions that belong to take blocks; those + delete at _TAKE_DELETE_COST (0) so a skipped take costs nothing to + cross. Returns (matches, best_distance) where matches is a list of + (i, j, c) triples: pending[i] aligned to window[j] at cost c <= 0.5 on + the minimum-cost path to the best prefix end (ties break to the + shortest prefix; a substitution tie beats a free take deletion, so a + take read aloud still registers its matches). + """ + m, n = len(pending), len(window) + if take_flags is None: + take_flags = [False] * n + dele = [_TAKE_DELETE_COST if t else _DELETE_COST for t in take_flags] + dist = [[0.0] * (n + 1) for _ in range(m + 1)] + for i in range(1, m + 1): + dist[i][0] = float(i) + for j in range(1, n + 1): + dist[0][j] = dist[0][j - 1] + dele[j - 1] + cost = [[0.0] * n for _ in range(m)] + for i in range(1, m + 1): + for j in range(1, n + 1): + c = _word_cost(pending[i - 1], window[j - 1]) + cost[i - 1][j - 1] = c + dist[i][j] = min(dist[i - 1][j - 1] + c, + dist[i - 1][j] + 1.0, + dist[i][j - 1] + dele[j - 1]) + best_j = min(range(n + 1), key=lambda j: (dist[m][j], j)) + matches = [] + i, j = m, best_j + while i > 0 and j > 0: + c = cost[i - 1][j - 1] + if dist[i][j] == dist[i - 1][j - 1] + c: + if c <= _FUZZY_COST: + matches.append((i - 1, j - 1, c)) + i, j = i - 1, j - 1 + elif dist[i][j] == dist[i][j - 1] + dele[j - 1]: + j -= 1 + else: + i -= 1 + matches.reverse() + return matches, dist[m][best_j] + + +class Aligner: + """Monotonic committed-anchor aligner over the script's speakable words.""" + + def __init__(self, words, k_provisional=4, window_base=10, window_mult=2, + hold_threshold=0.5, take_ranges=None): + self.k_provisional = k_provisional + self.window_base = window_base + self.window_mult = window_mult + self.hold_threshold = hold_threshold + # Take-block words (script_ingest.take_word_ranges, [start, end) + # pairs in the same global word indexing as `words`). + self._take_words = set() + for a, b in (take_ranges or []): + self._take_words.update(range(max(0, a), min(b, len(words)))) + # Expanded script: flat normalized tokens, each carrying its source + # global word index (number expansion makes this one-to-many). + self._exp = [] + self._word_last_exp = [] + last = -1 + for i, word in enumerate(words): + for tok in normalize_word(word): + self._exp.append((tok, i)) + last = len(self._exp) - 1 + self._word_last_exp.append(last) + self._exp_take = [i in self._take_words for _, i in self._exp] + self._n_words = len(words) + self._exp_anchor = -1 + self._anchor_word = -1 + self._segment = None + self._consumed = 0 + self._last_len = 0 + + @property + def anchor(self): + """Committed global word index, -1 before any match.""" + return self._anchor_word + + def set_anchor(self, word_index): + """Jump the anchor anywhere (backwards allowed) and clear tracking. + + The current hypothesis is consumed so already-spoken words are not + re-matched against the new position; the next feed matches only + newly committed words. + """ + word_index = max(-1, min(int(word_index), self._n_words - 1)) + if word_index < 0: + self._exp_anchor = -1 + self._anchor_word = -1 + else: + self._exp_anchor = self._word_last_exp[word_index] + self._anchor_word = word_index + # Consume everything already heard in the current segment so the + # next feed matches only words spoken after the jump. + self._consumed = self._last_len + + def feed(self, tokens, segment, final): + """Consume one partial or final hypothesis for a segment. + + tokens is the full BPE token list for the segment's current + hypothesis. Returns {"anchor": int, "moved": bool, "held": bool}. + """ + words = merge_bpe(tokens) + if segment != self._segment: + self._segment = segment + self._consumed = 0 + self._last_len = len(words) + if final: + commit = len(words) + else: + commit = max(self._consumed, len(words) - self.k_provisional) + raw_pending = words[self._consumed:commit] + self._consumed = max(self._consumed, commit) + # Clamp to the last _PENDING_CAP raw words: a mid-segment Aligner + # rebuild (script edit) resets _consumed and would otherwise commit + # the whole hypothesis as one batch, and _match is quadratic in the + # batch size (a 300-word batch blocks the event loop for seconds). + # Raw words, not normalized tokens, so number expansion cannot + # reinflate the bound; the dropped words are stale context the + # window matcher does not need. + raw_pending = raw_pending[-_PENDING_CAP:] + + pending = [] + for w in raw_pending: + pending.extend(normalize_word(w)) + + moved = False + held = False + if pending: + start = self._exp_anchor + 1 + width = self.window_mult * len(pending) + self.window_base + end, truncated = self._window_end(start, width) + window_toks = [tok for tok, _ in self._exp[start:end]] + take_flags = self._exp_take[start:end] + matches, _ = _match(pending, window_toks, take_flags) + ok = self._gate_passes(pending, matches, take_flags) + if not ok and truncated and matches: + # Window truncated by document end: the batch may be the + # script tail plus overflow speech. Overflow past the end + # cannot match anything, so retry the gate over only the + # pending words up to the last matched one. + last_i = max(i for i, _, _ in matches) + ok = self._gate_passes(pending[:last_i + 1], matches, + take_flags) + if not ok: + held = True + elif matches: + new_exp = start + max(j for _, j, _ in matches) + if new_exp > self._exp_anchor: + self._exp_anchor = new_exp + word = self._exp[new_exp][1] + # Never broadcast a take-block word as the anchor: the + # UI dims or hides takes, so snap to the next visible + # word (max() keeps the anchor monotonic when the + # fallback resolves backwards at end of script). + self._anchor_word = max(self._anchor_word, + self._visible_word(word)) + moved = True + return {"anchor": self._anchor_word, "moved": moved, "held": held} + + def _window_end(self, start, width): + """Window end so [start:end) holds `width` non-take tokens. + + Take-block tokens ride along for free (they delete at cost 0 in + _match), so a take span of any length never consumes window budget + and the script after it stays reachable. Returns (end, truncated) + where truncated means the document ended before the budget filled. + """ + n = len(self._exp) + end = start + budget = width + while end < n and budget > 0: + if not self._exp_take[end]: + budget -= 1 + end += 1 + return end, budget > 0 + + def _gate_passes(self, pending, matches, take_flags): + """The hold gate: is this batch evidence of on-script progress? + + Only informative matches count (see _informative); when the batch + has informative tokens, at least one must match exactly. Unmatched + non-take window words strictly between the first and last matched + position (interior deletions) are charged against the ratio, so + scattered stopword coincidences fail while a genuine contiguous + skip (deletions before the first match) passes. See the module + docstring, Hold rule. + """ + if not matches: + return False + inf_idx = {k for k, tok in enumerate(pending) if _informative(tok)} + if inf_idx: + relevant = [m for m in matches if m[0] in inf_idx] + # Require an exact informative match, UNLESS every informative + # pending token matched: a batch whose content words all match, + # merely fuzzily, is an ASR garble of an on-script read (the + # recorded streams produce e.g. "workshopod" for "workshop"), + # not an ad-lib; off-script speech carries unmatched content + # words and fails here or on the ratio below. + if (not any(c == 0.0 for _, _, c in relevant) + and {m[0] for m in relevant} != inf_idx): + return False + num = len(relevant) + denom = len(inf_idx) + else: + num = len(matches) + denom = len(pending) + matched_j = {j for _, j, _ in matches} + interior = sum(1 for j in range(min(matched_j) + 1, max(matched_j)) + if j not in matched_j and not take_flags[j]) + return num / (denom + interior) >= self.hold_threshold + + def _visible_word(self, word): + """Snap a take-block word to the nearest visible (non-take) word. + + Prefers the next non-take word (the one the presenter reads next, + and the one the UI can scroll to); falls back to the previous one + when the take runs to the end of the script. Returns -1 only when + the whole script is take blocks. + """ + if word not in self._take_words: + return word + w = word + 1 + while w < self._n_words and w in self._take_words: + w += 1 + if w < self._n_words: + return w + w = word - 1 + while w >= 0 and w in self._take_words: + w -= 1 + return w diff --git a/skills/mc-prompter/scripts/server/script_ingest.py b/skills/mc-prompter/scripts/server/script_ingest.py index e730ddf..5d90156 100644 --- a/skills/mc-prompter/scripts/server/script_ingest.py +++ b/skills/mc-prompter/scripts/server/script_ingest.py @@ -272,6 +272,47 @@ def _ensure_section(heading=None, level=0): } +def speakable_words(doc): + """Return the doc's speakable words in global word-index order. + + Flattens sections, then para/take blocks, then runs, splitting each + run's text on whitespace. Note blocks are excluded. This ordering MUST + match the UI's data-i word indexing (static/js/model.js wraps exactly + the non-whitespace chunks of each run in in + the same document order), so len(speakable_words(doc)) equals + doc["word-count"] and the Phase B aligner's anchor indexes agree with + the rendered spans. + """ + words = [] + for section in doc["sections"]: + for block in section["blocks"]: + if block["type"] in ("para", "take"): + for run in block["runs"]: + words.extend(run["text"].split()) + return words + + +def take_word_ranges(doc): + """Return [start, end) global word-index pairs for the doc's take blocks. + + Indexes are in speakable_words(doc) order (the UI's data-i indexing): + the ranges partition off exactly the words that speakable_words yields + from take blocks, one pair per take block in document order. Empty take + blocks (no runs survive ingestion) produce no range. Phase B's aligner + uses these to treat take words as free to skip. + """ + ranges = [] + i = 0 + for section in doc["sections"]: + for block in section["blocks"]: + if block["type"] in ("para", "take"): + n = sum(len(run["text"].split()) for run in block["runs"]) + if block["type"] == "take" and n: + ranges.append([i, i + n]) + i += n + return ranges + + def main(argv=None): parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) parser.add_argument("file", help="path to the script file") diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/adlib.json b/skills/mc-prompter/scripts/tests/fixtures/align/adlib.json new file mode 100644 index 0000000..fd3041d --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/adlib.json @@ -0,0 +1,703 @@ +{ + "script": "Cloud bills compound quietly. A local rig is a one-time cost that keeps paying you back. I ran the numbers last month and the break-even point was under ninety days.\n", + "events": [ + { + "tokens": [ + " C", + "l", + "ou", + "d" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is", + " a", + " one", + " tim", + "e" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is", + " a", + " one", + " tim", + "e", + " co", + "st" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is", + " a", + " one", + " tim", + "e", + " co", + "st", + " th", + "at", + " ke", + "ep", + "s", + " pay" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is", + " a", + " one", + " tim", + "e", + " co", + "st", + " th", + "at", + " ke", + "ep", + "s", + " pay", + "ing", + " you", + " b", + "ack" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " C", + "l", + "ou", + "d", + "ill", + "s", + " com", + "p", + "ound", + " qu", + "i", + "et", + "ly", + ".", + " A", + " loc", + "al", + " r", + "ig", + " is", + " a", + " one", + " tim", + "e", + " co", + "st", + " th", + "at", + " ke", + "ep", + "s", + " pay", + "ing", + " you", + " b", + "ack" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " The" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is", + "le", + " ex", + "per", + "im", + "ent" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is", + "le", + " ex", + "per", + "im", + "ent", + " wh", + "ich", + " was", + " ext", + "re" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is", + "le", + " ex", + "per", + "im", + "ent", + " wh", + "ich", + " was", + " ext", + "re", + "ly" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is", + "le", + " ex", + "per", + "im", + "ent", + " wh", + "ich", + " was", + " ext", + "re", + "ly", + " gen", + "er", + "ous", + " of", + " him" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " The", + " way", + " my", + " ne", + "igh", + "b", + "or", + " l", + "ent", + " me", + " his", + " g", + "ar", + "age", + " for", + " th", + "is", + "le", + " ex", + "per", + "im", + "ent", + " wh", + "ich", + " was", + " ext", + "re", + "ly", + " gen", + "er", + "ous", + " of", + " him" + ], + "segment": 1, + "final": true + }, + { + "tokens": [ + " I", + " r", + "an" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + "o", + " po", + "int", + " was", + " u", + "nder" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + "o", + " po", + "int", + " was", + " u", + "nder", + " n", + "in", + "ety", + " day", + "s" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + "o", + " po", + "int", + " was", + " u", + "nder", + " n", + "in", + "ety", + " day", + "s" + ], + "segment": 2, + "final": true + } + ], + "notes": "Mid-paragraph ad-lib: an off-script sentence about a neighbor's garage inserted between sentence two and sentence three, then return to script. Captured live from the nemotron-streaming int8 560ms model (sherpa-onnx 1.13.4, greedy_search, 120 ms chunks) over macOS say + ffmpeg synthesized speech; timestamps stripped; 9 word-level tail revisions present (a later partial's merged words are not an extension of the previous partial's)." +} diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/digits.json b/skills/mc-prompter/scripts/tests/fixtures/align/digits.json new file mode 100644 index 0000000..c3af557 --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/digits.json @@ -0,0 +1,343 @@ +{ + "script": "I ran the numbers last month and the break-even point was under 90 days. We shipped the first build in 2026 and the response was huge.\n", + "events": [ + { + "tokens": [ + " I", + " r", + "an" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + " was", + " u", + "nder" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + " was", + " u", + "nder", + " n", + "in", + "ety", + " day", + "s" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " I", + " r", + "an", + " the", + " num", + "bers", + " l", + "ast", + " mon", + "th", + "e", + " bre", + "ake", + "ven", + " po", + "int", + " was", + " u", + "nder", + " n", + "in", + "ety", + " day", + "s" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " We" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in", + " tw", + "ent", + "y", + " tw", + "ent", + "y" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in", + " tw", + "ent", + "y", + " tw", + "ent", + "y", + " six" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in", + " tw", + "ent", + "y", + " tw", + "ent", + "y", + " six", + " and", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in", + " tw", + "ent", + "y", + " tw", + "ent", + "y", + " six", + " and", + " the", + " res", + "p", + "on", + "se", + " was" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " We", + " sh", + "i", + "pp", + "ed", + " the", + " f", + "irst", + " bu", + "ild", + " in", + " tw", + "ent", + "y", + " tw", + "ent", + "y", + " six", + " and", + " the", + " res", + "p", + "on", + "se", + " was", + " h", + "u", + "ge" + ], + "segment": 1, + "final": true + } + ], + "notes": "Digits in script spoken as words: script has 90 and 2026, the reader says ninety and twenty twenty six. Captured live from the nemotron-streaming int8 560ms model (sherpa-onnx 1.13.4, greedy_search, 120 ms chunks) over macOS say + ffmpeg synthesized speech; timestamps stripped; 1 word-level tail revisions present (a later partial's merged words are not an extension of the previous partial's)." +} diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/skip.json b/skills/mc-prompter/scripts/tests/fixtures/align/skip.json new file mode 100644 index 0000000..64d4588 --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/skip.json @@ -0,0 +1,288 @@ +{ + "script": "Round trips to a data center add up. Local models answer before your finger leaves the key. That speed changes how you work, because you stop batching your thoughts and start thinking out loud.\n", + "events": [ + { + "tokens": [ + " R", + "ound", + " t", + "ri", + "ps", + " to", + " a" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " R", + "ound", + " t", + "ri", + "ps", + " to", + " a", + " d", + "ata", + " c", + "ent", + "er" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " R", + "ound", + " t", + "ri", + "ps", + " to", + " a", + " d", + "ata", + " c", + "ent", + "er", + " add", + " up" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " R", + "ound", + " t", + "ri", + "ps", + " to", + " a", + " d", + "ata", + " c", + "ent", + "er", + " add", + " up" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you", + " st", + "op" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you", + " st", + "op", + "ing", + " you", + "r", + " th", + "oug", + "ht" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you", + " st", + "op", + "ing", + " you", + "r", + " th", + "oug", + "ht", + "h", + "oug", + "ht", + "s", + " and" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you", + " st", + "op", + "ing", + " you", + "r", + " th", + "oug", + "ht", + "h", + "oug", + "ht", + "s", + " and", + " st", + "art", + " th", + "ink", + "ing" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " S", + "pe", + "ed", + " ch", + "ang", + "es", + " how", + " you", + " wor", + "k", + " be", + "ca", + "use", + " you", + " st", + "op", + "ing", + " you", + "r", + " th", + "oug", + "ht", + "h", + "oug", + "ht", + "s", + " and", + " st", + "art", + " th", + "ink", + "ing", + " out", + " l", + "ou", + "d" + ], + "segment": 1, + "final": true + } + ], + "notes": "Skipped sentence: the middle sentence of the script (nine words) is not spoken; the reader jumps from sentence one to sentence three. Captured live from the nemotron-streaming int8 560ms model (sherpa-onnx 1.13.4, greedy_search, 120 ms chunks) over macOS say + ffmpeg synthesized speech; timestamps stripped; 2 word-level tail revisions present (a later partial's merged words are not an extension of the previous partial's)." +} diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/stopword.json b/skills/mc-prompter/scripts/tests/fixtures/align/stopword.json new file mode 100644 index 0000000..6a2dd78 --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/stopword.json @@ -0,0 +1,302 @@ +{ + "script": "Today we will review the plan and the budget and the schedule and then we can go to the demo at the end.\n", + "events": [ + { + "tokens": [ + " T", + "od", + "ay" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the", + " pl", + "an" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the", + " pl", + "an", + " and", + " al", + "so", + " the" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the", + " pl", + "an", + " and", + " al", + "so", + " the", + " o", + "ther", + " th", + "ing", + " we" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the", + " pl", + "an", + " and", + " al", + "so", + " the", + " o", + "ther", + " th", + "ing", + " we", + " did" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " T", + "od", + "ay", + " we", + " w", + "ill", + " re", + "v", + "iew", + " the", + " pl", + "an", + " and", + " al", + "so", + " the", + " o", + "ther", + " th", + "ing", + " we", + " did" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " B", + "ud", + "get" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the", + " sch", + "ed", + "u", + "le" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the", + " sch", + "ed", + "u", + "le", + " and", + " the", + "n", + " we" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the", + " sch", + "ed", + "u", + "le", + " and", + " the", + "n", + " we", + " can", + " go", + " to", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the", + " sch", + "ed", + "u", + "le", + " and", + " the", + "n", + " we", + " can", + " go", + " to", + " the", + " de", + "m", + "o", + " at" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " B", + "ud", + "get", + " end", + " the", + " sch", + "ed", + "u", + "le", + " and", + " the", + "n", + " we", + " can", + " go", + " to", + " the", + " de", + "m", + "o", + " at", + " the", + " end" + ], + "segment": 1, + "final": true + } + ], + "notes": "Stopword-dense off-script ad-lib ('and also the other thing we did') spoken after 'the plan' in a script full of forward stopword duplicates; the aligner must hold through it and recover on 'and the budget and the schedule'. [capture: nemotron-streaming int8 560ms, 120ms chunks, 0 tail revisions observed]" +} \ No newline at end of file diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/tailrev.json b/skills/mc-prompter/scripts/tests/fixtures/align/tailrev.json new file mode 100644 index 0000000..1a0a99f --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/tailrev.json @@ -0,0 +1,304 @@ +{ + "script": "So that is the case: your money, your latency, and your data all point the same direction. Try one local tool this week and tell me how it goes in the comments. Catch you in the next one.\n", + "events": [ + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase", + ",", + " you", + "r", + " m", + "oney" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase", + ",", + " you", + "r", + " m", + "oney", + ",", + " you", + "r", + " l", + "at", + "ency", + ",", + " and" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase", + ",", + " you", + "r", + " m", + "oney", + ",", + " you", + "r", + " l", + "at", + "ency", + ",", + " and", + " you", + "r", + " d", + "ata" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase", + ",", + " you", + "r", + " m", + "oney", + ",", + " you", + "r", + " l", + "at", + "ency", + ",", + " and", + " you", + "r", + " d", + "ata", + " all", + " po", + "int", + " the", + " s", + "ame" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " Th", + "at", + " is", + " the", + " c", + "ase", + ",", + " you", + "r", + " m", + "oney", + ",", + " you", + "r", + " l", + "at", + "ency", + ",", + " and", + " you", + "r", + " d", + "ata", + " all", + " po", + "int", + " the", + " s", + "ame" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " One", + " loc" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " One", + " loc", + "al", + " to", + "ol", + " th", + "is" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " One", + " loc", + "al", + " to", + "ol", + " th", + "is", + " and", + " te", + "ll", + " me", + " how", + " it", + " go", + "es" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " One", + " loc", + "al", + " to", + "ol", + " th", + "is", + " and", + " te", + "ll", + " me", + " how", + " it", + " go", + "es", + " in", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " One", + " loc", + "al", + " to", + "ol", + " th", + "is", + " and", + " te", + "ll", + " me", + " how", + " it", + " go", + "es", + " in", + " the", + " com", + "ment", + "s" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " One", + " loc", + "al", + " to", + "ol", + " th", + "is", + " and", + " te", + "ll", + " me", + " how", + " it", + " go", + "es", + " in", + " the", + " com", + "ment", + "s" + ], + "segment": 1, + "final": true + }, + { + "tokens": [ + " C", + "atch", + " you", + " in", + " the", + " ne", + "xt", + " one" + ], + "segment": 2, + "final": false + }, + { + "tokens": [ + " C", + "atch", + " you", + " in", + " the", + " ne", + "xt", + " one" + ], + "segment": 2, + "final": true + } + ], + "notes": "Fast read (say -r 210) to stress tail revisions and merges. Captured live from the nemotron-streaming int8 560ms model (sherpa-onnx 1.13.4, greedy_search, 120 ms chunks) over macOS say + ffmpeg synthesized speech; timestamps stripped; 3 word-level tail revisions present (a later partial's merged words are not an extension of the previous partial's)." +} diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/takeskip.json b/skills/mc-prompter/scripts/tests/fixtures/align/takeskip.json new file mode 100644 index 0000000..3531d1c --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/takeskip.json @@ -0,0 +1,376 @@ +{ + "script": "Local inference is having a moment and the hardware finally caught up to the hype this year.\n\n[TAKE int1 42s-51s] This whole paragraph is already recorded footage from the interview where our guest explains the memory bandwidth story in careful and thorough detail for the audience at home.\n\nSo let us get the rig on the bench and see what it can actually do today.\n", + "events": [ + { + "tokens": [ + " L", + "oc", + "al" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard", + " fin" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard", + " fin", + "ally", + " ca", + "ught", + " up", + " to" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard", + " fin", + "ally", + " ca", + "ught", + " up", + " to", + " the", + " h", + "y", + "pe" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard", + " fin", + "ally", + " ca", + "ught", + " up", + " to", + " the", + " h", + "y", + "pe", + " th", + "is", + " y", + "ear" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " L", + "oc", + "al", + " inf", + "ere", + "nce", + " is", + " ha", + "ving", + " a", + " mom", + "ent", + " and", + " the", + " h", + "ard", + " fin", + "ally", + " ca", + "ught", + " up", + " to", + " the", + " h", + "y", + "pe", + " th", + "is", + " y", + "ear" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " So" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the", + " r", + "ig", + " on", + " the", + " b", + "en", + "ch" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the", + " r", + "ig", + " on", + " the", + " b", + "en", + "ch", + " and", + " see" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the", + " r", + "ig", + " on", + " the", + " b", + "en", + "ch", + " and", + " see", + " wh", + "at", + " it", + " can" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the", + " r", + "ig", + " on", + " the", + " b", + "en", + "ch", + " and", + " see", + " wh", + "at", + " it", + " can", + " act", + "u", + "ally", + " do" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " So", + " let", + " us", + " get", + " the", + " r", + "ig", + " on", + " the", + " b", + "en", + "ch", + " and", + " see", + " wh", + "at", + " it", + " can", + " act", + "u", + "ally", + " do", + " to", + "day" + ], + "segment": 1, + "final": true + } + ], + "notes": "A 30-word TAKE paragraph (longer than the match window) sits between intro and outro; the reader skips it aloud. Drive with take_word_ranges: the anchor must cross the take and never land inside it. [capture: nemotron-streaming int8 560ms, 120ms chunks, 0 tail revisions observed]" +} \ No newline at end of file diff --git a/skills/mc-prompter/scripts/tests/fixtures/align/verbatim.json b/skills/mc-prompter/scripts/tests/fixtures/align/verbatim.json new file mode 100644 index 0000000..acc452d --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/align/verbatim.json @@ -0,0 +1,570 @@ +{ + "script": "Hey everyone, welcome back to the workshop. Today we are digging into why running your tools locally beats renting them by the minute.\n\nStick around to the end, because the last demo surprised even me.\n", + "events": [ + { + "tokens": [ + " He", + "y", + " e", + "very" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol", + "s" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol", + "s", + " loc", + "ally" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol", + "s", + " loc", + "ally", + " be", + "at", + "s", + " re", + "nt", + "ing", + " the", + "m" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol", + "s", + " loc", + "ally", + " be", + "at", + "s", + " re", + "nt", + "ing", + " the", + "m", + " by", + " the", + " min", + "ute" + ], + "segment": 0, + "final": false + }, + { + "tokens": [ + " He", + "y", + " e", + "very", + "one", + ",", + " we", + "l", + "c", + "ome", + " b", + "ack", + " to", + " the", + " wor", + "ks", + "h", + "op", + "od", + ".", + " We", + " are", + " d", + "ig", + "g", + "ing", + " int", + "o", + " why", + " run", + "ning", + " you", + "r", + " to", + "ol", + "s", + " loc", + "ally", + " be", + "at", + "s", + " re", + "nt", + "ing", + " the", + "m", + " by", + " the", + " min", + "ute" + ], + "segment": 0, + "final": true + }, + { + "tokens": [ + " St", + "ick", + " ar", + "ound", + " to", + " the" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " St", + "ick", + " ar", + "ound", + " to", + " the", + " end" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " St", + "ick", + " ar", + "ound", + " to", + " the", + " end", + " be", + "ca", + "use", + " the", + " l", + "ast" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " St", + "ick", + " ar", + "ound", + " to", + " the", + " end", + " be", + "ca", + "use", + " the", + " l", + "ast", + " sur", + "pr", + "ised" + ], + "segment": 1, + "final": false + }, + { + "tokens": [ + " St", + "ick", + " ar", + "ound", + " to", + " the", + " end", + " be", + "ca", + "use", + " the", + " l", + "ast", + " sur", + "pr", + "ised", + " e", + "ven", + " me" + ], + "segment": 1, + "final": true + } + ], + "notes": "Verbatim read of two paragraphs; second paragraph after a 1.6 s pause (endpoint expected). Captured live from the nemotron-streaming int8 560ms model (sherpa-onnx 1.13.4, greedy_search, 120 ms chunks) over macOS say + ffmpeg synthesized speech; timestamps stripped; 5 word-level tail revisions present (a later partial's merged words are not an extension of the previous partial's)." +} diff --git a/skills/mc-prompter/scripts/tests/test-align.py b/skills/mc-prompter/scripts/tests/test-align.py new file mode 100644 index 0000000..3cb75e7 --- /dev/null +++ b/skills/mc-prompter/scripts/tests/test-align.py @@ -0,0 +1,607 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# /// +"""Tests for server/align.py (mc-prompter Phase B voice-follow). + +Run directly: + uv run skills/mc-prompter/scripts/tests/test-align.py + +Pure stdlib unittest; no network, no models, no downloads. The fixture files +under fixtures/align/ are REAL partial event streams captured offline from +the nemotron-streaming model (sherpa-onnx 1.13.4, greedy_search, 120 ms +chunks) over synthesized speech; each fixture carries the exact script +snippet it was read against. Tests drive the aligner from those committed +JSON streams and never touch the models. +""" + +import json +import sys +import time +import unittest +from pathlib import Path + +TESTS_DIR = Path(__file__).resolve().parent +sys.path.insert(0, str(TESTS_DIR.parent / "server")) + +from align import Aligner, merge_bpe, normalize_word # noqa: E402 +from script_ingest import ingest, speakable_words, take_word_ranges # noqa: E402 + +FIXTURES = TESTS_DIR / "fixtures" / "align" + + +def load_fixture(name): + fx = json.loads((FIXTURES / f"{name}.json").read_text(encoding="utf-8")) + doc = ingest(fx["script"]) + return fx, speakable_words(doc), take_word_ranges(doc) + + +def run_fixture(aligner, events): + """Feed every event; return the list of feed() results.""" + return [aligner.feed(e["tokens"], e["segment"], e["final"]) + for e in events] + + +def bpe(text): + """Turn plain words into space-prefixed BPE-ish pieces.""" + return [" " + w for w in text.split()] + + +def assert_monotonic(testcase, results): + anchors = [r["anchor"] for r in results] + for a, b in zip(anchors, anchors[1:]): + testcase.assertLessEqual(a, b, f"anchor retreated: {anchors}") + + +class TestBpeMerge(unittest.TestCase): + + def test_space_prefixed_pieces_start_words(self): + tokens = [" He", "y", " e", "very", "one", ","] + self.assertEqual(merge_bpe(tokens), ["Hey", "everyone,"]) + + def test_first_piece_without_space_starts_a_word(self): + self.assertEqual(merge_bpe(["We", " are"]), ["We", "are"]) + + def test_punctuation_pieces_attach_to_previous_word(self): + tokens = [" wor", "ks", "h", "op", ".", " We"] + self.assertEqual(merge_bpe(tokens), ["workshop.", "We"]) + + def test_empty_input(self): + self.assertEqual(merge_bpe([]), []) + + def test_whitespace_only_pieces_dropped(self): + self.assertEqual(merge_bpe([" ", " a"]), ["a"]) + + +class TestNormalization(unittest.TestCase): + + def test_casefold_and_punctuation(self): + self.assertEqual(normalize_word("Workshop."), ["workshop"]) + self.assertEqual(normalize_word("quietly,"), ["quietly"]) + + def test_hyphen_splits(self): + self.assertEqual(normalize_word("one-time"), ["one", "time"]) + self.assertEqual(normalize_word("break-even"), ["break", "even"]) + + def test_apostrophe_removed_not_split(self): + self.assertEqual(normalize_word("don't"), ["dont"]) + + def test_small_numbers(self): + self.assertEqual(normalize_word("90"), ["ninety"]) + self.assertEqual(normalize_word("42"), ["forty", "two"]) + self.assertEqual(normalize_word("7"), ["seven"]) + self.assertEqual(normalize_word("15"), ["fifteen"]) + + def test_hundreds(self): + self.assertEqual(normalize_word("300"), ["three", "hundred"]) + self.assertEqual(normalize_word("250"), ["two", "hundred", "fifty"]) + + def test_years_read_as_pairs(self): + self.assertEqual(normalize_word("2026"), ["twenty", "twenty", "six"]) + self.assertEqual(normalize_word("1995"), + ["nineteen", "ninety", "five"]) + + def test_year_special_cases(self): + self.assertEqual(normalize_word("2000"), ["two", "thousand"]) + self.assertEqual(normalize_word("2007"), ["two", "thousand", "seven"]) + self.assertEqual(normalize_word("1900"), ["nineteen", "hundred"]) + self.assertEqual(normalize_word("1907"), + ["nineteen", "oh", "seven"]) + + def test_thousands_separator_comma(self): + self.assertEqual(normalize_word("1,000"), ["one", "thousand"]) + + def test_ordinals(self): + self.assertEqual(normalize_word("1st"), ["first"]) + self.assertEqual(normalize_word("12th"), ["twelfth"]) + self.assertEqual(normalize_word("42nd"), ["forty", "two"]) + + def test_percent(self): + self.assertEqual(normalize_word("50%"), ["fifty", "percent"]) + + def test_symbol_only_word_yields_nothing(self): + self.assertEqual(normalize_word("..."), []) + + def test_apostrophe_family_stripped(self): + # ASCII, right/left curly quote, U+02BC modifier letter apostrophe + # (category Lm: survives isalnum and must be stripped explicitly), + # and U+201B reversed quote all normalize like a plain apostrophe. + for apo in ("'", "’", "‘", "ʼ", "‛"): + self.assertEqual(normalize_word(f"won{apo}t"), ["wont"], repr(apo)) + self.assertEqual(normalize_word(f"don{apo}t"), ["dont"], repr(apo)) + + +class TestAlignerSynthetic(unittest.TestCase): + """Unit behavior with hand-built token streams (no fixtures).""" + + WORDS = ("alpha bravo charlie delta echo foxtrot golf hotel india " + "juliet kilo lima").split() + + @staticmethod + def tokens(text): + """Turn plain words into space-prefixed BPE-ish pieces.""" + return [" " + w for w in text.split()] + + def test_anchor_starts_at_minus_one(self): + self.assertEqual(Aligner(self.WORDS).anchor, -1) + + def test_provisional_tail_not_committed(self): + al = Aligner(self.WORDS, k_provisional=4) + r = al.feed(self.tokens("alpha bravo charlie delta"), 0, False) + self.assertEqual(r["anchor"], -1) # all 4 words provisional + r = al.feed(self.tokens("alpha bravo charlie delta echo"), 0, False) + self.assertEqual(r["anchor"], 0) # only alpha committed + + def test_final_commits_whole_hypothesis(self): + al = Aligner(self.WORDS) + r = al.feed(self.tokens("alpha bravo charlie delta"), 0, True) + self.assertEqual(r["anchor"], 3) + self.assertTrue(r["moved"]) + + def test_tail_revision_absorbed(self): + al = Aligner(self.WORDS, k_provisional=2) + al.feed(self.tokens("alpha bravo charlie del"), 0, False) + r = al.feed(self.tokens("alpha bravo charlie delta echo"), 0, False) + self.assertEqual(r["anchor"], 2) + r = al.feed(self.tokens("alpha bravo charlie delta echo"), 0, True) + self.assertEqual(r["anchor"], 4) + + def test_new_segment_resets_tracking_not_anchor(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo charlie"), 0, True) + self.assertEqual(al.anchor, 2) + r = al.feed(self.tokens("delta echo"), 1, True) + self.assertEqual(r["anchor"], 4) + + def test_offscript_speech_holds(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo"), 0, True) + r = al.feed(self.tokens("penguin zebra walrus"), 1, True) + self.assertTrue(r["held"]) + self.assertFalse(r["moved"]) + self.assertEqual(r["anchor"], 1) + + def test_recovers_after_hold(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo"), 0, True) + al.feed(self.tokens("penguin zebra walrus"), 1, True) + r = al.feed(self.tokens("charlie delta echo"), 2, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 4) + + def test_skip_followed_within_window(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo"), 0, True) + r = al.feed(self.tokens("golf hotel india"), 1, True) + self.assertEqual(r["anchor"], 8) + + def test_jump_beyond_window_holds(self): + words = [f"w{i}" for i in range(200)] + al = Aligner(words, window_base=10, window_mult=2) + al.feed(self.tokens("w0 w1"), 0, True) + r = al.feed(self.tokens("w150 w151"), 1, True) + self.assertTrue(r["held"]) + self.assertEqual(r["anchor"], 1) + + def test_set_anchor_backwards_and_clamped(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo charlie delta echo"), 0, True) + al.set_anchor(1) + self.assertEqual(al.anchor, 1) + al.set_anchor(9999) + self.assertEqual(al.anchor, len(self.WORDS) - 1) + al.set_anchor(-5) + self.assertEqual(al.anchor, -1) + + def test_set_anchor_consumes_current_hypothesis(self): + al = Aligner(self.WORDS) + al.feed(self.tokens("alpha bravo charlie delta echo"), 0, False) + al.set_anchor(7) + # Re-feeding the same hypothesis must not re-match old words. + r = al.feed(self.tokens("alpha bravo charlie delta echo"), 0, True) + self.assertEqual(r["anchor"], 7) + self.assertFalse(r["moved"]) + # New words after the jump match from the new position. + r = al.feed(self.tokens("india juliet"), 1, True) + self.assertEqual(r["anchor"], 9) + + def test_number_in_script_matches_spoken_words(self): + al = Aligner("the answer is 42 exactly".split()) + r = al.feed(self.tokens("the answer is forty two exactly"), 0, True) + self.assertEqual(r["anchor"], 4) + + def test_spoken_digits_match_worded_script(self): + al = Aligner("wait ninety days now".split()) + r = al.feed(self.tokens("wait 90 days now"), 0, True) + self.assertEqual(r["anchor"], 3) + + +class TestTakeSkipping(unittest.TestCase): + """Take blocks are free to skip and the anchor never sits inside one. + + Take paragraphs are already-recorded footage, dimmed or hidden in the + UI; a presenter normally reads straight past them. The take here is 28 + words, far wider than the match window for small pending batches, so + without take-aware skipping the aligner would stall at it forever. + """ + + INTRO = ("the local rig finally paid for itself after ninety days " + "running quietly").split() + TAKE = ("this whole paragraph is already recorded footage from the " + "interview where our guest explains the memory bandwidth story " + "in careful and thorough detail for the audience at home").split() + OUTRO = "so let us get back to the bench and wrap things up".split() + + def build(self, words=None, ranges=None): + if words is None: + words = self.INTRO + self.TAKE + self.OUTRO + ranges = [[len(self.INTRO), len(self.INTRO) + len(self.TAKE)]] + return Aligner(words, take_ranges=ranges), set(range(*ranges[0])) + + def test_long_take_skipped_in_one_final(self): + al, take_words = self.build() + al.feed(bpe(" ".join(self.INTRO)), 0, True) + self.assertEqual(al.anchor, len(self.INTRO) - 1) + r = al.feed(bpe(" ".join(self.OUTRO)), 1, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], + len(self.INTRO) + len(self.TAKE) + len(self.OUTRO) - 1) + + def test_take_longer_than_window_crossed_by_small_batch(self): + # 3 pending words give a window of 2*3+10 = 16 non-take tokens, + # much narrower than the 28-word take; the take must not consume + # window budget. + al, take_words = self.build() + al.feed(bpe(" ".join(self.INTRO)), 0, True) + r = al.feed(bpe("so let us"), 1, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], len(self.INTRO) + len(self.TAKE) + 2) + + def test_take_read_aloud_still_tracks(self): + # If the presenter DOES read the take, substitution matching works + # as usual; the broadcast anchor snaps to the next visible word. + al, take_words = self.build() + al.feed(bpe(" ".join(self.INTRO)), 0, True) + r = al.feed(bpe(" ".join(self.TAKE)), 1, True) + self.assertFalse(r["held"]) + self.assertTrue(r["moved"]) + self.assertEqual(r["anchor"], len(self.INTRO) + len(self.TAKE)) + r = al.feed(bpe(" ".join(self.OUTRO)), 2, True) + self.assertEqual(r["anchor"], + len(self.INTRO) + len(self.TAKE) + len(self.OUTRO) - 1) + + def test_anchor_never_inside_take(self): + for feeds in ( + [self.INTRO, self.OUTRO], + [self.INTRO, self.TAKE, self.OUTRO], + [self.INTRO, self.TAKE[:9], self.OUTRO], + ): + al, take_words = self.build() + for seg, chunk in enumerate(feeds): + # Word-by-word partials plus a final, like a real stream. + for k in range(1, len(chunk) + 1): + r = al.feed(bpe(" ".join(chunk[:k])), seg, False) + self.assertNotIn(r["anchor"], take_words) + r = al.feed(bpe(" ".join(chunk)), seg, True) + self.assertNotIn(r["anchor"], take_words) + + def test_take_at_end_of_script_falls_back_to_previous_word(self): + words = self.INTRO + self.TAKE + ranges = [[len(self.INTRO), len(words)]] + al, take_words = self.build(words, ranges) + al.feed(bpe(" ".join(self.INTRO)), 0, True) + r = al.feed(bpe(" ".join(self.TAKE)), 1, True) + self.assertNotIn(r["anchor"], take_words) + self.assertEqual(r["anchor"], len(self.INTRO) - 1) + + +class TestStopwordGate(unittest.TestCase): + """Stopword coincidences must not move the anchor (reviewer's repro).""" + + SCRIPT = ("first we review the plan and the budget and the schedule " + "and then we can go to the demo today").split() + + def test_stopword_dense_adlib_holds_then_recovers(self): + al = Aligner(self.SCRIPT) + al.feed(bpe("first we review the plan"), 0, True) + self.assertEqual(al.anchor, 4) + # Fully off-script, stopword-dense ad-lib: "and", "the", "we" all + # appear ahead in the script and "other" fuzzy-matches "the"; the + # old ratio counted those at full weight and jumped the anchor 9 + # words into unspoken script, after which the resumed read matched + # later duplicates and compounded the drift. + r = al.feed(bpe("and also the other thing we did"), 1, True) + self.assertTrue(r["held"]) + self.assertEqual(r["anchor"], 4) + # Resuming on script matches the ORIGINAL duplicates, not later + # ones, because the anchor never drifted. + r = al.feed(bpe("and the budget and the schedule"), 2, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], self.SCRIPT.index("schedule")) + + def test_lone_stopword_on_script_still_advances(self): + # Small partial commits are often pure stopwords ("and the"); a + # contiguous match at the window start is on-script continuation + # and must not hold, or verbatim reads would stutter. + al = Aligner(self.SCRIPT) + al.feed(bpe("first we review the plan"), 0, True) + r = al.feed(bpe("and the"), 1, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 6) + + def test_garbled_content_word_still_advances(self): + # An ASR garble of an on-script read ("workshopod" for "workshop") + # has every content word matched, merely fuzzily; that is not an + # ad-lib and must not hold (the verbatim capture contains exactly + # this). + al = Aligner("welcome to the workshop everyone".split()) + al.feed(bpe("welcome to the"), 0, True) + r = al.feed(bpe("workshopod"), 1, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 3) + + +class TestPendingClamp(unittest.TestCase): + """feed() clamps each batch to the last 24 raw words. + + A mid-segment Aligner rebuild (script edit while a segment is in + flight) makes the next partial commit the whole hypothesis as one + batch; _match is quadratic in batch size, so an unclamped 300-word + batch blocks the server's event loop for seconds. + """ + + @staticmethod + def alpha_words(n): + syl = [c + v for c in "bcdfghjklmnprst" for v in "aeiou"] + return [a + b for a in syl for b in syl][:n] + + def test_300_word_batch_bounded_wall_time(self): + words = self.alpha_words(600) + al = Aligner(words) + batch = bpe(" ".join(words[:300])) + t0 = time.perf_counter() + al.feed(batch, 0, True) + elapsed = time.perf_counter() - t0 + # Generous CI-safe bound; unclamped this took seconds. + self.assertLess(elapsed, 0.2) + + def test_clamp_keeps_recent_speech_matching(self): + words = self.alpha_words(600) + al = Aligner(words) + al.set_anchor(275) + # 200-word batch: only the last 24 words survive the clamp, and + # they are exactly the recent speech that continues the anchor. + r = al.feed(bpe(" ".join(words[100:300])), 0, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 299) + + +class TestEndOfScript(unittest.TestCase): + """Script tail plus overflow speech in one batch still anchors the tail.""" + + def test_tail_and_overflow_in_same_batch(self): + al = Aligner("hello world".split()) + r = al.feed(bpe("hello world thanks for watching everyone"), 0, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 1) + + def test_tail_shorter_than_half_the_batch(self): + al = Aligner("okay folks welcome back".split()) + r = al.feed(bpe("okay folks welcome back thanks so much for " + "watching everyone goodbye"), 0, True) + self.assertFalse(r["held"]) + self.assertEqual(r["anchor"], 3) + + def test_overflow_after_anchored_tail_holds(self): + al = Aligner("hello world".split()) + al.feed(bpe("hello world"), 0, True) + r = al.feed(bpe("thanks for watching everyone"), 1, True) + self.assertTrue(r["held"]) + self.assertEqual(r["anchor"], 1) + + +class FixtureCase(unittest.TestCase): + """Shared helpers for the recorded-stream scenarios.""" + + def drive(self, name): + fx, words, ranges = load_fixture(name) + aligner = Aligner(words, take_ranges=ranges) + results = run_fixture(aligner, fx["events"]) + assert_monotonic(self, results) + take_words = {i for a, b in ranges for i in range(a, b)} + for r in results: + self.assertNotIn(r["anchor"], take_words, + "anchor landed inside a take block") + return fx, words, results + + +class TestVerbatimFixture(FixtureCase): + + def test_reaches_near_end_and_monotonic(self): + fx, words, results = self.drive("verbatim") + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + def test_no_holds_on_verbatim_read(self): + fx, words, results = self.drive("verbatim") + self.assertEqual(sum(r["held"] for r in results), 0) + + +class TestAdlibFixture(FixtureCase): + + def test_holds_during_adlib_then_recovers(self): + fx, words, results = self.drive("adlib") + held = [r for r in results if r["held"]] + self.assertGreaterEqual(len(held), 3, + "ad-lib produced no holds") + # The ad-lib starts right after "back." in the script; while held, + # the anchor must not run away into unspoken script. + boundary = words.index("back.") + for r in held: + self.assertLessEqual(r["anchor"], boundary + 3) + # After the ad-lib the aligner recovers to (near) the end. + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestSkipFixture(FixtureCase): + + def test_skip_followed_within_window(self): + fx, words, results = self.drive("skip") + # The middle sentence (ending at "key.") was never spoken; the + # anchor must jump across it and reach near the end anyway. + skipped_end = words.index("key.") + self.assertGreater(results[-1]["anchor"], skipped_end) + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestDigitsFixture(FixtureCase): + + def test_numbers_do_not_stall(self): + fx, words, results = self.drive("digits") + anchors = [r["anchor"] for r in results] + # Anchor moves past "90" (spoken "ninety") and past "2026" (spoken + # "twenty twenty six") and reaches near the end. + self.assertGreater(max(anchors), words.index("90")) + self.assertGreater(max(anchors), words.index("2026")) + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestTailRevisionFixture(FixtureCase): + + def test_fixture_contains_word_level_tail_revisions(self): + fx, words, _ = load_fixture("tailrev") + prev = {} + revisions = 0 + for e in fx["events"]: + merged = merge_bpe(e["tokens"]) + before = prev.get(e["segment"], []) + if merged[:len(before)] != before: + revisions += 1 + prev[e["segment"]] = merged + self.assertGreaterEqual(revisions, 1) + + def test_anchor_never_exceeds_then_retreats(self): + fx, words, results = self.drive("tailrev") + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestStopwordAdlibFixture(FixtureCase): + """Captured stream: stopword-dense ad-lib after 'plan' in a script full + of forward stopword duplicates (the drift reproduction from review).""" + + def test_holds_through_adlib_without_drift(self): + fx, words, results = self.drive("stopword") + self.assertGreaterEqual(sum(r["held"] for r in results), 1) + # While the ad-lib plays (segment 0 carries it), the anchor must + # not run ahead of the resume point: "budget" and everything after + # are unspoken until segment 1. + budget = words.index("budget") + seg0 = [r for r, e in zip(results, fx["events"]) + if e["segment"] == 0] + for r in seg0: + self.assertLess(r["anchor"], budget) + # The resumed read recovers to the end of the script. + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestTakeSkipFixture(FixtureCase): + """Captured stream: a 28-word TAKE paragraph (longer than the window) + between intro and outro, skipped aloud by the reader.""" + + def test_take_crossed_and_never_entered(self): + fx, words, results = self.drive("takeskip") + # drive() already asserts the anchor is never inside the take. + _, _, ranges = load_fixture("takeskip") + self.assertEqual(len(ranges), 1) + take_start, take_end = ranges[0] + self.assertGreater(take_end - take_start, 20) + anchors = [r["anchor"] for r in results] + self.assertGreaterEqual(max(anchors), take_end) + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + # The outro is followed promptly: no holds after crossing the take. + crossed = next(i for i, a in enumerate(anchors) if a >= take_end) + for r in results[crossed:]: + self.assertFalse(r["held"]) + + +class TestSetAnchorRecovery(FixtureCase): + + def test_backward_jump_then_reread_recovers(self): + fx, words, _ = load_fixture("verbatim") + aligner = Aligner(words) + run_fixture(aligner, fx["events"]) + self.assertGreaterEqual(aligner.anchor, len(words) - 3) + aligner.set_anchor(5) + self.assertEqual(aligner.anchor, 5) + # Re-read the same passage (new segment ids simulate the re-take). + results = [aligner.feed(e["tokens"], e["segment"] + 100, e["final"]) + for e in fx["events"]] + assert_monotonic(self, results) + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + def test_manual_jump_at_skip_point(self): + fx, words, _ = load_fixture("skip") + aligner = Aligner(words) + events = fx["events"] + first_final = next(i for i, e in enumerate(events) if e["final"]) + for e in events[:first_final + 1]: + aligner.feed(e["tokens"], e["segment"], e["final"]) + # The creator clicks the last word of the skipped sentence. + target = words.index("key.") + aligner.set_anchor(target) + self.assertEqual(aligner.anchor, target) + results = [aligner.feed(e["tokens"], e["segment"], e["final"]) + for e in events[first_final + 1:]] + assert_monotonic(self, results) + self.assertGreaterEqual(results[-1]["anchor"], len(words) - 3) + + +class TestFixtureHygiene(unittest.TestCase): + """The committed fixtures stay CI-safe and schema-correct.""" + + NAMES = ("verbatim", "adlib", "skip", "digits", "tailrev", + "stopword", "takeskip") + + def test_schema_and_size(self): + for name in self.NAMES: + path = FIXTURES / f"{name}.json" + self.assertLess(path.stat().st_size, 100_000) + fx = json.loads(path.read_text(encoding="utf-8")) + self.assertIn("script", fx) + self.assertIn("events", fx) + self.assertIn("notes", fx) + for e in fx["events"]: + self.assertEqual(set(e), {"tokens", "segment", "final"}) + self.assertIsInstance(e["tokens"], list) + self.assertIsInstance(e["segment"], int) + self.assertIsInstance(e["final"], bool) + + def test_segments_are_monotonic_in_every_stream(self): + for name in self.NAMES: + fx = json.loads((FIXTURES / f"{name}.json") + .read_text(encoding="utf-8")) + segs = [e["segment"] for e in fx["events"]] + self.assertEqual(segs, sorted(segs)) + + +if __name__ == "__main__": + unittest.main() diff --git a/skills/mc-prompter/scripts/tests/test-script_ingest.py b/skills/mc-prompter/scripts/tests/test-script_ingest.py index 7922186..bc079e7 100644 --- a/skills/mc-prompter/scripts/tests/test-script_ingest.py +++ b/skills/mc-prompter/scripts/tests/test-script_ingest.py @@ -21,7 +21,7 @@ TESTS_DIR = Path(__file__).resolve().parent sys.path.insert(0, str(TESTS_DIR.parent / "server")) -from script_ingest import ingest # noqa: E402 +from script_ingest import ingest, speakable_words, take_word_ranges # noqa: E402 FIXTURE = TESTS_DIR / "fixtures" / "sample-script.md" @@ -312,5 +312,112 @@ def test_empty_brackets_never_leak(self): self.assertIn("Also here.", text) +class TestSpeakableWords(unittest.TestCase): + """speakable_words() must mirror the UI's data-i word indexing. + + static/js/model.js renders each run's text by splitting on whitespace + (String.split(/(\\s+)/), keeping only non-whitespace chunks) and wraps + each chunk in a data-i span, incrementing one global counter across + para and take blocks in document order; note blocks are skipped. That + is exactly run["text"].split() flattened in the same order, so the + list length must equal doc["word-count"]. + """ + + def test_length_equals_word_count_on_fixture(self): + doc = ingest(FIXTURE.read_text(encoding="utf-8")) + words = speakable_words(doc) + self.assertEqual(len(words), doc["word-count"]) + + def test_order_spans_sections_blocks_and_runs(self): + doc = ingest( + "## A\n\nOne two. [INVENTED] Three four.\n\n" + "[TAKE int1 1.0s-2.0s]\nFive six.\n\n" + "[a note that is never indexed]\n\n## B\n\nSeven.\n" + ) + words = speakable_words(doc) + self.assertEqual(words, ["One", "two.", "Three", "four.", + "Five", "six.", "Seven."]) + self.assertEqual(len(words), doc["word-count"]) + + def test_notes_excluded(self): + doc = ingest("Speak this [not this] aloud.\n") + self.assertEqual(speakable_words(doc), ["Speak", "this", "aloud."]) + + +class TestTakeWordRanges(unittest.TestCase): + """take_word_ranges() must use speakable_words() indexing exactly. + + The Phase B aligner treats take words as free to skip; the ranges are + [start, end) pairs into the same global word index the UI renders as + data-i spans, so for every range, speakable_words(doc)[start:end] must + be precisely that take block's words. + """ + + def take_block_words(self, doc): + out = [] + for section in doc["sections"]: + for block in section["blocks"]: + if block["type"] == "take": + words = [] + for run in block["runs"]: + words.extend(run["text"].split()) + out.append(words) + return out + + def assert_parity(self, doc): + words = speakable_words(doc) + ranges = take_word_ranges(doc) + expected = self.take_block_words(doc) + self.assertEqual(len(ranges), len(expected)) + for (start, end), block_words in zip(ranges, expected): + self.assertEqual(words[start:end], block_words) + + def test_no_takes_yields_empty(self): + self.assertEqual(take_word_ranges(ingest("Just plain talk.\n")), []) + + def test_single_take_between_paras(self): + doc = ingest( + "Intro of four words.\n\n" + "[TAKE int1 1.0s-2.0s]\nRecorded middle bit here.\n\n" + "Outro words.\n" + ) + self.assertEqual(take_word_ranges(doc), [[4, 8]]) + self.assert_parity(doc) + + def test_take_at_document_start_and_end(self): + doc = ingest( + "[TAKE a 0s-1s]\nFirst take.\n\nMiddle para.\n\n" + "[TAKE b 2s-3s]\nLast take here.\n" + ) + self.assertEqual(take_word_ranges(doc), [[0, 2], [4, 7]]) + self.assert_parity(doc) + + def test_ranges_span_sections_and_skip_notes(self): + doc = ingest( + "## A\n\nOne two. [INVENTED] Three four.\n\n" + "[TAKE int1 1.0s-2.0s]\nFive six.\n\n" + "[a note that is never indexed]\n\n## B\n\nSeven.\n" + ) + self.assertEqual(take_word_ranges(doc), [[4, 6]]) + self.assert_parity(doc) + + def test_parity_on_fixture(self): + doc = ingest(FIXTURE.read_text(encoding="utf-8")) + ranges = take_word_ranges(doc) + self.assertEqual(len(ranges), 2) + self.assert_parity(doc) + words = speakable_words(doc) + for start, end in ranges: + self.assertGreaterEqual(start, 0) + self.assertGreater(end, start) + self.assertLessEqual(end, len(words)) + + def test_plain_fmt(self): + doc = ingest("Spoken bit here.\n\n[TAKE t1 0.5s-2.5s]\nDone now.\n", + fmt="plain") + self.assertEqual(take_word_ranges(doc), [[3, 5]]) + self.assert_parity(doc) + + if __name__ == "__main__": unittest.main() From 09aaaf603c28e34591d68bf1c4adff2f3b6f0b9e Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 21:32:46 -0500 Subject: [PATCH 3/8] Add streaming ASR engine, capture ownership protocol, and aligner wiring --- skills/mc-prompter/scripts/server/asr.py | 313 ++++++++++ skills/mc-prompter/scripts/server/main.py | 403 ++++++++++++- skills/mc-prompter/scripts/server/replay.py | 293 +++++++++ .../mc-prompter/scripts/tests/test-server.py | 570 +++++++++++++++++- 4 files changed, 1544 insertions(+), 35 deletions(-) create mode 100644 skills/mc-prompter/scripts/server/asr.py create mode 100644 skills/mc-prompter/scripts/server/replay.py diff --git a/skills/mc-prompter/scripts/server/asr.py b/skills/mc-prompter/scripts/server/asr.py new file mode 100644 index 0000000..91f6807 --- /dev/null +++ b/skills/mc-prompter/scripts/server/asr.py @@ -0,0 +1,313 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# dependencies = ["numpy", "sherpa-onnx==1.13.4"] +# /// +"""Streaming ASR engine for mc-prompter (Phase B, voice-follow). + +This module is imported LAZILY by server/main.py, only when the server is +launched with --models-dir and an implemented --asr-provider. It normally +runs inside the prompter-lab workspace venv (which pins sherpa-onnx==1.13.4, +numpy); the PEP 723 header above documents the dependencies and allows a +direct `uv run asr.py` sanity import. CI never imports this module. + +Contract (binding, see the Phase B contract): + AsrEngine(models_dir, on_event, provider="nemotron-streaming", + num_threads=2) + models_dir the workspace models directory containing + nemotron-streaming/{encoder.int8.onnx, decoder.int8.onnx, + joiner.int8.onnx, tokens.txt} and silero_vad.onnx + on_event callable taking one dict; called from the WORKER thread. + The caller wraps it with loop.call_soon_threadsafe. + provider "nemotron-streaming" is the only implemented provider; + anything else raises ValueError (planned lanes never + pretend to run). + engine.start() spawns the worker thread; model load happens on the + worker so start() returns immediately. ready flips true + (status event) once the recognizer and VAD are built. + Restart after stop() is supported: start() begins from + fresh state (queue drained, ready/behind/dropped reset) + so a leftover stop sentinel or stale pre-stop audio can + never reach a new worker. + engine.stop() stops and joins the worker. + engine.alive True while the worker thread is running. + engine.feed(pcm16_bytes) called from the event loop with raw little + endian PCM16 mono 16 kHz bytes of any length (the + browser sends ~3840 bytes per 120 ms). Bounded queue + (maxsize 50); on overflow the OLDEST frame is dropped + and behind=True until the queue drains below half. + engine.stats {"queue": n, "behind": bool, "ready": bool} + +Events emitted through on_event: + {"kind": "partial", "segment": n, "text": str, "tokens": [...]} + on every hypothesis change. tokens are BPE pieces; a piece starting + with a space starts a new word. The tail of the hypothesis may + revise between partials; the aligner handles that. + {"kind": "final", "segment": n, "text": str, "tokens": [...]} + at an endpoint, before the recognizer resets (only when the + hypothesis is non-empty; endpoints on pure silence are not final + events). The next partial starts segment n+1 with a fresh + hypothesis. + {"kind": "vad", "speaking": bool} + on speaking/silence transitions (silero VAD, 512-sample windows). + {"kind": "status", "ready": bool, "behind": bool, "queue": int} + on ready/behind changes. If model load fails (imports included) or + the decode loop crashes mid-session, an error line goes to stderr, + ready flips False, the queue is drained (stats reflect death), a + final status event is emitted, and the worker exits; the engine + never fakes readiness. + +Pipeline per frame (all on the worker thread): PCM16 LE bytes -> float32 in +[-1, 1); 512-sample windows into the VAD (drained each window so its buffer +never grows); the full chunk into the recognizer stream; decode while ready; +endpoint -> final + reset + segment increment. +""" + +import json +import queue +import sys +import threading +from pathlib import Path + +PROVIDERS = ("nemotron-streaming",) +SAMPLE_RATE = 16000 +VAD_WINDOW = 512 +QUEUE_MAX = 50 + + +class AsrEngine: + """Dedicated-thread streaming recognizer with a bounded drop-oldest queue.""" + + def __init__(self, models_dir, on_event, provider="nemotron-streaming", + num_threads=2): + if provider not in PROVIDERS: + raise ValueError( + f"asr provider {provider!r} is not implemented; " + f"implemented: {', '.join(PROVIDERS)}" + ) + self.models_dir = Path(models_dir) + self.on_event = on_event + self.provider = provider + self.num_threads = num_threads + self._queue = queue.Queue(maxsize=QUEUE_MAX) + self._behind = False + self._ready = False + self._dropped = 0 + self._stop = threading.Event() + self._thread = None + + def start(self): + """Spawn the worker thread (idempotent). Model load happens there. + + Every start begins from fresh state: a prior stop() leaves its None + sentinel (and possibly stale pre-stop audio) in the queue, which a + restarted worker must never inherit or it would decode old audio and + then die on the sentinel while reporting ready. + """ + if self._thread is not None: + return + self._drain_queue() + self._ready = False + self._behind = False + self._dropped = 0 + self._stop.clear() + self._thread = threading.Thread( + target=self._run, name="mc-prompter-asr", daemon=True + ) + self._thread.start() + + def stop(self): + """Signal the worker and join it (safe to call more than once).""" + self._stop.set() + try: + self._queue.put_nowait(None) + except queue.Full: + pass + if self._thread is not None: + self._thread.join(timeout=10) + self._thread = None + + def feed(self, pcm16_bytes): + """Enqueue a PCM16 frame; drop the oldest frame on overflow.""" + try: + self._queue.put_nowait(pcm16_bytes) + except queue.Full: + try: + self._queue.get_nowait() + except queue.Empty: + pass + try: + self._queue.put_nowait(pcm16_bytes) + except queue.Full: + pass + self._dropped += 1 + if not self._behind: + self._behind = True + self._emit_status() + + @property + def stats(self): + return { + "queue": self._queue.qsize(), + "behind": self._behind, + "ready": self._ready, + } + + @property + def alive(self): + """True while the worker thread is running.""" + return self._thread is not None and self._thread.is_alive() + + def _drain_queue(self): + """Empty the audio queue (start reset and worker-death cleanup).""" + while True: + try: + self._queue.get_nowait() + except queue.Empty: + return + + def _emit(self, event): + try: + self.on_event(event) + except Exception as exc: # a sink bug must never kill the worker + print(f"asr: on_event raised: {exc}", file=sys.stderr) + + def _emit_status(self): + self._emit( + { + "kind": "status", + "ready": self._ready, + "behind": self._behind, + "queue": self._queue.qsize(), + } + ) + + def _build(self, sherpa_onnx): + """Construct the recognizer and VAD (worker thread only).""" + model_dir = self.models_dir / self.provider + recognizer = sherpa_onnx.OnlineRecognizer.from_transducer( + tokens=str(model_dir / "tokens.txt"), + encoder=str(model_dir / "encoder.int8.onnx"), + decoder=str(model_dir / "decoder.int8.onnx"), + joiner=str(model_dir / "joiner.int8.onnx"), + num_threads=self.num_threads, + sample_rate=SAMPLE_RATE, + feature_dim=80, + enable_endpoint_detection=True, + rule1_min_trailing_silence=2.4, + rule2_min_trailing_silence=1.2, + rule3_min_utterance_length=300, + decoding_method="greedy_search", + model_type="nemo_transducer", + ) + vad_cfg = sherpa_onnx.VadModelConfig() + vad_cfg.silero_vad.model = str(self.models_dir / "silero_vad.onnx") + vad_cfg.silero_vad.threshold = 0.5 + vad_cfg.silero_vad.min_silence_duration = 0.4 + vad_cfg.sample_rate = SAMPLE_RATE + vad = sherpa_onnx.VoiceActivityDetector( + vad_cfg, buffer_size_in_seconds=30 + ) + return recognizer, vad + + def _load_modules(self): + """Import the heavy dependencies (worker thread only; test seam).""" + import numpy + import sherpa_onnx + return numpy, sherpa_onnx + + def _run(self): + # Imports sit inside the guard: a half-installed venv (numpy or + # sherpa-onnx missing) must fail with a ready=False status, not a + # silent thread death via the default excepthook. + try: + np, sherpa_onnx = self._load_modules() + recognizer, vad = self._build(sherpa_onnx) + except Exception as exc: + print(f"asr: model load failed: {exc}", file=sys.stderr) + self._ready = False + self._drain_queue() + self._emit_status() + return + stream = recognizer.create_stream() + self._ready = True + self._emit_status() + try: + self._decode_loop(np, recognizer, vad, stream) + except Exception as exc: + # A mid-session crash (onnxruntime error, bad frame) must never + # leave ready=True on a dead worker: flag it, empty the dead + # queue so stats reflect death, tell the clients, exit. + print(f"asr: decode loop crashed: {exc}", file=sys.stderr) + self._ready = False + self._drain_queue() + self._emit_status() + + def _decode_loop(self, np, recognizer, vad, stream): + segment = 0 + prev_text = "" + speaking = False + vad_buf = np.zeros(0, dtype=np.float32) + + while not self._stop.is_set(): + try: + item = self._queue.get(timeout=0.2) + except queue.Empty: + continue + if item is None: + break + if self._behind and self._queue.qsize() < QUEUE_MAX // 2: + self._behind = False + self._emit_status() + + data = item + if len(data) % 2: + data = data[:-1] + if not data: + continue + samples = ( + np.frombuffer(data, dtype="= VAD_WINDOW: + vad.accept_waveform(vad_buf[:VAD_WINDOW]) + vad_buf = vad_buf[VAD_WINDOW:] + while not vad.empty(): + vad.pop() + now_speaking = bool(vad.is_speech_detected()) + if now_speaking != speaking: + speaking = now_speaking + self._emit({"kind": "vad", "speaking": speaking}) + + # Recognizer: any chunk length is fine. + stream.accept_waveform(SAMPLE_RATE, samples) + while recognizer.is_ready(stream): + recognizer.decode_stream(stream) + result = json.loads(recognizer.get_result_as_json_string(stream)) + text = result.get("text", "") + tokens = result.get("tokens", []) + if text != prev_text: + self._emit( + { + "kind": "partial", + "segment": segment, + "text": text, + "tokens": tokens, + } + ) + prev_text = text + if recognizer.is_endpoint(stream): + if text: + self._emit( + { + "kind": "final", + "segment": segment, + "text": text, + "tokens": tokens, + } + ) + recognizer.reset(stream) + segment += 1 + prev_text = "" diff --git a/skills/mc-prompter/scripts/server/main.py b/skills/mc-prompter/scripts/server/main.py index 5878f72..ce20568 100644 --- a/skills/mc-prompter/scripts/server/main.py +++ b/skills/mc-prompter/scripts/server/main.py @@ -3,7 +3,7 @@ # requires-python = ">=3.11" # dependencies = ["aiohttp==3.12.15"] # /// -"""mc-prompter aiohttp server (Phase A: classic teleprompter). +"""mc-prompter aiohttp server (Phase A classic teleprompter + Phase B voice-follow). Launched by run_prompter.py as `python -m server.main` with cwd set to the scripts directory (the PEP 723 header above also allows a direct @@ -14,6 +14,17 @@ python -m server.main --port 8770 --host 127.0.0.1|0.0.0.0 --script --owner-wpm --session-file --token + [--models-dir ] [--asr-provider none|nemotron-streaming] + +ASR (Phase B): when --models-dir is set AND --asr-provider is an implemented +provider (nemotron-streaming), server/asr.py is imported lazily and an +AsrEngine runs on a dedicated worker thread; server/align.py is imported +lazily too and an Aligner is rebuilt from script_ingest.speakable_words(doc) +on every script load or edit. Without --models-dir (or with provider none) +the server is pure tier 1: neither asr nor align nor numpy is ever imported, +and it runs with only aiohttp installed. Provider zipformer-small is a +documented planned lane: selecting it exits 3 at startup, it never pretends +to run. HTTP routes: GET / home page (static/home.html) @@ -28,7 +39,9 @@ "started": ""} GET /api/state {"snapshot": , "doc-version": n, "script": {"path", "title", "word-count"} or null, - "config": {"owner-wpm": N or null}} + "config": {"owner-wpm": N or null}, + "asr": {"available": bool, "provider": str, + "ready": bool}} GET /api/source {"path": , "raw": "", "doc": diff --git a/skills/mc-prompter/scripts/server/static/prompt.html b/skills/mc-prompter/scripts/server/static/prompt.html index 5d7965f..969a526 100644 --- a/skills/mc-prompter/scripts/server/static/prompt.html +++ b/skills/mc-prompter/scripts/server/static/prompt.html @@ -31,6 +31,9 @@
150 wpm manual + + + @@ -130,6 +133,13 @@

Display

+
+

Voice follow

+
+ + +
+

Script markers

@@ -158,8 +168,9 @@

Keyboard shortcuts

ccountdown, then play ssection list dsettings drawer + vtoggle voice follow (this machine only) mouse wheelspeed up / down 2 wpm - click a wordjump the eyeline to that word + click a wordjump the eyeline to that word (re-anchors in voice follow) ?this help escclose panels @@ -168,6 +179,35 @@

Keyboard shortcuts

+ +