From e2f88cec621d7c658e850473af791c875da9692f Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 19:16:43 -0500 Subject: [PATCH 1/6] Add mc-prompter implementation plan (working document, drops before PR) --- PLAN-mc-prompter.md | 300 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 300 insertions(+) create mode 100644 PLAN-mc-prompter.md diff --git a/PLAN-mc-prompter.md b/PLAN-mc-prompter.md new file mode 100644 index 0000000..ecfecc5 --- /dev/null +++ b/PLAN-mc-prompter.md @@ -0,0 +1,300 @@ +# mc-prompter: teleprompter and AI producer, implementation plan + +Working document for the feat-prompter branch. Not intended to merge; it guides the build and gets deleted (or distilled into skill references) before the PR. Written 2026-07-09 from a recon pass over the module and external research on real-time local ASR, teleprompter prior art, and live cueing UX. Revised the same day after an adversarial three-lens review (technical feasibility, module conventions, product); the review's fixes are folded in throughout and marked where they changed a decision. + +## Decisions record (2026-07-09, BMad) + +- LLM runtime for producer mode: Ollama is the default provider of a new `[llm]` config lane, with the standard provider-ladder pattern (local-first, other rungs opt-in later). +- Scope: all three tiers built on this one branch, in phases, with the classic teleprompter working end to end first. +- Skill shape: one service skill, `mc-prompter`, following the mc-audio pattern (no stage, no gate, no project.json state). The `record` stage stays creator-owned. +- Cue channels: visual-only in v1. The kokoro spoken tier is designed here but ships as a fast-follow behind a config flag. +- Landing (approved 2026-07-09): stacked PRs. PR 1 = Phase A plus minimal docs and the help row, PR 2 = Phase B, PR 3 = Phases C+D. All developed on this branch. + +## Product overview: three tiers + +Tier 1 is a classic teleprompter with the full standard feature set, no AI, no model downloads. Tier 2 is voice-follow: the scroll tracks the speaker through a known script using streaming local ASR plus a deterministic alignment algorithm, no LLM. Tier 3 is producer mode: a rundown file (duration, ordered talking points, intro, wrap), a rolling transcript, and a small local LLM that keeps the speaker on track with rate-limited visual cues. + +Each tier is progressive enhancement over the one below it. Tier 1 works on any machine with only the module installed. Tier 2 requires the prompter-lab workspace (ASR models, consent-gated download). Tier 3 additionally requires a running Ollama. + +Everything chosen is cross-platform (Windows, macOS, Linux): browser UI, browser mic capture, sherpa-onnx, silero VAD, kokoro-onnx, Ollama. This is the module's first fully cross-platform lane, and the sherpa-onnx dependency incidentally opens a path to the cross-platform transcription lane already on TODO.md. + +## Architecture + +### Skill shape + +New folder `skills/mc-prompter/`: + +``` +skills/mc-prompter/ + SKILL.md service skill: what it does, how to launch, tier gating + customize.toml [workflow] block + prompter defaults + references/ + cueing.md the cue design contract: escalation ladder, budget, vocabulary, + replan rules, coverage semantics, headphones/AEC dependency + rundown-spec.md the rundown.md file format specification (time math included) + scripts/ + ensure_workspace.py prompter-lab builder (mirrors mc-audio's, consent-gated) + run_prompter.py stdlib launcher: validates workspace, port probe, launches server + server/ the application package (both launch paths run python -m server.main) + __init__.py + main.py aiohttp app: HTTP + WebSocket + static UI; lazy tier-2/3 imports + asr.py sherpa-onnx streaming recognizer + silero VAD (workspace only) + align.py script-follow alignment engine (pure stdlib) + producer.py rundown state machine, replanner, cue engine, Ollama client + rundown.py rundown.md parser (pure stdlib) + script_ingest.py script.md / markdown / plain text ingestion (pure stdlib) + static/ vanilla HTML/JS/CSS, no build step + tests/ self-running test-*.py files (convention below) + fixtures/ +``` + +Conventions honored: SKILL.md frontmatter with exactly `name` and `description`; scripts invoked only via `uv run` with PEP 723 headers; scripts take explicit resolved arguments and do no config discovery; the skill reads only its own folder, `_bmad/scripts/`, and project files; nothing user-specific ships in the module; config keys kebab-case. + +### The prompter-lab workspace + +Heavy dependencies live in a persistent venv at `{engines-path}/prompter-lab`, exactly like audio-lab: + +``` +/.venv/ aiohttp (same pin as the PEP 723 header), numpy, + sherpa-onnx (pinned), soundfile +/models/ ASR models (below); later the kokoro pair +/out/ session artifacts (take logs, session transcripts, script backups) +``` + +`ensure_workspace.py --check` exits 0/4 like mc-audio's; the build asks consent before any download. Tier 1 does not require the workspace at all: `run_prompter.py` launches the server with ASR disabled when the workspace is absent, using `uv run` with aiohttp as a PEP 723 dependency. When the workspace exists, the launcher runs the venv interpreter instead, resolved portably (`.venv/bin/python` on POSIX, `.venv\Scripts\python.exe` on Windows). + +Two-launch-path discipline (review finding): + +- Both paths execute the server identically as `python -m server.main` from the scripts directory, so intra-package imports resolve the same way in both. A test executes main.py under a bare env with only aiohttp and asserts the tier-1 routes come up. +- main.py imports asr.py (and anything touching numpy/soundfile) lazily, only inside the workspace-present branch. align.py, rundown.py, and script_ingest.py stay pure stdlib. +- The aiohttp version is pinned identically in the PEP 723 header and the workspace venv, bumped together. + +Port handling (review finding): default port 8770 (mc-ograf's ephemeral verifier uses 8771). On launch, probe the port; if occupied, query `/health` (which reports session and script identity), then offer kill-and-replace or auto-increment to the next free port. The chosen port is printed and written to a session file the skill reads back. `--port` overrides. The server binds 127.0.0.1 by default; `--lan` opts into 0.0.0.0 for the phone remote and tablet displays, and the docs note the Windows Firewall consent dialog this triggers. The home page shows the LAN URL only when it is actually reachable. + +ASR models downloaded into `models/`: + +- Primary: sherpa-onnx export of nvidia nemotron-speech-streaming-en-0.6b, int8 (the streaming sibling of the parakeet family; NVIDIA Open Model License, commercial use permitted). Chunk setting 560 ms as the default latency point. +- Fallback for low-end hardware: a small streaming zipformer English export (Apache-2.0), selectable via config. +- VAD: silero VAD via sherpa-onnx's built-in VoiceActivityDetector (one dependency covers both). + +### Server and UI + +One aiohttp process serving four pages plus a WebSocket: + +- `/` home: pick a source (project script.md, rundown.md, pasted/loaded text), configure, preflight, launch. Shows the reconciled rundown plan before a show starts. +- `/prompt` the prompter display: fullscreen scroll surface, all tier-1 features, the producer rail when tier 3 is active. +- `/remote` phone-as-remote over LAN: play/pause, speed, jump to marker, next/prev section, and in producer mode the point-list override controls. Control only, no mic. The remote URL embeds a per-session token and is presented as a QR code on the home page, so a random LAN device cannot drive the prompter mid-show. +- `/overlay` OBS browser source: transparent background, renders only the ambient rail and cue cards for live shows. + +Concurrency model (review finding, binding): sherpa-onnx decode is CPU-bound and blocking, so it never runs on the event loop. A dedicated ASR thread consumes a bounded queue of PCM frames and publishes recognition events back via `call_soon_threadsafe`. Ollama ticks run as background tasks with hard timeouts and never hold shared state across an await. WebSocket fan-out never awaits decode or LLM work. The server smoke test asserts `/remote` command round-trip latency stays low while the replay harness saturates the ASR path. + +Audio capture and ownership (review findings, binding): + +- Mic capture happens on the server machine. getUserMedia requires a secure context, which http over LAN is not, so a tablet pointed at `/prompt` is display-only. The WS protocol separates the display role from the audio-producer role: the server grants a capture token to exactly one localhost connection; frames without the current token are rejected and the UI names the mic owner. Mic-on-remote-device would require shipping TLS and is explicitly deferred. +- getUserMedia constraints request `echoCancellation: false, noiseSuppression: false, autoGainControl: false, channelCount: 1`; browser speech processing measurably degrades ASR input. The preflight screen reads back `track.getSettings()` and surfaces what was actually applied, since browsers may ignore constraints. When the kokoro spoken tier ships, headphones-only output is what keeps AEC unnecessary; cueing.md records that dependency. +- The AudioContext is created with `{ sampleRate: 16000 }` so the browser resamples; if the browser refuses the rate, a small resampler in the worklet handles the conversion (naive decimation from 44.1 kHz aliases into the speech band). The preflight screen verifies `context.sampleRate`. The worklet ships ~120 ms PCM16 mono frames over the WebSocket. +- Backpressure has a policy at both ends: the browser checks `ws.bufferedAmount` and drops frames past a threshold; the server frame queue is bounded and drops oldest on overflow while raising an "ASR behind real-time" state on the rail. The replay harness asserts queue-depth behavior. + +State flows back to all connected pages over the same WebSocket (scroll position, VAD state, transcript tail, cue events), so the remote and overlay stay in sync with the prompter display. + +UI is vanilla HTML/JS/CSS with no build step. Display settings (mirror flips, font, colors, margins, eyeline position) persist in localStorage per device, with the `[prompter]` config values as defaults; a beam-splitter rig and the operator's browser keep independent settings without reconfiguration each launch. + +### Tier 1: the standard feature checklist + +Scroll and timing: + +- Smooth continuous scroll, speed as WPM with live +/- adjustment (keyboard, wheel, remote) +- Timed mode: give total duration, speed is continuously re-derived (remaining words over remaining time, recomputed on resume and after any jump), with the timer display showing drift from plan +- Pause/resume (spacebar), jump forward/back, jump to marker, restart +- Countdown before scroll starts +- Elapsed and remaining time, estimated read time from word count at the creator's measured wpm (`[owner] wpm` from the studio config when available) + +Display: + +- Mirror flip horizontal, vertical, and both (beam-splitter rigs) +- Font family/size, text and background colors, margins, line height +- Adjustable eyeline/cue marker (position, style) +- Fullscreen; works on a second monitor or a tablet pointed at the same URL (display-only on remote devices, see audio ownership above) +- Per-device settings persistence (localStorage), config defaults underneath + +Script handling: + +- Markdown and plain text; project `script.md` ingestion (below) +- Inline bracket notes render dimmed and are never matched by voice-follow +- Named markers/sections for jumping +- Edit-in-place from the home page between takes; edits write back to the source file with a timestamped backup copied to the workspace `out/` first, so the prompted text and the pipeline artifact never silently diverge + +Remote: + +- Keyboard shortcuts throughout; bluetooth presenters and USB foot pedals work as keyboard emulators for free +- `/remote` phone page over LAN, session-token URL via QR code + +### Script ingestion (pipeline tie-in) + +`script.md` from a project is directly consumable: plain spoken prose. Ingestion handles the two inline marker types: + +- `[INVENTED]` flags render as a subtle badge, toggleable off +- `[TAKE s-s]` lines render dimmed with a "have it already" badge, since those lines were already spoken well in the interview footage and may not need re-recording; a toggle hides them entirely + +The prompter takes a path argument; the skill resolves it from the project when launched inside the pipeline flow ("record with the teleprompter") or accepts any file standalone. + +### Tier 2: voice-follow alignment engine + +Deterministic, no LLM. The prior art (bounded-window Levenshtein prefix matching, the PromptSmart hold-and-re-anchor behavior) consumed Web-Speech-style utterance partials; a streaming transducer behaves differently, so the contract is adapted for sherpa-onnx output (review finding, binding): + +- Input contract: the engine consumes token deltas since the last partial, not whole hypotheses. The last K tokens (K around 3 to 5) of the hypothesis are held provisional because beam search can revise the tail between partials; the anchor commit lags the hypothesis head by K tokens and absorbs revisions. On endpoint detection (`is_final`), the anchor hard-commits and tail-tracking state resets for the next segment. BPE pieces merge to words during normalization before matching. +- Normalize both script tokens and ASR tokens: lowercase, strip punctuation, expand common number/abbreviation forms at index-build time (a small normalization table; "2026" also indexes as "twenty twenty six") +- Maintain a monotonic anchor (last committed script token). On each delta batch, take a lookahead window from the anchor (window size proportional to utterance length plus a constant, on the order of 2x + 10 tokens) and find the window prefix minimizing Levenshtein distance to the pending recognized words; the best prefix end becomes the provisional anchor +- Silence (VAD) produces no partials, so the scroll holds; ad-libs fail to match and the anchor holds until speech re-matches within the window +- Escape hatches: click/tap any word to re-anchor, arrow keys nudge the anchor, and a paragraph-skip gesture jumps the window when the creator deliberately skips content +- The scroll controller eases toward the anchor position rather than jumping; the anchor leads the eyeline by the measured end-to-end latency times the current speaking rate, so the eyeline sits where the speaker actually is, not where ASR last confirmed +- Match state is visible: matched text subtly tinted behind the eyeline, so trust in the tracker is inspectable + +Latency: the honest budget is end-to-end and includes terms the naive sum misses: capture framing (~120 ms) + ASR chunk emission (560 ms configured) + decode compute (hardware-dependent, grows under OBS load) + alignment (<10 ms) + scroll easing (a deliberate time constant). Realistic eyeline-follows-voice latency is 1 to 2 s depending on hardware. The replay harness measures capture-timestamp-to-anchor-update wall time on each target platform, and the measured number feeds the eyeline lead default. No fixed latency claim ships in docs. + +Preflight (review finding, this is the try-once-never-again defense): before any take, the preflight screen enumerates input devices with a picker persisted per machine, shows a live level meter, reports the applied audio constraints and sample rate, and runs a 10-second "read this sentence" tracking test that demonstrates the match-tint following before a real take starts. If recognition confidence is garbage, it fails loudly and names the device in use. + +The engine is a pure-stdlib module. Its fixture suite is built from recorded real partial sequences (the actual partial/final event stream captured from the model over the replay WAV), not hand-written final transcripts, covering: verbatim read, ad-lib excursion and return, skipped paragraph, number/abbreviation mismatch, repeated-phrase script traps, and tail-revision events. + +### Tier 3: producer mode + +Inputs: a rundown file, the rolling transcript from the same ASR stream, and the show clock. + +Show clock semantics (review finding, binding): the plan's arithmetic never keys off server or page start. An explicit GO LIVE control (on `/prompt` and `/remote`) starts the show clock after any pre-roll, and a plan-hold control freezes elapsed time and the state machine during BRB or technical trouble while VAD and the transcript keep running so context is not lost. Both are fixture-tested scenarios. + +The producer is two cooperating parts: + +- A deterministic state machine (code, not LLM). It tracks elapsed time against per-segment budgets, and it re-plans rather than merely flagging lateness: on every tick, remaining show time is redistributed across uncovered segments proportionally to their original budgets, with wrap-minutes protected as a hard reserve. Green/yellow/red state is always computed against the current re-plan, never the original rundown. When redistribution would push any segment below a feasibility floor, the state machine emits a card-tier CUT suggestion ("DROP: point 4, or 90s each"). This replan behavior is the core producer value (a countdown that only turns red is a nag, not a producer) and is specified in cueing.md as part of the binding contract. +- An LLM tick (Ollama): scheduled adaptively, and at VAD pause events, it receives a compact state block (rundown with per-point coverage, the current re-plan, elapsed vs plan, the last ~60 seconds of transcript) and returns structured JSON: proposed coverage transitions, current-topic guess, an optional suggested cue with tier and text, and a one-line reason. Cheap keyword/fuzzy matching runs continuously between ticks as a first-pass coverage signal the LLM confirms or overrides. + +Coverage semantics (review finding, binding): coverage is sticky and monotonic in the state machine; the LLM may only propose uncovered-to-covered transitions, never reversions, so the rail cannot flicker. "Next" is defined as the first uncovered point in rundown order, which stays well-defined when the creator covers points out of order. The human is the final authority: `/remote` gains producer controls, a tappable point list with mark-covered, skip, and make-current, so one tap mid-show rescues any model misjudgment. + +LLM tick budget (review finding, binding): the tick must never starve the ASR thread. Concretely: + +- Requests use Ollama structured outputs (`format` with a JSON schema), `think: false`, temperature 0, a hard `num_predict` cap, and `keep_alive` so the model stays resident +- The prompt keeps a stable prefix (system + rundown first, rolling transcript last) so Ollama prefix caching skips reprocessing +- Cadence is adaptive: the next tick is scheduled at `max(15 s, 3x last tick wall time)`, and ticks are skipped entirely while the ASR queue depth signals CPU pressure +- Default model is `qwen3:4b` where Ollama reports GPU/Metal offload; on CPU-only machines the producer startup check recommends and falls back to a sub-2B tag (`qwen3:1.7b`). Every request carries a hard timeout; a timed-out tick is dropped, not queued + +The cue engine (code) is the final authority on delivery: it applies the density setting, the one-active-cue rule, per-interval budget, and tier gating. The LLM proposes; the state machine disposes. A wrong suggestion costs nothing because rate limiting, tiering, and coverage stickiness are deterministic. + +Visual cue surface (v1, from the cueing research): + +- Ambient tier, always on: a rail showing current point, next point, and a green/yellow/red segment-time state computed against the re-plan (Toastmasters vocabulary), plus overall show progress. No motion, no reading required +- Card tier: a single quiet card ("NEXT: pricing demo", "STRETCH: 4 min left, 1 point to go", "DROP: point 4, or 90s each"), released at pauses, auto-expiring, never stacked +- Attention tier: the card flashes/enlarges for time-critical states ("WRAP", "2:00 OVER"), the one tier allowed to appear mid-sentence +- Vocabulary is the broadcast lexicon (standby, wrap, stretch, hard wrap, time remaining); `references/cueing.md` is the binding contract + +Free-talk support: a rundown with no script body per point is exactly the "5 ideas in this order plus intro and wrap" show. The intro and wrap can carry full scripted text (prompted via tier 1/2) while the middle segments run producer-only. The segment handoff is explicit UI behavior, not hand-waved: when a scripted segment ends (anchor reaches section end, or manual next-section), `/prompt` switches to a large-type rail view showing the current bullet set; entering the next scripted segment switches back to the scroll surface. + +Spoken tier (designed now, shipped later behind `spoken-cues = false`): short formulaic kokoro phrases only ("thirty seconds", "wrap"), synthesized by a persistent kokoro instance (~150 to 300 ms for a short cue on CPU), released only at pauses, hard requirement that output routes to headphones (this is also what keeps browser AEC unnecessary). Never speech-over-speech except a true emergency tier. The mix-minus principle from IFB practice: the speaker must never hear their own voice back. + +### The rundown artifact + +New file format, specified in `references/rundown-spec.md`. The starter template ships through mc-setup's `assets/` into the studio like tokens and format profiles, so Manny and any skill can read it as a project file without crossing skill-folder boundaries; mc-prompter owns the spec and does any template-based drafting, and Manny routes to it (review finding). + +```markdown +--- +show: "Why local models win" +duration-minutes: 30 +cue-density: normal # hands-off | minimal | normal | chatty +wrap-minutes: 3 +--- + +## Intro (3 min) + +Full scripted intro text here, prompted normally. + +## Point 1: The cost argument (5 min) + +- cloud bills compound, local is capex +- the 4090 anecdote + +## Point 2: Latency (5 min) +... + +## Wrap (3 min) + +Scripted wrap text. +``` + +Time math (review finding, in the spec): per-segment minutes are optional; unbudgeted segments split the remaining time evenly. If explicit minutes exceed duration-minutes, duration-minutes wins and the parser warns at load; the home page shows the reconciled plan before the show starts. The parser accepts exactly `(N min)` and `(Nm)` heading suffixes and rejects anything else with a line-numbered error, because hand-written and Manny-drafted rundowns will produce creative variants on day one. Frontmatter `cue-density` overrides the config value (most specific wins). + +Segments with prose bodies prompt as script; segments with only bullets run producer-only. For pipeline projects the file lives at `{projects-path}//rundown.md`; standalone shows pass any path. This artifact also fills the episode-plan gap the livestream formats already reference (the 1.0.x per-episode stream-pack fast-follow needs the same file), and the planned mc-research skill becomes its natural upstream. + +### Config: new studio sub-tables + +Two new tables in `[modules.manticore]`, seeded from mc-setup's `[defaults]`: + +```toml +[defaults.prompter] +workspace = "prompter-lab" # resolved {engines-path}/{prompter.workspace} +asr-provider = "nemotron-streaming" # nemotron-streaming (default) | zipformer-small | none +cue-density = "normal" # hands-off | minimal | normal | chatty +spoken-cues = false # kokoro tier, fast-follow +port = 8770 + +[defaults.llm] +provider = "ollama" # the only implemented rung; others planned, opt-in +model = "qwen3:4b" # any ollama tag; producer falls back to a sub-2B tag on CPU-only +endpoint = "http://localhost:11434" +api-key-env = "" # stays empty for local lanes, pattern-consistent +``` + +mc-setup changes: + +- Add both tables to `[defaults]`, a short optional interview step (offer the teleprompter, ask about producer mode and Ollama only if wanted), and both table names to the step 8 write list +- Migration (review finding, important): do NOT add these tables to the step 1a 0.x classifier list; that list defines what makes a studio 0.x, and adding 1.1 tables to it would mislabel every current 1.0.x studio as 0.x and run the full migration flow on it. Instead add a separate backfill rule alongside it: a config that has the 1.0 tables but is missing `[prompter]` or `[llm]` is a current studio predating the teleprompter; backfill both tables surgically from `[defaults]` and offer the optional prompter interview +- check_deps.py gains an optional `ollama` row (producer mode only) with install pointers +- PIPELINE.md's conventions line mentions the new sub-tables + +The `[llm]` lane follows the enforcement pattern: the producer accepts only `provider = "ollama"` and exits 3 for anything else, so no planned lane ever pretends to work. + +## Module integration checklist + +- `skills/module-help.csv`: one new row with the full 13-column schema (module, skill, display-name, menu-code, description, action, args, phase, preceded-by, followed-by, required, output-location, outputs); phase `anytime`, preceded-by and followed-by left empty per the mc-audio/mc-ograf service-skill precedent +- `.claude-plugin/marketplace.json`: add `./skills/mc-prompter` to the skills array (version bump at release, not in this PR) +- `skills/mc-pipeline/PIPELINE.md`: record stage row gains a sentence noting mc-prompter as the optional tool for the creator-owned record stage; no stage table change, no gate change +- mc-script SKILL.md step 6: the "now record" handoff names the teleprompter option (naming only, no cross-skill read) +- mc-agent (Manny) capabilities: route "teleprompter", "prompt me", "run my show", "producer mode" to mc-prompter; for rundown drafting Manny routes to mc-prompter or works from the studio-installed template, never reads mc-prompter's folder +- docs/user-guide.md: new section; README skill table row; TODO.md: add the spoken-cue fast-follow and the sherpa-onnx transcription-lane opportunity + +## Phasing + +- Phase A, classic teleprompter: workspace-less launch path, aiohttp server with the concurrency skeleton, `/prompt` with the full tier-1 checklist, `/remote` with session token, script.md ingestion with marker handling, home page, settings persistence, port handling. Verifiable end to end with zero models +- Phase B, voice-follow: ensure_workspace.py, browser audio path with the constraint/ownership/backpressure rules, sherpa-onnx + VAD integration on the ASR thread, align.py with the transducer contract and its recorded-partials fixture suite, preflight screen with device picker and tracking test, tracking UX (hold, re-anchor, click-to-anchor, match tinting, eyeline lead) +- Phase C, producer: rundown parser + spec, producer state machine with replanner and cue engine plus time-warped fixtures (running long, running short, out-of-order coverage, GO LIVE / hold), Ollama tick with the budget rules, ambient rail + cards on `/prompt` and `/overlay`, `/remote` producer controls, `[prompter]`/`[llm]` config plumbing, mc-setup + check_deps integration +- Phase D, finish: docs, help catalog row, marketplace entry, Manny routing, take-log hook, cross-platform smoke instructions, changelog + +Landing (approved 2026-07-09): develop everything on this branch, land as stacked PRs: PR 1 = Phase A plus minimal docs and the help row (a complete, useful teleprompter on its own), PR 2 = Phase B, PR 3 = Phases C+D. Each is independently green and valuable, and review feedback on the foundations arrives before the producer is built on top of them. + +## Testing and verification + +Test convention (review finding, matches the quality gate CI): the CI discovers `skills/*/scripts/tests/test-*.py` (hyphenated) and runs each file directly via `uv run`; there is no pytest. Every test file carries its own PEP 723 header, uses stdlib unittest with `unittest.main()`, and runs with no models, no network, no downloads. + +- `test-align.py`: the transducer-contract fixtures (recorded partial sequences committed as small JSON fixtures; the five failure-mode cases plus tail-revision events) +- `test-producer.py`: replanner and cue engine under time-warped scenarios (running long with replan and CUT suggestion, running short with STRETCH, point skipped, out-of-order coverage, GO LIVE and hold semantics, budget/ladder enforcement, coverage stickiness) +- `test-rundown.py`: parser, time-math reconciliation, heading-suffix rejection with line numbers +- `test-script_ingest.py`: marker handling, bracket-note exclusion +- `test-server.py`: aiohttp smoke via its own PEP 723 aiohttp dependency; routes up under the bare tier-1 env; WS state fan-out; remote-latency-under-ASR-load assertion lives here but auto-skips (exit 0 with a message) when the workspace is absent, so CI never needs models +- One small recorded fixture WAV (a few seconds, 16 kHz mono) committed under `scripts/tests/fixtures/`; nothing generates audio in-test +- The ASR replay harness (feed the fixture WAV through the same code path the WebSocket uses, measure capture-to-anchor wall time) is a documented manual check per platform, not a CI assertion; CI has no workspace and wall-clock assertions are flaky by construction +- Manual test matrix documented in the skill: macOS (reference), Windows, Linux; beam-splitter mirror check; phone remote; OBS overlay; device-picker preflight + +## Risks and mitigations + +- Nemotron streaming model quality/latency on low-end CPUs: zipformer-small fallback behind `asr-provider`, and tier 1 works with `none` +- ASR token timestamps are start-only in sherpa-onnx: alignment keys on token text order, not timestamps, so this costs nothing +- CPU contention between ASR, the LLM tick, and OBS on one machine: the tick budget rules above (adaptive cadence, pressure-skip, small-model fallback, timeouts) plus the bounded-queue drop policy keep voice-follow latency from compounding; the rail surfaces "ASR behind real-time" instead of silently lagging +- Ollama absent or model not pulled: producer mode degrades to the deterministic rail (timing and replan cues still work, coverage judgments off); the UI says exactly what is missing +- Repeated phrases in scripts confusing alignment: bounded window plus monotonic anchor limits damage; fixture-tested +- Browser mic pitfalls (wrong device, processing constraints ignored, wrong sample rate): the preflight screen with device picker, level meter, applied-settings readback, and the 10-second tracking test catches all of these before the first real take +- Scope creep: the kokoro spoken tier and the mc-cut take-log consumption are explicitly out of v1 + +## Explicitly out of scope (future hooks) + +- Spoken kokoro cue tier (designed, config key ships false; fast-follow) +- mc-cut consuming the take log (`out/take-log.json`, script positions and timestamps per take) to pre-anchor cut plans +- Chat/vision inputs to the producer (reading live chat is a natural producer input later) +- Cloud LLM rungs for `[llm]`, paid ASR rungs +- TLS for mic capture on remote devices (tablet-as-mic) +- Cross-platform batch transcription lane via sherpa-onnx parakeet-tdt offline export (separate TODO item this branch makes cheaper) From 318ce5dbad17948e9aaed06b0c30b1ebeddce3eb Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 19:16:43 -0500 Subject: [PATCH 2/6] Add mc-prompter script ingestion: sections, TAKE/INVENTED/note markers --- .../scripts/server/script_ingest.py | 301 +++++++++++++++++ .../scripts/tests/fixtures/sample-script.md | 46 +++ .../scripts/tests/test-script_ingest.py | 316 ++++++++++++++++++ 3 files changed, 663 insertions(+) create mode 100644 skills/mc-prompter/scripts/server/script_ingest.py create mode 100644 skills/mc-prompter/scripts/tests/fixtures/sample-script.md create mode 100644 skills/mc-prompter/scripts/tests/test-script_ingest.py diff --git a/skills/mc-prompter/scripts/server/script_ingest.py b/skills/mc-prompter/scripts/server/script_ingest.py new file mode 100644 index 0000000..e730ddf --- /dev/null +++ b/skills/mc-prompter/scripts/server/script_ingest.py @@ -0,0 +1,301 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# /// +"""Script ingestion for mc-prompter (Phase A). + +Usage: + uv run {skill-root}/scripts/server/script_ingest.py \ + [--fmt markdown|plain] + +Contract: + input a script file: markdown (default) or plain text. The pipeline's + script.md is directly consumable; any standalone file works too. + output the script model as JSON on stdout: + { + "title": "first h1 text or null", + "word-count": 1234, + "sections": [ + {"id": "s0", "heading": "Intro or null", "level": 2, + "blocks": [ + {"type": "para", "runs": [ + {"text": "Speakable text.", "flags": []}, + {"text": "Flagged sentence.", "flags": ["invented"]} + ]}, + {"type": "note", "text": "pause here"}, + {"type": "take", "source": "int1", "start": 42.0, + "end": 51.5, "runs": [...]} + ]} + ] + } + markers three inline marker types are recognized: + [TAKE s-s] the whole paragraph becomes + a take block; the marker is stripped from the display text + and source/start/end are parsed (float seconds). A + malformed TAKE marker degrades to note handling and never + crashes ingestion. + [INVENTED] flags the sentence it immediately follows; the + marker is stripped and that sentence's run gets the flag + "invented". + [any other bracketed text] becomes a note block: in place if + the bracket is a whole line on its own, otherwise the note + text is extracted from the paragraph and emitted as note + blocks immediately after it. + sections markdown sections split on headings at any level; text before + the first heading is a section with heading null (level 0). The + first h1 becomes the document title and is not a section by + itself; its content flows into the current heading-null section. + Section ids are "s0", "s1", ... in document order. + fmt "markdown" (default) or "plain". Plain skips heading parsing: + the whole document is one heading-null section of blank-line + separated paragraphs, with the same marker rules. + counting word-count counts only speakable words (whitespace-split words + of para and take runs); note text is never counted. + +Sentence boundary rule (for [INVENTED]): the flagged sentence ends at the +marker; a sentence terminator (".", "!", "?") immediately before the marker +(ignoring whitespace) belongs to the flagged sentence, since the marker flags +the sentence it follows. The sentence starts just after the previous +terminator, or at the paragraph start when none exists. This is a +deliberately simple, deterministic rule; abbreviations and decimal points are +treated as sentence boundaries like any other period. + +Exit codes: 0 ok, 2 file missing or unreadable, 2 bad --fmt (argparse usage). +Pure stdlib; importable by server/main.py via `from . import script_ingest` +or as a plain module from the scripts/server directory. +""" + +import argparse +import json +import re +import sys +from pathlib import Path + +TAKE_RE = re.compile(r"\[TAKE\s+(\S+)\s+(\d+(?:\.\d+)?)s-(\d+(?:\.\d+)?)s\]") +INVENTED = "[INVENTED]" +INVENTED_RE = re.compile(r"\[INVENTED\]") +BRACKET_RE = re.compile(r"\[([^\[\]]+)\]") +HEADING_RE = re.compile(r"^(#{1,6})\s+(.*?)\s*#*\s*$") +WHOLE_LINE_NOTE_RE = re.compile(r"^\[([^\[\]]+)\]$") +EMPTY_BRACKET_RE = re.compile(r"\[\s*\]") +SENTENCE_ENDS = ".!?" + + +def _clean(text): + """Collapse whitespace runs to single spaces and strip the ends.""" + return re.sub(r"\s+", " ", text).strip() + + +def _build_runs(text): + """Split paragraph text into runs, attributing [INVENTED] flags. + + Each [INVENTED] marker flags the sentence it immediately follows. A + sentence terminator (".", "!", "?") sitting right before the marker + (ignoring whitespace) is part of the flagged sentence; the sentence + starts after the previous terminator or at the start of the unconsumed + text. Everything else becomes unflagged runs. Markers are stripped. + Empty runs are dropped. + """ + runs = [] + pos = 0 + for m in INVENTED_RE.finditer(text): + segment = text[pos:m.start()] + stripped = segment.rstrip() + if stripped and stripped[-1] in SENTENCE_ENDS: + search = stripped[:-1] + else: + search = stripped + boundary = max(search.rfind(c) for c in SENTENCE_ENDS) + if boundary >= 0: + before = _clean(segment[:boundary + 1]) + flagged = _clean(segment[boundary + 1:]) + else: + before = "" + flagged = _clean(segment) + if before: + runs.append({"text": before, "flags": []}) + if flagged: + runs.append({"text": flagged, "flags": ["invented"]}) + pos = m.end() + tail = _clean(text[pos:]) + if tail: + runs.append({"text": tail, "flags": []}) + return runs + + +def _paragraph_blocks(para_text): + """Turn one paragraph's raw text into an ordered list of blocks. + + Order of operations: whole-line bracket notes are pulled out first (they + become note blocks; a paragraph that is only note lines yields only + notes), then a well-formed TAKE marker is stripped and remembered, then + remaining non-INVENTED brackets are extracted as trailing note blocks, + and finally [INVENTED] attribution builds the runs. + """ + notes = [] + kept_lines = [] + for line in para_text.splitlines(): + m = WHOLE_LINE_NOTE_RE.match(line.strip()) + if m: + content = m.group(1) + if content == INVENTED[1:-1] or TAKE_RE.fullmatch(line.strip()): + kept_lines.append(line) + else: + notes.append(content.strip()) + else: + kept_lines.append(line) + text = "\n".join(kept_lines) + + take = None + tm = TAKE_RE.search(text) + if tm: + take = { + "source": tm.group(1), + "start": float(tm.group(2)), + "end": float(tm.group(3)), + } + text = text[:tm.start()] + text[tm.end():] + + def _extract_note(m): + content = m.group(1) + if content == INVENTED[1:-1]: + return m.group(0) + notes.append(content.strip()) + return " " + + # Repeat until stable so nested brackets are fully extracted: each pass + # handles the innermost pairs and exposes the next level. Empty bracket + # pairs (literal or left over once inner content is consumed) are dropped + # so no bracket residue reaches speakable text. Bounded: every pass that + # changes the text removes characters. + while True: + new_text = BRACKET_RE.sub(_extract_note, text) + new_text = EMPTY_BRACKET_RE.sub(" ", new_text) + if new_text == text: + break + text = new_text + + blocks = [] + runs = _build_runs(text) + if runs: + if take: + blocks.append({"type": "take", **take, "runs": runs}) + else: + blocks.append({"type": "para", "runs": runs}) + for note in notes: + blocks.append({"type": "note", "text": note}) + return blocks + + +def _split_paragraphs(lines): + """Split a list of lines into blank-line separated paragraph strings.""" + paras = [] + buf = [] + for line in lines: + if line.strip(): + buf.append(line) + elif buf: + paras.append("\n".join(buf)) + buf = [] + if buf: + paras.append("\n".join(buf)) + return paras + + +def _word_count(sections): + total = 0 + for section in sections: + for block in section["blocks"]: + if block["type"] in ("para", "take"): + for run in block["runs"]: + total += len(run["text"].split()) + return total + + +def ingest(text, fmt="markdown"): + """Ingest script text and return the script model dict. + + fmt "markdown" splits sections on headings; the first h1 becomes title. + fmt "plain" produces a single heading-null section. Marker rules are + identical in both formats. Raises ValueError on an unknown fmt. + """ + if fmt not in ("markdown", "plain"): + raise ValueError(f"unknown fmt: {fmt}") + + # Defensive BOM strip: pasted text (POST /api/source) may carry a leading + # U+FEFF that would otherwise break heading detection on the first line. + text = text.lstrip("") + + title = None + sections = [] + current = None + + def _ensure_section(heading=None, level=0): + nonlocal current + current = {"heading": heading, "level": level, + "lines": [], "blocks": []} + sections.append(current) + + if fmt == "plain": + _ensure_section() + current["lines"] = text.splitlines() + else: + for line in text.splitlines(): + hm = HEADING_RE.match(line) + if hm: + level = len(hm.group(1)) + heading_text = hm.group(2).strip() + if level == 1 and title is None: + title = heading_text + continue + _ensure_section(heading_text, level) + else: + if current is None: + if not line.strip(): + continue + _ensure_section() + current["lines"].append(line) + + for section in sections: + for para in _split_paragraphs(section.pop("lines")): + section["blocks"].extend(_paragraph_blocks(para)) + + out_sections = [ + {"id": f"s{i}", "heading": s["heading"], "level": s["level"], + "blocks": s["blocks"]} + for i, s in enumerate(sections) + ] + return { + "title": title, + "word-count": _word_count(out_sections), + "sections": out_sections, + } + + +def main(argv=None): + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("file", help="path to the script file") + parser.add_argument("--fmt", default="markdown", + choices=("markdown", "plain"), + help="input format (default: markdown)") + args = parser.parse_args(argv) + + path = Path(args.file) + if not path.is_file(): + print(f"error: script not found: {path}", file=sys.stderr) + return 2 + try: + text = path.read_text(encoding="utf-8-sig") + except (OSError, UnicodeDecodeError) as exc: + print(f"error: cannot read {path}: {exc}", file=sys.stderr) + return 2 + + doc = ingest(text, fmt=args.fmt) + print(json.dumps(doc, indent=2, ensure_ascii=False)) + return 0 + + +if __name__ == "__main__": + code = main() + if code: + sys.exit(code) diff --git a/skills/mc-prompter/scripts/tests/fixtures/sample-script.md b/skills/mc-prompter/scripts/tests/fixtures/sample-script.md new file mode 100644 index 0000000..b02fa5b --- /dev/null +++ b/skills/mc-prompter/scripts/tests/fixtures/sample-script.md @@ -0,0 +1,46 @@ +# Why Local Tools Win + +A quick note before we start. This opening line sits above the first section +heading on purpose, so ingestion has a preamble to deal with. + +## Cold open + +Hey everyone, welcome back to the workshop. Today we are digging into why +running your tools locally beats renting them by the minute. + +[smile, beat, then lean in] + +Stick around to the end, because the last demo surprised even me. + +## The cost argument + +Cloud bills compound quietly. A local rig is a one-time cost that keeps +paying you back. I ran the numbers last month and the break-even point was +under ninety days. [INVENTED] That figure assumes you already own a decent +GPU, which most of you watching probably do. + +[TAKE int1 12.0s-19.5s] +The first time I saw the invoice I genuinely thought it was a typo, and +that was the moment I started moving everything back on-prem. + +## Latency + +Round trips to a data center add up. When the model sits on your own +machine, the answer starts streaming before your finger leaves the key. +Local inference latency has dropped forty percent this year. [INVENTED] + +That speed changes how you work [pause for the b-roll insert] because you +stop batching your thoughts and start thinking out loud. + +[TAKE int1 41.25s-52.0s] +My editor asked me why the demo footage looked sped up, and I had to +explain that no, that is just what local feels like. + +## Wrap + +So that is the case: your money, your latency, and your data all point the +same direction. Every major toolchain will ship a local-first mode within a +year. [INVENTED] Try one local tool this week and tell me how it goes in +the comments. + +Catch you in the next one. diff --git a/skills/mc-prompter/scripts/tests/test-script_ingest.py b/skills/mc-prompter/scripts/tests/test-script_ingest.py new file mode 100644 index 0000000..7922186 --- /dev/null +++ b/skills/mc-prompter/scripts/tests/test-script_ingest.py @@ -0,0 +1,316 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# /// +"""Tests for script_ingest.py (mc-prompter Phase A). + +Run directly: + uv run skills/mc-prompter/scripts/tests/test-script_ingest.py + +Pure stdlib unittest; no network, no models, no downloads. Uses the committed +fixture at fixtures/sample-script.md plus small inline documents. +""" + +import json +import subprocess +import sys +import tempfile +import unittest +from pathlib import Path + +TESTS_DIR = Path(__file__).resolve().parent +sys.path.insert(0, str(TESTS_DIR.parent / "server")) + +from script_ingest import ingest # noqa: E402 + +FIXTURE = TESTS_DIR / "fixtures" / "sample-script.md" + + +def blocks_of(doc, section_index): + return doc["sections"][section_index]["blocks"] + + +def speakable_text(doc): + parts = [] + for section in doc["sections"]: + for block in section["blocks"]: + if block["type"] in ("para", "take"): + for run in block["runs"]: + parts.append(run["text"]) + return " ".join(parts) + + +class TestSectioning(unittest.TestCase): + + def test_title_from_first_h1(self): + doc = ingest("# My Title\n\nHello there.\n") + self.assertEqual(doc["title"], "My Title") + + def test_first_h1_is_not_a_section(self): + doc = ingest("# My Title\n\nHello there.\n\n## One\n\nBody.\n") + headings = [s["heading"] for s in doc["sections"]] + self.assertNotIn("My Title", headings) + + def test_h1_content_lands_in_null_heading_section(self): + doc = ingest("# My Title\n\nHello there.\n\n## One\n\nBody.\n") + first = doc["sections"][0] + self.assertIsNone(first["heading"]) + self.assertEqual(first["level"], 0) + self.assertEqual(first["blocks"][0]["runs"][0]["text"], + "Hello there.") + + def test_preamble_before_any_heading(self): + doc = ingest("Lead-in text.\n\n## One\n\nBody.\n") + self.assertIsNone(doc["title"]) + self.assertIsNone(doc["sections"][0]["heading"]) + self.assertEqual(doc["sections"][1]["heading"], "One") + + def test_section_ids_are_sequential(self): + doc = ingest("Pre.\n\n## A\n\nx.\n\n## B\n\ny.\n\n### C\n\nz.\n") + self.assertEqual([s["id"] for s in doc["sections"]], + ["s0", "s1", "s2", "s3"]) + self.assertEqual(doc["sections"][3]["level"], 3) + + def test_heading_levels_recorded(self): + doc = ingest("## Two\n\na.\n\n### Three\n\nb.\n") + self.assertEqual(doc["sections"][0]["level"], 2) + self.assertEqual(doc["sections"][1]["level"], 3) + + def test_second_h1_is_a_normal_section(self): + doc = ingest("# Title\n\n# Another\n\nBody.\n") + self.assertEqual(doc["title"], "Title") + self.assertEqual(doc["sections"][0]["heading"], "Another") + self.assertEqual(doc["sections"][0]["level"], 1) + + +class TestTakeParsing(unittest.TestCase): + + def test_take_paragraph_parsed(self): + doc = ingest("[TAKE int1 12.0s-19.5s]\nAlready recorded line.\n") + block = blocks_of(doc, 0)[0] + self.assertEqual(block["type"], "take") + self.assertEqual(block["source"], "int1") + self.assertEqual(block["start"], 12.0) + self.assertEqual(block["end"], 19.5) + self.assertEqual(block["runs"][0]["text"], "Already recorded line.") + + def test_take_marker_inline_mid_paragraph(self): + doc = ingest("Some text [TAKE cam2 5s-9s] more text.\n") + block = blocks_of(doc, 0)[0] + self.assertEqual(block["type"], "take") + self.assertEqual(block["source"], "cam2") + self.assertEqual(block["start"], 5.0) + self.assertEqual(block["end"], 9.0) + self.assertEqual(block["runs"][0]["text"], "Some text more text.") + + def test_malformed_take_degrades_to_note(self): + doc = ingest("[TAKE int1]\nStill speakable text.\n") + blocks = blocks_of(doc, 0) + types = [b["type"] for b in blocks] + self.assertIn("para", types) + self.assertIn("note", types) + note = next(b for b in blocks if b["type"] == "note") + self.assertEqual(note["text"], "TAKE int1") + + def test_malformed_take_missing_s_suffix(self): + doc = ingest("Text here. [TAKE int1 12.0-19.5]\n") + blocks = blocks_of(doc, 0) + self.assertEqual(blocks[0]["type"], "para") + self.assertEqual(blocks[1]["type"], "note") + + +class TestInventedFlag(unittest.TestCase): + + def test_flags_the_sentence_it_follows(self): + doc = ingest("First sentence. Second sentence. [INVENTED] Third.\n") + runs = blocks_of(doc, 0)[0]["runs"] + self.assertEqual(runs[0]["text"], "First sentence.") + self.assertEqual(runs[0]["flags"], []) + self.assertEqual(runs[1]["text"], "Second sentence.") + self.assertEqual(runs[1]["flags"], ["invented"]) + self.assertEqual(runs[2]["text"], "Third.") + self.assertEqual(runs[2]["flags"], []) + + def test_flags_from_paragraph_start(self): + doc = ingest("Only sentence with no terminator [INVENTED] tail.\n") + runs = blocks_of(doc, 0)[0]["runs"] + self.assertEqual(runs[0]["flags"], ["invented"]) + self.assertEqual(runs[0]["text"], + "Only sentence with no terminator") + self.assertEqual(runs[1]["text"], "tail.") + + def test_marker_mid_sentence_flags_partial_span(self): + doc = ingest("Solid claim. Shaky number [INVENTED] and onward.\n") + runs = blocks_of(doc, 0)[0]["runs"] + self.assertEqual(runs[1]["text"], "Shaky number") + self.assertEqual(runs[1]["flags"], ["invented"]) + + def test_marker_not_leaked_into_text(self): + doc = ingest("A claim. [INVENTED] More text.\n") + self.assertNotIn("INVENTED", speakable_text(doc)) + + +class TestNotes(unittest.TestCase): + + def test_whole_line_note_becomes_note_block(self): + doc = ingest("[pause here]\n") + block = blocks_of(doc, 0)[0] + self.assertEqual(block["type"], "note") + self.assertEqual(block["text"], "pause here") + + def test_inline_note_extracted_after_paragraph(self): + doc = ingest("Keep talking [look at camera] without stopping.\n") + blocks = blocks_of(doc, 0) + self.assertEqual(blocks[0]["type"], "para") + self.assertEqual(blocks[0]["runs"][0]["text"], + "Keep talking without stopping.") + self.assertEqual(blocks[1]["type"], "note") + self.assertEqual(blocks[1]["text"], "look at camera") + + def test_note_text_excluded_from_word_count(self): + with_note = ingest("One two three. [a very long stage direction]\n") + without = ingest("One two three.\n") + self.assertEqual(with_note["word-count"], 3) + self.assertEqual(with_note["word-count"], without["word-count"]) + + +class TestWordCount(unittest.TestCase): + + def test_counts_para_and_take_runs_only(self): + doc = ingest( + "## S\n\nFour words right here.\n\n" + "[TAKE int1 1.0s-2.0s]\nThree more words.\n\n" + "[note words never counted at all]\n" + ) + self.assertEqual(doc["word-count"], 7) + + def test_invented_runs_are_counted(self): + doc = ingest("Two words. [INVENTED] Two more.\n") + self.assertEqual(doc["word-count"], 4) + + +class TestPlainFormat(unittest.TestCase): + + def test_no_heading_parsing(self): + doc = ingest("# not a title\n\nBody text.\n", fmt="plain") + self.assertIsNone(doc["title"]) + self.assertEqual(len(doc["sections"]), 1) + self.assertIsNone(doc["sections"][0]["heading"]) + self.assertIn("# not a title", speakable_text(doc)) + + def test_marker_rules_still_apply(self): + doc = ingest("Spoken bit [aside] here.\n\n[TAKE t1 0.5s-2.5s]\nDone.\n", + fmt="plain") + blocks = blocks_of(doc, 0) + types = [b["type"] for b in blocks] + self.assertEqual(types, ["para", "note", "take"]) + + def test_unknown_fmt_raises(self): + with self.assertRaises(ValueError): + ingest("text", fmt="html") + + +class TestFixture(unittest.TestCase): + + def setUp(self): + self.doc = ingest(FIXTURE.read_text(encoding="utf-8")) + + def test_title_and_sections(self): + self.assertEqual(self.doc["title"], "Why Local Tools Win") + headings = [s["heading"] for s in self.doc["sections"]] + self.assertEqual(headings, [None, "Cold open", "The cost argument", + "Latency", "Wrap"]) + + def test_two_takes(self): + takes = [b for s in self.doc["sections"] for b in s["blocks"] + if b["type"] == "take"] + self.assertEqual(len(takes), 2) + self.assertEqual(takes[0]["source"], "int1") + self.assertEqual(takes[0]["start"], 12.0) + self.assertEqual(takes[0]["end"], 19.5) + self.assertEqual(takes[1]["start"], 41.25) + self.assertEqual(takes[1]["end"], 52.0) + + def test_three_invented_flags(self): + flagged = [r for s in self.doc["sections"] for b in s["blocks"] + if b["type"] in ("para", "take") for r in b["runs"] + if "invented" in r["flags"]] + self.assertEqual(len(flagged), 3) + + def test_invented_after_terminator_flags_previous_sentence(self): + latency = next(s for s in self.doc["sections"] + if s["heading"] == "Latency") + flagged = [r for b in latency["blocks"] if b["type"] == "para" + for r in b["runs"] if "invented" in r["flags"]] + self.assertEqual( + flagged[0]["text"], + "Local inference latency has dropped forty percent this year.") + + def test_notes_present_and_not_counted(self): + notes = [b for s in self.doc["sections"] for b in s["blocks"] + if b["type"] == "note"] + self.assertGreaterEqual(len(notes), 2) + self.assertNotIn("smile, beat", speakable_text(self.doc)) + self.assertGreater(self.doc["word-count"], 100) + + def test_no_markers_leak_into_display_text(self): + text = speakable_text(self.doc) + self.assertNotIn("[", text) + self.assertNotIn("TAKE", text) + self.assertNotIn("INVENTED", text) + + +class TestRobustness(unittest.TestCase): + + def test_bom_prefixed_title_parses(self): + doc = ingest("# My Title\n\nHello there.\n") + self.assertEqual(doc["title"], "My Title") + self.assertNotIn("#", speakable_text(doc)) + self.assertEqual(doc["word-count"], 2) + + def test_bom_file_read_via_main(self): + with tempfile.TemporaryDirectory() as tmp: + path = Path(tmp) / "bom.md" + path.write_bytes("# My Title\n\nHello there.\n".encode("utf-8-sig")) + proc = subprocess.run( + [sys.executable, str(TESTS_DIR.parent / "server" + / "script_ingest.py"), str(path)], + capture_output=True, text=True) + self.assertEqual(proc.returncode, 0, proc.stderr) + doc = json.loads(proc.stdout) + self.assertEqual(doc["title"], "My Title") + + def test_non_utf8_file_exits_2_with_message(self): + with tempfile.TemporaryDirectory() as tmp: + path = Path(tmp) / "cp1252.md" + path.write_bytes("# Title\n\nCurly ’quote’.\n" + .encode("cp1252")) + proc = subprocess.run( + [sys.executable, str(TESTS_DIR.parent / "server" + / "script_ingest.py"), str(path)], + capture_output=True, text=True) + self.assertEqual(proc.returncode, 2) + self.assertIn("error: cannot read", proc.stderr) + + def test_nested_brackets_fully_extracted(self): + doc = ingest("Talk here [note [nested] stuff] more.\n") + blocks = blocks_of(doc, 0) + notes = [b["text"] for b in blocks if b["type"] == "note"] + self.assertIn("nested", notes) + self.assertTrue(any("note" in n and "stuff" in n for n in notes)) + text = speakable_text(doc) + self.assertNotIn("[", text) + self.assertNotIn("]", text) + self.assertEqual(text, "Talk here more.") + + def test_empty_brackets_never_leak(self): + doc = ingest("Before [] after.\n\nAlso [[]] here.\n") + text = speakable_text(doc) + self.assertNotIn("[", text) + self.assertNotIn("]", text) + self.assertIn("Before after.", text) + self.assertIn("Also here.", text) + + +if __name__ == "__main__": + unittest.main() From 5bb3de2a19f761826fcdbc670bec8fdf0e05840c Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 19:16:43 -0500 Subject: [PATCH 3/6] Add mc-prompter server and launcher: aiohttp app, WS hub, session tokens --- skills/mc-prompter/scripts/run_prompter.py | 337 +++++++++++ skills/mc-prompter/scripts/server/__init__.py | 7 + skills/mc-prompter/scripts/server/main.py | 542 ++++++++++++++++++ .../scripts/tests/test-run_prompter.py | 231 ++++++++ .../mc-prompter/scripts/tests/test-server.py | 374 ++++++++++++ 5 files changed, 1491 insertions(+) create mode 100644 skills/mc-prompter/scripts/run_prompter.py create mode 100644 skills/mc-prompter/scripts/server/__init__.py create mode 100644 skills/mc-prompter/scripts/server/main.py create mode 100644 skills/mc-prompter/scripts/tests/test-run_prompter.py create mode 100644 skills/mc-prompter/scripts/tests/test-server.py diff --git a/skills/mc-prompter/scripts/run_prompter.py b/skills/mc-prompter/scripts/run_prompter.py new file mode 100644 index 0000000..424ca2c --- /dev/null +++ b/skills/mc-prompter/scripts/run_prompter.py @@ -0,0 +1,337 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# /// +"""Launch the mc-prompter teleprompter server (Phase A). + +Pure stdlib launcher. Probes the port, generates the session token, spawns +the aiohttp server as a child process, confirms it is up, writes the session +file, prints the URLs, and waits on the child (Ctrl-C terminates it). + +Usage: + uv run {skill-root}/scripts/run_prompter.py --script + [--port 8770] [--lan] [--owner-wpm 150] [--no-open] + +Behavior: + port default 8770. If busy, /health on 127.0.0.1 is queried (1 s + timeout); when it answers app == "mc-prompter" the running + session info is printed. Without an explicit --port the + launcher auto-increments to the next free port; with an + explicit --port it exits 5 with guidance. + spawn `uv run --with aiohttp==3.12.15 python -m server.main ` + with cwd set to this scripts directory so the server package + resolves. Server args are all explicit: --port --host --script + --owner-wpm --session-file --token. + session token = secrets.token_urlsafe(16). After /health confirms the + server is up, /mc-prompter/session-.json is + written: {"port": N, "pid": ..., "token": "...", + "script": "...", "started": ""}. The file is removed + when the launcher exits. + --lan binds 0.0.0.0 and prints the LAN URL (best-effort local IP); + expect a Windows Firewall consent dialog on Windows. Without + it the server binds 127.0.0.1 and only localhost URLs print. + --no-open skip opening the browser at the home page. + +Exit codes: 0 ok, 1 server failed to start, 2 usage, 4 script path missing +or unreadable, 5 port conflict with an explicit --port. +""" + +import argparse +import contextlib +import datetime +import json +import os +import secrets +import signal +import socket +import subprocess +import sys +import tempfile +import time +import urllib.error +import urllib.request +import webbrowser +from pathlib import Path + +SCRIPTS_DIR = Path(__file__).resolve().parent +DEFAULT_PORT = 8770 +AIOHTTP_PIN = "aiohttp==3.12.15" +HEALTH_TIMEOUT = 1.0 +# Generous cap: a cold uv cache must resolve, download, and build aiohttp +# before the server can even start, which can take minutes on a slow +# network. wait_for_health fails immediately if the child exits. +STARTUP_TIMEOUT = 300.0 +PORT_SCAN_LIMIT = 20 + + +class PortConflict(Exception): + """Requested port is busy and --port was explicit (exit 5).""" + + def __init__(self, port, info=None): + self.port = port + self.info = info + super().__init__(f"port {port} is busy") + + +def port_is_free(port, host="127.0.0.1"): + """True when the port can be bound on the given host.""" + with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as sock: + sock.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) + try: + sock.bind((host, port)) + return True + except OSError: + return False + + +def health_identity(port, timeout=HEALTH_TIMEOUT): + """GET /health on loopback; return its JSON when it is an mc-prompter.""" + url = f"http://127.0.0.1:{port}/health" + try: + with urllib.request.urlopen(url, timeout=timeout) as resp: + data = json.loads(resp.read().decode("utf-8")) + except (urllib.error.URLError, OSError, ValueError): + return None + if isinstance(data, dict) and data.get("app") == "mc-prompter": + return data + return None + + +def describe_session(info): + return ( + f"session {info.get('session')} serving " + f"{info.get('script') or '(no script)'} since {info.get('started')}" + ) + + +def pick_port(requested, explicit, probe=port_is_free, + identify=health_identity, limit=PORT_SCAN_LIMIT): + """Resolve the port to use, auto-incrementing unless explicit. + + Raises PortConflict when the explicitly requested port is busy, and + RuntimeError when no free port is found within the scan limit. + """ + port = requested + for _ in range(limit): + if probe(port): + return port + info = identify(port) + if explicit: + raise PortConflict(port, info) + if info: + print( + f"port {port} runs an mc-prompter " + f"({describe_session(info)}); trying {port + 1}" + ) + else: + print(f"port {port} is busy; trying {port + 1}") + port += 1 + raise RuntimeError( + f"no free port found in {requested}..{requested + limit - 1}" + ) + + +def session_file_path(port): + return Path(tempfile.gettempdir()) / "mc-prompter" / f"session-{port}.json" + + +def write_session_file(path, port, pid, token, script): + """Write the session JSON the skill reads back; returns the path.""" + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + payload = { + "port": port, + "pid": pid, + "token": token, + "script": str(script), + "started": datetime.datetime.now().isoformat(timespec="seconds"), + } + path.write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8") + return path + + +def build_server_cmd(port, host, script, owner_wpm, session_file, token): + """The child process command; runs with cwd = the scripts directory.""" + # Phase B seam: when the prompter-lab workspace exists, swap the + # `uv run --with` prefix for `/.venv python -m server.main` + # (POSIX .venv/bin/python, Windows .venv\Scripts\python.exe) with the + # same cwd and the same explicit args. TODO(phase-b): implement the + # workspace-present branch here; do not add config discovery. + return [ + "uv", "run", "--with", AIOHTTP_PIN, "python", "-m", "server.main", + "--port", str(port), + "--host", host, + "--script", str(script) if script else "", + "--owner-wpm", str(owner_wpm or 0), + "--session-file", str(session_file), + "--token", token, + ] + + +def spawn_server(cmd): + """Spawn the server command in its own process group. + + The child is a `uv run` wrapper whose grandchild is the actual python + server; isolating the tree in its own group lets terminate_server + signal all of it, so no orphan keeps the port after shutdown. + """ + if sys.platform == "win32": + return subprocess.Popen( + cmd, cwd=SCRIPTS_DIR, + creationflags=subprocess.CREATE_NEW_PROCESS_GROUP, + ) + return subprocess.Popen(cmd, cwd=SCRIPTS_DIR, start_new_session=True) + + +def terminate_server(child, timeout=10.0): + """Terminate the whole server process tree and reap the child. + + POSIX: SIGTERM the process group, escalate to SIGKILL after timeout. + Windows: terminate()/kill() the child (spawned with its own process + group; uv forwards termination on Windows). + """ + if child.poll() is not None: + return + if sys.platform == "win32": + child.terminate() + try: + child.wait(timeout=timeout) + except subprocess.TimeoutExpired: + child.kill() + child.wait() + return + try: + pgid = os.getpgid(child.pid) + except ProcessLookupError: + child.wait() + return + with contextlib.suppress(ProcessLookupError): + os.killpg(pgid, signal.SIGTERM) + try: + child.wait(timeout=timeout) + except subprocess.TimeoutExpired: + with contextlib.suppress(ProcessLookupError): + os.killpg(pgid, signal.SIGKILL) + child.wait() + + +def local_ip(): + """Best-effort LAN IP via the UDP connect trick (no packets sent).""" + try: + with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as sock: + sock.connect(("192.0.2.1", 80)) + return sock.getsockname()[0] + except OSError: + return "127.0.0.1" + + +def wait_for_health(port, child, timeout=STARTUP_TIMEOUT): + """Poll /health until the child answers as mc-prompter or dies.""" + deadline = time.monotonic() + timeout + while time.monotonic() < deadline: + if child.poll() is not None: + return None + info = health_identity(port) + if info: + return info + time.sleep(0.25) + return None + + +def main(argv=None): + parser = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + parser.add_argument("--script", required=True, + help="path to the script file to prompt") + parser.add_argument("--port", type=int, default=None, + help=f"port (default {DEFAULT_PORT}, " + "auto-increments when busy)") + parser.add_argument("--lan", action="store_true", + help="bind 0.0.0.0 and print the LAN remote URL") + parser.add_argument("--owner-wpm", type=int, default=0, + help="creator wpm for read-time estimates, 0 = unset") + parser.add_argument("--no-open", action="store_true", + help="do not open the browser at the home page") + args = parser.parse_args(argv) + + script = Path(args.script).expanduser().resolve() + if not script.is_file(): + print(f"error: script not found: {script}", file=sys.stderr) + return 4 + try: + script.read_text(encoding="utf-8") + except (OSError, UnicodeDecodeError) as exc: + print(f"error: cannot read {script}: {exc}", file=sys.stderr) + return 4 + + explicit = args.port is not None + requested = args.port if explicit else DEFAULT_PORT + try: + port = pick_port(requested, explicit) + except PortConflict as exc: + print(f"error: port {exc.port} is already in use.", file=sys.stderr) + if exc.info: + print( + f" an mc-prompter is running there: " + f"{describe_session(exc.info)}", + file=sys.stderr, + ) + print( + " pick another --port, or omit --port to auto-select one.", + file=sys.stderr, + ) + return 5 + except RuntimeError as exc: + print(f"error: {exc}", file=sys.stderr) + return 1 + + host = "0.0.0.0" if args.lan else "127.0.0.1" + token = secrets.token_urlsafe(16) + session_file = session_file_path(port) + cmd = build_server_cmd(port, host, script, args.owner_wpm, + session_file, token) + + child = spawn_server(cmd) + try: + info = wait_for_health(port, child) + if info is None: + print("error: server failed to start", file=sys.stderr) + terminate_server(child) + return 1 + + write_session_file(session_file, port, child.pid, token, script) + + base_url = f"http://127.0.0.1:{port}" + local_url = f"{base_url}/?token={token}" + print(f"mc-prompter is up (session {token[:8]})") + print(f" home: {local_url}") + print(f" prompt: {base_url}/prompt") + print(f" remote: http://127.0.0.1:{port}/remote?token={token}") + if args.lan: + ip = local_ip() + print(f" LAN remote: http://{ip}:{port}/remote?token={token}") + print(" note: on Windows the first --lan launch triggers a " + "Windows Firewall consent dialog; allow it for the LAN " + "remote to reach the server.") + print(f" session file: {session_file}", flush=True) + + if not args.no_open: + webbrowser.open(local_url) + + def _on_terminate(signum, frame): + raise KeyboardInterrupt + + signal.signal(signal.SIGTERM, _on_terminate) + try: + child.wait() + except KeyboardInterrupt: + terminate_server(child) + return 0 + return 0 if child.returncode == 0 else 1 + finally: + # Never leave a stale session file advertising a dead pid/token. + with contextlib.suppress(OSError): + session_file.unlink(missing_ok=True) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/skills/mc-prompter/scripts/server/__init__.py b/skills/mc-prompter/scripts/server/__init__.py new file mode 100644 index 0000000..e4045c6 --- /dev/null +++ b/skills/mc-prompter/scripts/server/__init__.py @@ -0,0 +1,7 @@ +"""mc-prompter server package (Phase A). + +Launched as `python -m server.main` with cwd set to the scripts directory, +so the package resolves identically under both launch paths (uv run with +aiohttp as a PEP 723 dependency now; the prompter-lab workspace venv in +Phase B). +""" diff --git a/skills/mc-prompter/scripts/server/main.py b/skills/mc-prompter/scripts/server/main.py new file mode 100644 index 0000000..5878f72 --- /dev/null +++ b/skills/mc-prompter/scripts/server/main.py @@ -0,0 +1,542 @@ +#!/usr/bin/env python3 +# /// script +# requires-python = ">=3.11" +# dependencies = ["aiohttp==3.12.15"] +# /// +"""mc-prompter aiohttp server (Phase A: classic teleprompter). + +Launched by run_prompter.py as `python -m server.main` with cwd set to the +scripts directory (the PEP 723 header above also allows a direct +`uv run main.py` for debugging). All arguments are explicit; the server does +no config discovery. + +Usage: + python -m server.main --port 8770 --host 127.0.0.1|0.0.0.0 + --script --owner-wpm + --session-file --token + +HTTP routes: + GET / home page (static/home.html) + GET /prompt prompter display (static/prompt.html) + GET /remote phone remote (static/remote.html) + GET /overlay OBS overlay (static/overlay.html) + GET /static/* static assets + GET /health identity JSON, no auth: + {"app": "mc-prompter", "version": "0.1.0", + "session": "", + "script": "", "port": N, + "started": ""} + GET /api/state {"snapshot": , "doc-version": n, + "script": {"path", "title", "word-count"} or null, + "config": {"owner-wpm": N or null}} + GET /api/source {"path": , "raw": "", + "doc": + + + + diff --git a/skills/mc-prompter/scripts/server/static/js/home.js b/skills/mc-prompter/scripts/server/static/js/home.js new file mode 100644 index 0000000..97c4c3e --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/home.js @@ -0,0 +1,221 @@ +/* mc-prompter home page glue. + * + * Source picker (path load or paste), edit-in-place with save (writes the + * file after a timestamped backup) and session-only apply, script info, + * links to /prompt and /overlay, and the remote URL with token plus copy. + * + * Token discovery: the launcher opens the home page with ?token=... in the + * URL. We keep it in sessionStorage so in-app navigation survives. If no + * token reaches this page, the remote URL is shown without one (read-only on + * other devices) and a warning explains where to find the full URL. + * SEAM(token): if the server later exposes the token to loopback callers in + * GET /api/state config, read it in fetchState() below. + */ +(function () { + 'use strict'; + var MC = window.MC; + + var els = { + session: document.getElementById('session'), + conn: document.getElementById('conn'), + connLabel: document.getElementById('conn-label'), + pathInput: document.getElementById('path-input'), + btnLoad: document.getElementById('btn-load'), + infoPath: document.getElementById('info-path'), + infoTitle: document.getElementById('info-title'), + infoWords: document.getElementById('info-words'), + infoEst: document.getElementById('info-est'), + remoteUrl: document.getElementById('remote-url'), + linkPrompt: document.getElementById('link-prompt'), + linkOverlay: document.getElementById('link-overlay'), + btnCopy: document.getElementById('btn-copy'), + tokenWarning: document.getElementById('token-warning'), + editor: document.getElementById('editor'), + btnApply: document.getElementById('btn-apply'), + btnSave: document.getElementById('btn-save'), + btnRevert: document.getElementById('btn-revert'), + dirtyChip: document.getElementById('dirty-chip'), + msg: document.getElementById('msg') + }; + + var TOKEN_STORE = 'mc-prompter-token'; + var token = new URLSearchParams(location.search).get('token') || null; + if (token) { + try { sessionStorage.setItem(TOKEN_STORE, token); } catch (e) { /* fine */ } + } else { + try { token = sessionStorage.getItem(TOKEN_STORE); } catch (e) { /* fine */ } + } + + var docVersion = null; + var loadedRaw = ''; + var dirty = false; + var currentWpm = 150; + var wordCount = null; + + var ws = MC.createWS({ role: 'home', token: token }); + + // ---------- messages ---------- + + function say(text, kind) { + els.msg.textContent = text || ''; + els.msg.className = kind || ''; + } + + function setDirty(d) { + dirty = d; + els.dirtyChip.classList.toggle('hidden', !d); + } + + // ---------- remote URL ---------- + + function renderRemoteUrl() { + var url = location.origin + '/remote' + (token ? '?token=' + encodeURIComponent(token) : ''); + els.remoteUrl.textContent = url; + els.tokenWarning.classList.toggle('hidden', !!token); + // The prompter display and overlay links need the token too, or they are + // rejected when the home page is opened from another device on the LAN. + els.linkPrompt.href = MC.model.withToken('/prompt', token); + els.linkOverlay.href = MC.model.withToken('/overlay', token); + } + + els.btnCopy.addEventListener('click', function () { + var text = els.remoteUrl.textContent; + function fallbackCopy() { + var range = document.createRange(); + range.selectNodeContents(els.remoteUrl); + var sel = window.getSelection(); + sel.removeAllRanges(); + sel.addRange(range); + try { document.execCommand('copy'); say('copied', 'ok'); } + catch (e) { say('copy failed, select it manually', 'bad'); } + } + if (navigator.clipboard && navigator.clipboard.writeText) { + navigator.clipboard.writeText(text).then(function () { say('copied', 'ok'); }, fallbackCopy); + } else { + fallbackCopy(); + } + }); + + // ---------- script info ---------- + + function renderInfo(path, doc) { + els.infoPath.textContent = path || 'none loaded'; + els.infoTitle.textContent = (doc && doc.title) || '-'; + wordCount = doc ? doc['word-count'] : null; + els.infoWords.textContent = wordCount !== null && wordCount !== undefined ? String(wordCount) : '-'; + renderEst(); + } + + function renderEst() { + var mins = MC.model.estimateMinutes(wordCount, currentWpm); + els.infoEst.textContent = mins === null ? '-' : + MC.model.fmtClock(mins * 60) + ' at ' + currentWpm + ' wpm'; + } + + // ---------- source loading ---------- + + function refreshSource(overwriteEditor) { + return MC.model.fetchSource(token).then(function (src) { + docVersion = src['doc-version']; + loadedRaw = src.raw || ''; + renderInfo(src.path, src.doc); + if (overwriteEditor || !dirty) { + els.editor.value = loadedRaw; + setDirty(false); + } + }).catch(function (err) { + say('could not fetch script: ' + err.message, 'bad'); + }); + } + + els.btnLoad.addEventListener('click', function () { + var path = els.pathInput.value.trim(); + if (!path) { say('enter an absolute path first', 'bad'); return; } + say('loading...'); + MC.model.postJSON('/api/source/load', { path: path }, token).then(function () { + say('loaded', 'ok'); + return refreshSource(true); + }).catch(function (err) { + say('load failed: ' + err.message, 'bad'); + }); + }); + + els.pathInput.addEventListener('keydown', function (e) { + if (e.key === 'Enter') els.btnLoad.click(); + }); + + // ---------- editor ---------- + + els.editor.addEventListener('input', function () { + setDirty(els.editor.value !== loadedRaw); + }); + + function pushSource(save) { + say(save ? 'saving...' : 'applying...'); + MC.model.postJSON('/api/source', { raw: els.editor.value, save: save }, token) + .then(function (res) { + docVersion = res ? res['doc-version'] : docVersion; + loadedRaw = els.editor.value; + setDirty(false); + if (save && res && res.backup) { + say('saved (backup: ' + res.backup + ')', 'ok'); + } else if (save) { + say('saved', 'ok'); + } else { + say('applied for this session', 'ok'); + } + return refreshSource(false); + }) + .catch(function (err) { + say((save ? 'save' : 'apply') + ' failed: ' + err.message, 'bad'); + }); + } + + els.btnApply.addEventListener('click', function () { pushSource(false); }); + els.btnSave.addEventListener('click', function () { pushSource(true); }); + els.btnRevert.addEventListener('click', function () { + els.editor.value = loadedRaw; + setDirty(false); + say('reverted to the last loaded text'); + }); + + // ---------- WS wiring ---------- + + ws.onStatus(function (s) { + els.conn.classList.toggle('on', s.connected); + els.connLabel.textContent = s.connected ? 'connected' : (s.rejected ? 'token rejected' : 'reconnecting'); + els.session.textContent = s.session || ''; + }); + + ws.on('doc-updated', function (msg) { + if (msg['doc-version'] !== docVersion) { + refreshSource(false); + if (dirty) say('script changed elsewhere; your unsaved edits are kept in the editor', 'bad'); + } + }); + + ws.on('state', function (msg) { + if (typeof msg.wpm === 'number' && msg.wpm > 0) { + currentWpm = msg.wpm; + renderEst(); + } + }); + + // ---------- boot ---------- + + renderRemoteUrl(); + MC.model.fetchState(token).then(function (state) { + if (state && state.config && state.config['owner-wpm']) { + currentWpm = state.config['owner-wpm']; + } + if (state && state.snapshot && typeof state.snapshot.wpm === 'number') { + currentWpm = state.snapshot.wpm; + } + if (state && state.script && state.script.path) { + els.pathInput.value = state.script.path; + } + renderEst(); + }).catch(function () { /* defaults are fine */ }).then(function () { + return refreshSource(true); + }); +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/model.js b/skills/mc-prompter/scripts/server/static/js/model.js new file mode 100644 index 0000000..61040f1 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/model.js @@ -0,0 +1,203 @@ +/* mc-prompter script model client (classic script, attaches to window.MC). + * + * Fetches the ingested script model from the server and renders it to DOM. + * + * Model shape (produced by script_ingest.py, see the Phase A contract): + * { title, "word-count", sections: [ { id, heading, level, blocks: [ + * {type:"para", runs:[{text, flags}]}, + * {type:"note", text}, + * {type:"take", source, start, end, runs:[...]} ] } ] } + * + * Rendering contract (Phase B alignment and click-to-anchor build on this): + * - every speakable word is wrapped in + * - take paragraphs get class "take" (dimmed, "have it already" badge, + * hidden entirely when body has class hide-takes) + * - invented runs get class "invented" (subtle badge, muted when body has + * class hide-invented) + * - notes render as class "note" (dimmed, italic, never counted) + */ +(function () { + 'use strict'; + window.MC = window.MC || {}; + + function withToken(url, token) { + if (!token) return url; + return url + (url.indexOf('?') >= 0 ? '&' : '?') + 'token=' + encodeURIComponent(token); + } + + function fetchJSON(url, token) { + return fetch(withToken(url, token), { cache: 'no-store' }).then(function (r) { + if (!r.ok) { + return r.text().then(function (body) { + var err = new Error('HTTP ' + r.status + ' for ' + url + (body ? ': ' + body : '')); + err.status = r.status; + throw err; + }); + } + return r.json(); + }); + } + + function postJSON(url, body, token) { + return fetch(withToken(url, token), { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(body) + }).then(function (r) { + return r.text().then(function (text) { + var data = null; + try { data = text ? JSON.parse(text) : null; } catch (e) { /* non-JSON error body */ } + if (!r.ok) { + var msg = (data && data.message) ? data.message : ('HTTP ' + r.status); + var err = new Error(msg); + err.status = r.status; + throw err; + } + return data; + }); + }); + } + + // GET /api/source -> {path, raw, doc, "doc-version"} + function fetchSource(token) { return fetchJSON('/api/source', token); } + + // GET /api/state -> {snapshot, "doc-version", script, config} + function fetchState(token) { return fetchJSON('/api/state', token); } + + /* Render the script model into `container` (emptied first). + * Returns an index used by the scroll engine and jump logic: + * { wordCount, takeWordCount, words: [span], + * paragraphs: [{el, wordStart, sectionId}], + * sections: [{id, heading, level, el, wordStart}] } + * takeWordCount is the share of wordCount inside TAKE paragraphs, so the + * page can pace against the visible count when hide-takes is on. + */ + function renderDoc(doc, container) { + while (container.firstChild) container.removeChild(container.firstChild); + + var index = { wordCount: 0, takeWordCount: 0, words: [], paragraphs: [], sections: [] }; + if (!doc || !doc.sections) return index; + + var wordIndex = 0; + + function appendWords(text, parent) { + // Split preserving whitespace so spacing survives verbatim. + var chunks = String(text).split(/(\s+)/); + for (var i = 0; i < chunks.length; i++) { + var chunk = chunks[i]; + if (!chunk) continue; + if (/^\s+$/.test(chunk)) { + parent.appendChild(document.createTextNode(chunk)); + } else { + var span = document.createElement('span'); + span.className = 'w'; + span.dataset.i = String(wordIndex); + span.textContent = chunk; + parent.appendChild(span); + index.words.push(span); + wordIndex += 1; + } + } + } + + function appendRuns(runs, parent) { + for (var i = 0; i < (runs || []).length; i++) { + var run = runs[i]; + var flags = run.flags || []; + if (flags.indexOf('invented') >= 0) { + var inv = document.createElement('span'); + inv.className = 'run invented'; + appendWords(run.text, inv); + parent.appendChild(inv); + } else { + appendWords(run.text, parent); + } + } + } + + for (var s = 0; s < doc.sections.length; s++) { + var sec = doc.sections[s]; + var secEl = document.createElement('section'); + secEl.className = 'sec'; + secEl.dataset.sid = sec.id; + var secEntry = { + id: sec.id, + heading: sec.heading || null, + level: sec.level || 2, + el: secEl, + wordStart: wordIndex + }; + + if (sec.heading) { + var h = document.createElement('h2'); + h.className = 'sec-h'; + h.textContent = sec.heading; + secEl.appendChild(h); + } + + var blocks = sec.blocks || []; + for (var b = 0; b < blocks.length; b++) { + var block = blocks[b]; + if (block.type === 'note') { + var noteEl = document.createElement('p'); + noteEl.className = 'note'; + noteEl.textContent = block.text || ''; + secEl.appendChild(noteEl); + } else if (block.type === 'take') { + var takeEl = document.createElement('p'); + takeEl.className = 'para take'; + // Phase B seam: source clip identity rides on data attributes. + if (block.source !== undefined) takeEl.dataset.source = String(block.source); + if (block.start !== undefined) takeEl.dataset.start = String(block.start); + if (block.end !== undefined) takeEl.dataset.end = String(block.end); + var takeStart = wordIndex; + appendRuns(block.runs, takeEl); + index.takeWordCount += wordIndex - takeStart; + secEl.appendChild(takeEl); + index.paragraphs.push({ el: takeEl, wordStart: takeStart, sectionId: sec.id }); + } else { + // para (default) + var paraEl = document.createElement('p'); + paraEl.className = 'para'; + var paraStart = wordIndex; + appendRuns(block.runs, paraEl); + secEl.appendChild(paraEl); + index.paragraphs.push({ el: paraEl, wordStart: paraStart, sectionId: sec.id }); + } + } + + container.appendChild(secEl); + index.sections.push(secEntry); + } + + index.wordCount = wordIndex; + return index; + } + + function estimateMinutes(wordCount, wpm) { + if (!wordCount || !wpm) return null; + return wordCount / wpm; + } + + // Format seconds as m:ss or h:mm:ss. Returns "--:--" for null/NaN. + function fmtClock(seconds) { + if (seconds === null || seconds === undefined || isNaN(seconds)) return '--:--'; + var t = Math.max(0, Math.round(seconds)); + var h = Math.floor(t / 3600); + var m = Math.floor((t % 3600) / 60); + var sec = t % 60; + var mm = (h > 0 && m < 10 ? '0' : '') + m; + var ss = (sec < 10 ? '0' : '') + sec; + return h > 0 ? h + ':' + mm + ':' + ss : mm + ':' + ss; + } + + window.MC.model = { + fetchSource: fetchSource, + fetchState: fetchState, + postJSON: postJSON, + withToken: withToken, + renderDoc: renderDoc, + estimateMinutes: estimateMinutes, + fmtClock: fmtClock + }; +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/overlay.js b/skills/mc-prompter/scripts/server/static/js/overlay.js new file mode 100644 index 0000000..e846f04 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/overlay.js @@ -0,0 +1,33 @@ +/* mc-prompter /overlay page glue (Phase A placeholder). + * + * OBS browser source: transparent background, a small "session connected" + * badge that hides itself 5 seconds after each (re)connect. Stays on the WS + * so Phase C can mount the producer rail and cue cards without changing the + * page contract. + */ +(function () { + 'use strict'; + var MC = window.MC; + + var badge = document.getElementById('badge'); + var hideTimer = null; + + var token = new URLSearchParams(location.search).get('token') || null; + var ws = MC.createWS({ role: 'overlay', token: token }); + + function showBadge(connected) { + badge.classList.remove('fade'); + badge.classList.toggle('disconnected', !connected); + clearTimeout(hideTimer); + if (connected) { + hideTimer = setTimeout(function () { badge.classList.add('fade'); }, 5000); + } + } + + ws.onStatus(function (s) { + showBadge(s.connected); + }); + + // Phase C seam: subscribe here for producer rail state. + // ws.on('state', function (msg) { ... }); +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/prompt.js b/skills/mc-prompter/scripts/server/static/js/prompt.js new file mode 100644 index 0000000..179f8f4 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/prompt.js @@ -0,0 +1,750 @@ +/* mc-prompter /prompt page glue. + * + * Owns the scroll engine, display settings, keyboard map, section list, + * settings drawer, countdown overlay, and the WS state loop. + * + * Leadership: the first prompt connection is the leader (the server says so + * in the welcome frame and promotes on disconnect via a role frame). Only + * the leader runs the engine autonomously and sends state snapshots (~4 Hz + * while playing plus on every change). Non-leader prompt pages are + * display-only followers: they apply incoming state frames to their own + * scroll surface so a tablet pointed at /prompt mirrors the leader. + */ +(function () { + 'use strict'; + var MC = window.MC; + + // ---------- elements ---------- + + var els = { + stage: document.getElementById('stage'), + surface: document.getElementById('surface'), + script: document.getElementById('script'), + eyeline: document.getElementById('eyeline'), + hud: document.getElementById('hud'), + conn: document.getElementById('conn'), + roleBadge: document.getElementById('role-badge'), + btnToggle: document.getElementById('btn-toggle'), + btnRestart: document.getElementById('btn-restart'), + clockElapsed: document.getElementById('clock-elapsed'), + clockRemaining: document.getElementById('clock-remaining'), + driftChip: document.getElementById('drift-chip'), + wpmVal: document.getElementById('wpm-val'), + modeChip: document.getElementById('mode-chip'), + countdown: document.getElementById('countdown'), + countdownNum: document.getElementById('countdown-num'), + sectionsDrawer: document.getElementById('sections-drawer'), + sectionsList: document.getElementById('sections-list'), + estTime: document.getElementById('est-time'), + settingsDrawer: document.getElementById('settings-drawer'), + helpOverlay: document.getElementById('help-overlay'), + tokenError: document.getElementById('token-error'), + toast: document.getElementById('toast') + }; + + // ---------- state ---------- + + var settings = MC.settings.load(); + var docIndex = { wordCount: 0, takeWordCount: 0, words: [], paragraphs: [], sections: [] }; + var docVersion = null; + var eyelineHidden = false; + var lastRemoteState = null; // latest leader snapshot seen while follower + var stateDirty = false; + var lastStateSent = 0; + var toastTimer = null; + var hudTimer = null; + + var token = new URLSearchParams(location.search).get('token') || null; + + var ws = MC.createWS({ role: 'prompt', token: token }); + + // Words the reader actually sees: hidden TAKE paragraphs contribute no + // scroll height, so they must not count toward pacing or time estimates. + // The engine reads this fresh every frame, so the hide-takes toggle is + // picked up immediately. + function visibleWordCount() { + var wc = docIndex.wordCount; + if (settings['hide-takes']) wc -= (docIndex.takeWordCount || 0); + return Math.max(wc, 0); + } + + var engine = MC.createEngine({ + surface: els.surface, + getWordCount: visibleWordCount, + onFrame: onEngineFrame, + onChange: onEngineChange, + onFinish: function () { toast('end of script'); } + }); + + // ---------- helpers ---------- + + function isLeader() { return ws.state.leader; } + + // Whether this page drives its own engine: the elected leader does, and so + // does a page with no server connection at all (standalone fallback, so the + // prompter never goes dead if the WS drops mid-take). + function drives() { return ws.state.leader || !ws.state.connected; } + + function toast(msg) { + els.toast.textContent = msg; + els.toast.classList.add('show'); + clearTimeout(toastTimer); + toastTimer = setTimeout(function () { els.toast.classList.remove('show'); }, 1800); + } + + function eyelinePx() { + return els.surface.clientHeight * (Number(settings['eyeline-percent']) / 100); + } + + // Document-space y of the eyeline (what the reader is looking at). + function eyelineDocY() { + return els.surface.scrollTop + eyelinePx(); + } + + function currentSectionId() { + var y = eyelineDocY(); + var id = null; + for (var i = 0; i < docIndex.sections.length; i++) { + if (docIndex.sections[i].el.offsetTop <= y + 2) id = docIndex.sections[i].id; + else break; + } + if (id === null && docIndex.sections.length) id = docIndex.sections[0].id; + return id; + } + + function jumpToEl(el) { + engine.jumpToPx(el.offsetTop - eyelinePx()); + } + + function jumpSection(id) { + for (var i = 0; i < docIndex.sections.length; i++) { + if (docIndex.sections[i].id === id) { jumpToEl(docIndex.sections[i].el); return; } + } + } + + // delta paragraphs: negative = previous, positive = next. + function jumpParagraphs(delta) { + var paras = docIndex.paragraphs; + if (!paras.length) return; + var y = eyelineDocY(); + // Index of the paragraph the eyeline currently sits in (last with top <= y). + var cur = -1; + for (var i = 0; i < paras.length; i++) { + if (paras[i].el.offsetTop <= y + 2) cur = i; else break; + } + var target = Math.min(Math.max(cur + delta, 0), paras.length - 1); + if (target === cur && delta < 0 && cur >= 0) { + // Already at this paragraph top? Snap to its top anyway (re-read). + target = Math.max(cur - 1, 0); + } + jumpToEl(paras[target].el); + } + + // ---------- display settings ---------- + + function applyDisplay() { + var s = settings; + els.script.style.fontFamily = s['font-family']; + els.script.style.fontSize = s['font-size'] + 'px'; + els.script.style.lineHeight = String(s['line-height']); + els.script.style.color = s['text-color']; + document.body.style.background = s['background-color']; + els.script.style.paddingLeft = s['margin-percent'] + '%'; + els.script.style.paddingRight = s['margin-percent'] + '%'; + // Lead-in and run-out so the first and last words can reach the eyeline. + els.script.style.paddingTop = s['eyeline-percent'] + 'vh'; + els.script.style.paddingBottom = (100 - Number(s['eyeline-percent'])) + 'vh'; + + var t = ''; + if (s['mirror-h']) t += ' scaleX(-1)'; + if (s['mirror-v']) t += ' scaleY(-1)'; + els.stage.style.transform = t ? t.trim() : 'none'; + + els.eyeline.style.top = s['eyeline-percent'] + '%'; + els.eyeline.classList.toggle('style-line', s['eyeline-style'] !== 'arrow'); + els.eyeline.classList.toggle('style-arrow', s['eyeline-style'] === 'arrow'); + els.eyeline.classList.toggle('off', eyelineHidden); + + document.body.classList.toggle('hide-takes', !!s['hide-takes']); + document.body.classList.toggle('hide-invented', !s['show-invented']); + } + + // Apply a settings mutation keeping the read position stable, since layout + // changes rescale the scroll geometry (the engine reads the ratio fresh + // every frame; we only need to preserve the position across the reflow). + function changeSetting(key, value) { + var ratio = engine.getPositionRatio(); + settings[key] = value; + MC.settings.save(settings); + applyDisplay(); + engine.setPositionRatio(ratio); + syncSettingsControls(); + if (key === 'hide-takes') updateEstTime(); + markDirty(); + } + + // ---------- settings drawer wiring ---------- + + var C = { + wpm: document.getElementById('set-wpm'), + mode: document.getElementById('set-mode'), + minutes: document.getElementById('set-minutes'), + rowMinutes: document.getElementById('row-minutes'), + countdown: document.getElementById('set-countdown'), + fontStack: document.getElementById('set-font-stack'), + fontFamily: document.getElementById('set-font-family'), + fontSize: document.getElementById('set-font-size'), + fontSizeN: document.getElementById('set-font-size-n'), + lineHeight: document.getElementById('set-line-height'), + lineHeightV: document.getElementById('set-line-height-v'), + margin: document.getElementById('set-margin'), + marginV: document.getElementById('set-margin-v'), + textColor: document.getElementById('set-text-color'), + bgColor: document.getElementById('set-bg-color'), + mirrorH: document.getElementById('set-mirror-h'), + mirrorV: document.getElementById('set-mirror-v'), + eyeline: document.getElementById('set-eyeline'), + eyelineV: document.getElementById('set-eyeline-v'), + eyelineStyle: document.getElementById('set-eyeline-style'), + hideTakes: document.getElementById('set-hide-takes'), + showInvented: document.getElementById('set-show-invented') + }; + + function initFontStackSelect() { + var stacks = MC.settings.FONT_STACKS; + for (var i = 0; i < stacks.length; i++) { + var opt = document.createElement('option'); + opt.value = stacks[i].value; + opt.textContent = stacks[i].label; + C.fontStack.appendChild(opt); + } + var custom = document.createElement('option'); + custom.value = ''; + custom.textContent = 'custom (type below)'; + C.fontStack.appendChild(custom); + } + + function syncSettingsControls() { + var s = settings; + C.countdown.value = s['countdown-seconds']; + C.fontFamily.value = s['font-family']; + var found = false; + for (var i = 0; i < C.fontStack.options.length; i++) { + if (C.fontStack.options[i].value === s['font-family']) { + C.fontStack.selectedIndex = i; + found = true; + break; + } + } + if (!found) C.fontStack.value = ''; + C.fontSize.value = s['font-size']; + C.fontSizeN.value = s['font-size']; + C.lineHeight.value = s['line-height']; + C.lineHeightV.textContent = Number(s['line-height']).toFixed(2); + C.margin.value = s['margin-percent']; + C.marginV.textContent = s['margin-percent'] + '%'; + C.textColor.value = s['text-color']; + C.bgColor.value = s['background-color']; + C.mirrorH.checked = !!s['mirror-h']; + C.mirrorV.checked = !!s['mirror-v']; + C.eyeline.value = s['eyeline-percent']; + C.eyelineV.textContent = s['eyeline-percent'] + '%'; + C.eyelineStyle.value = s['eyeline-style']; + C.hideTakes.checked = !!s['hide-takes']; + C.showInvented.checked = !!s['show-invented']; + } + + function wireSettingsControls() { + C.wpm.addEventListener('change', function () { + engine.setWpm(Number(C.wpm.value)); + }); + C.mode.addEventListener('change', applyModeControls); + C.minutes.addEventListener('change', applyModeControls); + C.countdown.addEventListener('change', function () { + changeSetting('countdown-seconds', Math.max(0, Number(C.countdown.value) || 0)); + }); + C.fontStack.addEventListener('change', function () { + if (C.fontStack.value) changeSetting('font-family', C.fontStack.value); + }); + C.fontFamily.addEventListener('change', function () { + if (C.fontFamily.value.trim()) changeSetting('font-family', C.fontFamily.value.trim()); + }); + C.fontSize.addEventListener('input', function () { + changeSetting('font-size', Number(C.fontSize.value)); + }); + C.fontSizeN.addEventListener('change', function () { + changeSetting('font-size', Number(C.fontSizeN.value)); + }); + C.lineHeight.addEventListener('input', function () { + changeSetting('line-height', Number(C.lineHeight.value)); + }); + C.margin.addEventListener('input', function () { + changeSetting('margin-percent', Number(C.margin.value)); + }); + C.textColor.addEventListener('input', function () { + changeSetting('text-color', C.textColor.value); + }); + C.bgColor.addEventListener('input', function () { + changeSetting('background-color', C.bgColor.value); + }); + C.mirrorH.addEventListener('change', function () { + changeSetting('mirror-h', C.mirrorH.checked); + }); + C.mirrorV.addEventListener('change', function () { + changeSetting('mirror-v', C.mirrorV.checked); + }); + C.eyeline.addEventListener('input', function () { + changeSetting('eyeline-percent', Number(C.eyeline.value)); + }); + C.eyelineStyle.addEventListener('change', function () { + changeSetting('eyeline-style', C.eyelineStyle.value); + }); + C.hideTakes.addEventListener('change', function () { + changeSetting('hide-takes', C.hideTakes.checked); + }); + C.showInvented.addEventListener('change', function () { + changeSetting('show-invented', C.showInvented.checked); + }); + } + + function applyModeControls() { + var timed = C.mode.value === 'timed'; + C.rowMinutes.classList.toggle('hidden', !timed); + engine.setMode(timed ? 'timed' : 'manual', Number(C.minutes.value)); + } + + // ---------- panels ---------- + + function togglePanel(el) { + el.classList.toggle('hidden'); + } + + function closePanels() { + els.sectionsDrawer.classList.add('hidden'); + els.settingsDrawer.classList.add('hidden'); + els.helpOverlay.classList.add('hidden'); + } + + document.addEventListener('click', function (e) { + var closer = e.target.closest('[data-close]'); + if (closer) document.getElementById(closer.dataset.close).classList.add('hidden'); + if (e.target === els.helpOverlay) els.helpOverlay.classList.add('hidden'); + }); + + // ---------- section list ---------- + + function renderSectionList() { + var list = els.sectionsList; + while (list.firstChild) list.removeChild(list.firstChild); + for (var i = 0; i < docIndex.sections.length; i++) { + (function (sec, i2) { + var li = document.createElement('li'); + li.dataset.sid = sec.id; + var name = document.createElement('span'); + name.textContent = sec.heading || (i2 === 0 ? 'Preamble' : 'Untitled section'); + var words = document.createElement('span'); + words.className = 'sec-words'; + var next = docIndex.sections[i2 + 1]; + var count = (next ? next.wordStart : docIndex.wordCount) - sec.wordStart; + words.textContent = count + ' w'; + li.appendChild(name); + li.appendChild(words); + li.addEventListener('click', function () { + if (drives()) jumpSection(sec.id); + else ws.cmd('jump-section', sec.id); + els.sectionsDrawer.classList.add('hidden'); + }); + list.appendChild(li); + })(docIndex.sections[i], i); + } + updateEstTime(); + } + + function updateEstTime() { + var wpm = engine.view().wpm || 150; + var wc = visibleWordCount(); + var mins = MC.model.estimateMinutes(wc, wpm); + els.estTime.textContent = wc + ' words, est. ' + + MC.model.fmtClock(mins === null ? null : mins * 60) + ' at ' + wpm + ' wpm'; + } + + function highlightCurrentSection(id) { + var items = els.sectionsList.children; + for (var i = 0; i < items.length; i++) { + items[i].classList.toggle('current', items[i].dataset.sid === id); + } + } + + // ---------- document loading ---------- + + function loadSource() { + return MC.model.fetchSource(token).then(function (src) { + var ratio = engine.getPositionRatio(); + docIndex = MC.model.renderDoc(src.doc, els.script); + docVersion = src['doc-version']; + applyDisplay(); + engine.setPositionRatio(ratio); + renderSectionList(); + document.title = 'mc-prompter | ' + ((src.doc && src.doc.title) || 'prompt'); + markDirty(); + }).catch(function (err) { + toast('script load failed: ' + err.message); + }); + } + + // ---------- WS: state out (leader) ---------- + + function markDirty() { stateDirty = true; maybeSendState(); } + + function buildSnapshot() { + var v = engine.view(); + return { + type: 'state', + position: Math.round(v.position * 10000) / 10000, + section: currentSectionId(), + playing: v.playing || (v.countdown !== null), + wpm: v.wpm, + mode: v.mode, + elapsed: Math.round(v.elapsed * 10) / 10, + remaining: v.remaining === null ? null : Math.round(v.remaining * 10) / 10, + countdown: v.countdown === null ? null : Math.round(v.countdown * 10) / 10 + }; + } + + // ~4 Hz while playing, prompt on changes (force), never a flood: forced + // sends still respect a 100 ms floor and leave the dirty flag for the + // heartbeat to flush. + function maybeSendState(force) { + if (!isLeader()) return; + var now = Date.now(); + var minGap = force ? 100 : 250; + if (now - lastStateSent < minGap) { stateDirty = true; return; } + if (ws.send(buildSnapshot())) { + lastStateSent = now; + stateDirty = false; + } + } + + // ---------- engine callbacks ---------- + + function updateHud(v) { + els.clockElapsed.textContent = MC.model.fmtClock(v.elapsed); + els.clockRemaining.textContent = MC.model.fmtClock(v.remaining); + els.wpmVal.textContent = String(v.wpm); + els.modeChip.textContent = v.mode; + els.btnToggle.textContent = (v.playing || v.countdown !== null) ? 'pause' : 'play'; + + if (v.mode === 'timed' && v.drift !== null) { + var d = Math.round(v.drift); + els.driftChip.classList.remove('hidden'); + if (Math.abs(d) <= 3) { + els.driftChip.textContent = 'on plan'; + els.driftChip.className = 'chip good'; + } else if (d > 0) { + els.driftChip.textContent = d + 's behind'; + els.driftChip.className = 'chip warn'; + } else { + els.driftChip.textContent = (-d) + 's ahead'; + els.driftChip.className = 'chip'; + } + } else { + els.driftChip.classList.add('hidden'); + } + + if (v.countdown !== null) { + els.countdown.classList.remove('hidden'); + els.countdownNum.textContent = String(Math.ceil(v.countdown)); + } else { + els.countdown.classList.add('hidden'); + } + + highlightCurrentSection(currentSectionId()); + } + + function onEngineFrame(v) { + if (everConnected && !ws.state.connected && (v.playing || v.countdown !== null)) { + droveWhileDisconnected = true; + } + if (drives()) updateHud(v); + if (v.playing) maybeSendState(); + } + + function onEngineChange(reason, v) { + if (everConnected && !ws.state.connected && reason !== 'seed') { + droveWhileDisconnected = true; + } + if (!drives()) return; + updateHud(v); + if (reason === 'speed' || reason === 'mode') updateEstTime(); + C.wpm.value = engine.getManualWpm(); + maybeSendState(true); + } + + // ---------- playback intents ---------- + + function startPlay() { + if (engine.getPositionRatio() < 0.001 && Number(settings['countdown-seconds']) > 0) { + engine.beginCountdown(Number(settings['countdown-seconds'])); + } else { + engine.play(); + } + } + + function handleCmd(msg) { + // Commands are relayed to everyone; only the leader executes them. + if (!isLeader()) return; + var v = msg.value; + switch (msg.cmd) { + case 'play': startPlay(); break; + case 'pause': engine.pause(); break; + case 'toggle': engine.toggle(Number(settings['countdown-seconds'])); break; + case 'restart': engine.restart(); break; + case 'speed-delta': engine.deltaWpm(typeof v === 'number' ? v : 2); break; + case 'speed-set': if (typeof v === 'number') engine.setWpm(v); break; + case 'jump-section': if (v) jumpSection(String(v)); break; + case 'jump-words': jumpParagraphs((Number(v) || 0) >= 0 ? 1 : -1); break; + case 'countdown': + engine.beginCountdown(typeof v === 'number' ? v : Number(settings['countdown-seconds'])); + break; + default: break; + } + } + + // ---------- follower: apply leader state ---------- + + function applyRemoteState(msg) { + lastRemoteState = msg; + if (isLeader()) return; + var scrollableH = Math.max(els.surface.scrollHeight - els.surface.clientHeight, 1); + els.surface.scrollTop = (Number(msg.position) || 0) * scrollableH; + els.clockElapsed.textContent = MC.model.fmtClock(msg.elapsed); + els.clockRemaining.textContent = MC.model.fmtClock(msg.remaining); + els.wpmVal.textContent = String(msg.wpm); + els.modeChip.textContent = msg.mode || 'manual'; + els.btnToggle.textContent = msg.playing ? 'pause' : 'play'; + if (msg.countdown !== null && msg.countdown !== undefined) { + els.countdown.classList.remove('hidden'); + els.countdownNum.textContent = String(Math.ceil(msg.countdown)); + } else { + els.countdown.classList.add('hidden'); + } + highlightCurrentSection(msg.section || null); + } + + // ---------- WS wiring ---------- + + var wasLeader = false; + var everConnected = false; // this page has reached the server before + var droveWhileDisconnected = false; // engine ran or the user drove it during an outage + + ws.onStatus(function (s) { + els.conn.classList.toggle('on', s.connected); + els.roleBadge.classList.toggle('hidden', s.leader || !s.connected); + if (s.rejected) { + // Token rejected (close 4403): the WS layer has stopped retrying, so + // surface the failure instead of silently running standalone. + els.tokenError.classList.remove('hidden'); + wasLeader = false; + return; + } + if (s.connected && !s.leader) { + // Connected follower (fresh join or demotion after a reconnect): + // exactly one driver per session, so stop a locally running engine + // before mirroring the leader. + if (engine.isPlaying()) engine.pause(); + // Seed the display from the welcome snapshot until live state frames + // arrive (a paused leader sends none). + if (!lastRemoteState && s.snapshot) applyRemoteState(s.snapshot); + } + if (s.connected && s.leader && !wasLeader) { + // Just became leader (first welcome, reconnect, or promotion after the + // previous leader dropped): seed the engine from the freshest snapshot + // we know about, then start reporting state. Skip the seed if this + // page kept driving through the outage (standalone fallback): the + // server's cached snapshot predates the disconnect and seeding from it + // would rewind the scroll mid-read. + var snap = lastRemoteState || s.snapshot; + if (snap && !droveWhileDisconnected) engine.seed(snap); + markDirty(); + } + if (s.connected) { + everConnected = true; + droveWhileDisconnected = false; + } + wasLeader = s.connected && s.leader; + }); + + ws.on('cmd', handleCmd); + ws.on('state', applyRemoteState); + ws.on('doc-updated', function (msg) { + if (msg['doc-version'] !== docVersion) loadSource(); + }); + ws.on('error', function (msg) { toast('server: ' + msg.message); }); + + // ---------- input: keyboard, wheel, click, scroll ---------- + + document.addEventListener('keydown', function (e) { + var tag = (e.target && e.target.tagName) || ''; + if (tag === 'INPUT' || tag === 'TEXTAREA' || tag === 'SELECT') return; + + var leaderAct = drives(); + function act(fn, cmdName, cmdValue) { + if (leaderAct) fn(); + else ws.cmd(cmdName, cmdValue); + } + + switch (e.key) { + case ' ': + e.preventDefault(); + act(function () { engine.toggle(Number(settings['countdown-seconds'])); }, 'toggle'); + break; + case 'ArrowUp': + case '+': + case '=': + e.preventDefault(); + act(function () { engine.deltaWpm(2); }, 'speed-delta', 2); + break; + case 'ArrowDown': + case '-': + e.preventDefault(); + act(function () { engine.deltaWpm(-2); }, 'speed-delta', -2); + break; + case 'ArrowLeft': + e.preventDefault(); + act(function () { jumpParagraphs(-1); }, 'jump-words', -1); + break; + case 'ArrowRight': + e.preventDefault(); + act(function () { jumpParagraphs(1); }, 'jump-words', 1); + break; + case 'Home': + e.preventDefault(); + act(function () { engine.restart(); }, 'restart'); + break; + case 'f': + case 'F': + e.preventDefault(); + if (document.fullscreenElement) document.exitFullscreen(); + else document.documentElement.requestFullscreen(); + break; + case 'm': + case 'M': { + e.preventDefault(); + // Cycle off -> h -> v -> both -> off. + var h = settings['mirror-h'], vv = settings['mirror-v']; + var next = !h && !vv ? [true, false] : h && !vv ? [false, true] : !h && vv ? [true, true] : [false, false]; + changeSetting('mirror-h', next[0]); + changeSetting('mirror-v', next[1]); + toast('mirror: ' + (next[0] && next[1] ? 'both' : next[0] ? 'horizontal' : next[1] ? 'vertical' : 'off')); + break; + } + case 'e': + case 'E': + e.preventDefault(); + eyelineHidden = !eyelineHidden; + els.eyeline.classList.toggle('off', eyelineHidden); + break; + case 'c': + case 'C': + e.preventDefault(); + act(function () { engine.beginCountdown(Number(settings['countdown-seconds'])); }, + 'countdown', Number(settings['countdown-seconds'])); + break; + case 's': + case 'S': + e.preventDefault(); + togglePanel(els.sectionsDrawer); + break; + case 'd': + case 'D': + e.preventDefault(); + togglePanel(els.settingsDrawer); + break; + case '?': + e.preventDefault(); + togglePanel(els.helpOverlay); + break; + case 'Escape': + closePanels(); + break; + default: + break; + } + }); + + els.surface.addEventListener('wheel', function (e) { + e.preventDefault(); + var d = e.deltaY < 0 ? 2 : -2; + if (drives()) engine.deltaWpm(d); + else ws.cmd('speed-delta', d); + }, { passive: false }); + + els.surface.addEventListener('click', function (e) { + var w = e.target.closest('.w'); + if (!w || !drives()) return; + // Click-to-jump: put the clicked word on the eyeline. + // Phase B seam: this same span is the click-to-anchor target. + engine.jumpToPx(w.offsetTop - eyelinePx()); + }); + + els.surface.addEventListener('scroll', function () { + engine.adoptScrollTop(); + }); + + window.addEventListener('resize', function () { + var ratio = engine.getPositionRatio(); + engine.recalc(); + engine.setPositionRatio(ratio); + }); + + // HUD auto-fade while playing. + function pokeHud() { + els.hud.classList.remove('faded'); + clearTimeout(hudTimer); + hudTimer = setTimeout(function () { + if (engine.isPlaying()) els.hud.classList.add('faded'); + }, 3000); + } + document.addEventListener('mousemove', pokeHud); + pokeHud(); + + // ---------- HUD buttons ---------- + + els.btnToggle.addEventListener('click', function () { + if (drives()) engine.toggle(Number(settings['countdown-seconds'])); + else ws.cmd('toggle'); + }); + els.btnRestart.addEventListener('click', function () { + if (drives()) engine.restart(); + else ws.cmd('restart'); + }); + document.getElementById('btn-sections').addEventListener('click', function () { + togglePanel(els.sectionsDrawer); + }); + document.getElementById('btn-settings').addEventListener('click', function () { + togglePanel(els.settingsDrawer); + }); + document.getElementById('btn-help').addEventListener('click', function () { + togglePanel(els.helpOverlay); + }); + + // Send-state heartbeat: 4 Hz while playing, on-change otherwise. + setInterval(function () { + if (isLeader() && (engine.isPlaying() || stateDirty)) maybeSendState(); + }, 250); + + // ---------- boot ---------- + + initFontStackSelect(); + applyDisplay(); + syncSettingsControls(); + wireSettingsControls(); + + MC.model.fetchState(token).then(function (state) { + var wpm = state && state.config && state.config['owner-wpm']; + if (wpm) { + engine.setWpm(wpm); + C.wpm.value = wpm; + } + }).catch(function () { /* defaults are fine */ }).then(loadSource); +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/remote.js b/skills/mc-prompter/scripts/server/static/js/remote.js new file mode 100644 index 0000000..5e61a4f --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/remote.js @@ -0,0 +1,170 @@ +/* mc-prompter /remote page glue. + * + * Phone-as-remote. Requires the session token from ?token= in the URL; a + * non-loopback WS connect without a valid token is closed with 4403 and the + * page goes read-only (error panel, no controls). Commands go out as cmd + * frames; the live display is driven by the leader's state frames. + */ +(function () { + 'use strict'; + var MC = window.MC; + + var els = { + app: document.getElementById('app'), + errorPanel: document.getElementById('error-panel'), + errorMsg: document.getElementById('error-msg'), + conn: document.getElementById('conn'), + session: document.getElementById('session'), + elapsed: document.getElementById('st-elapsed'), + remaining: document.getElementById('st-remaining'), + wpm: document.getElementById('st-wpm'), + posFill: document.getElementById('pos-fill'), + btnToggle: document.getElementById('btn-toggle'), + btnRestart: document.getElementById('btn-restart'), + btnSlower: document.getElementById('btn-slower'), + btnFaster: document.getElementById('btn-faster'), + btnPrev: document.getElementById('btn-prev'), + btnNext: document.getElementById('btn-next'), + sectionList: document.getElementById('section-list'), + secCount: document.getElementById('sec-count') + }; + + var token = new URLSearchParams(location.search).get('token') || null; + var sections = []; // [{id, heading}] + var currentSectionId = null; + var docVersion = null; + + function showError(msg) { + els.app.classList.add('hidden'); + els.errorPanel.classList.remove('hidden'); + if (msg) els.errorMsg.textContent = msg; + } + + if (!token) { + showError('This remote needs the session token. Open the exact remote URL ' + + 'shown on the prompter home page (it ends in ?token=...).'); + return; + } + + els.app.classList.remove('hidden'); + + var ws = MC.createWS({ role: 'remote', token: token }); + + // ---------- sections ---------- + + function loadSections() { + return MC.model.fetchSource(token).then(function (src) { + docVersion = src['doc-version']; + sections = []; + var docSections = (src.doc && src.doc.sections) || []; + for (var i = 0; i < docSections.length; i++) { + sections.push({ + id: docSections[i].id, + heading: docSections[i].heading || (i === 0 ? 'Preamble' : 'Untitled section') + }); + } + renderSections(); + }).catch(function (err) { + if (err.status === 401 || err.status === 403) { + showError('The session token was rejected. Grab a fresh remote URL from the home page.'); + } + }); + } + + function renderSections() { + var list = els.sectionList; + while (list.firstChild) list.removeChild(list.firstChild); + els.secCount.textContent = sections.length ? '(' + sections.length + ')' : ''; + for (var i = 0; i < sections.length; i++) { + (function (sec) { + var li = document.createElement('li'); + li.dataset.sid = sec.id; + li.textContent = sec.heading; + li.addEventListener('click', function () { + sendJumpSection(sec.id); + }); + list.appendChild(li); + })(sections[i]); + } + highlightSection(); + } + + function highlightSection() { + var items = els.sectionList.children; + for (var i = 0; i < items.length; i++) { + items[i].classList.toggle('current', items[i].dataset.sid === currentSectionId); + } + } + + function sectionIndex(id) { + for (var i = 0; i < sections.length; i++) { + if (sections[i].id === id) return i; + } + return -1; + } + + // Send a jump and update currentSectionId optimistically so rapid prev or + // next taps chain from the target instead of re-targeting the same section + // while the leader's next state frame is still in flight. The next state + // frame reconciles the real position. + function sendJumpSection(id) { + ws.cmd('jump-section', id); + currentSectionId = id; + highlightSection(); + } + + function jumpRelativeSection(delta) { + if (!sections.length) return; + var idx = sectionIndex(currentSectionId); + var target = idx < 0 ? 0 : Math.min(Math.max(idx + delta, 0), sections.length - 1); + sendJumpSection(sections[target].id); + } + + // ---------- live state ---------- + + function applyState(msg) { + currentSectionId = msg.section || null; + els.elapsed.textContent = MC.model.fmtClock(msg.elapsed); + els.remaining.textContent = MC.model.fmtClock(msg.remaining); + els.wpm.textContent = String(msg.wpm !== undefined ? msg.wpm : '--'); + els.posFill.style.width = (Math.min(Math.max(Number(msg.position) || 0, 0), 1) * 100) + '%'; + var playing = !!msg.playing; + els.btnToggle.textContent = playing ? 'pause' : 'play'; + els.btnToggle.classList.toggle('playing', playing); + if (msg.countdown !== null && msg.countdown !== undefined) { + els.btnToggle.textContent = 'in ' + Math.ceil(msg.countdown) + 's'; + } + highlightSection(); + } + + // ---------- wiring ---------- + + ws.onStatus(function (s) { + els.conn.classList.toggle('on', s.connected); + els.session.textContent = s.session || ''; + if (s.rejected) { + showError('The session token was rejected. Grab a fresh remote URL from the home page.'); + return; + } + if (s.connected && s.snapshot) applyState(s.snapshot); + }); + + ws.on('state', applyState); + ws.on('doc-updated', function (msg) { + if (msg['doc-version'] !== docVersion) loadSections(); + }); + + els.btnToggle.addEventListener('click', function () { ws.cmd('toggle'); }); + els.btnRestart.addEventListener('click', function () { ws.cmd('restart'); }); + els.btnSlower.addEventListener('click', function () { ws.cmd('speed-delta', -2); }); + els.btnFaster.addEventListener('click', function () { ws.cmd('speed-delta', 2); }); + els.btnPrev.addEventListener('click', function () { jumpRelativeSection(-1); }); + els.btnNext.addEventListener('click', function () { jumpRelativeSection(1); }); + + // ---------- boot ---------- + + loadSections(); + MC.model.fetchState(token).then(function (state) { + if (state && state.snapshot) applyState(state.snapshot); + }).catch(function () { /* WS state frames will fill in */ }); +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/scroll.js b/skills/mc-prompter/scripts/server/static/js/scroll.js new file mode 100644 index 0000000..15017b1 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/scroll.js @@ -0,0 +1,309 @@ +/* mc-prompter scroll engine (classic script, attaches to window.MC). + * + * requestAnimationFrame based smooth scroll of a scrollable surface. + * + * Speed model (Phase A contract): + * px/s = wpm / 60 * (scrollable height / word count) + * where scrollable height = scrollHeight - clientHeight. The ratio is read + * fresh every frame, so font, size, margin, line-height changes and window + * resizes are picked up automatically without an explicit recompute call. + * + * Position is reported as scrollTop / (scrollHeight - clientHeight), 0..1. + * + * Modes: + * manual wpm is whatever the user set (live +/- adjustment) + * timed given a total duration, wpm is continuously re-derived as + * remaining words / remaining minutes (recomputed every frame, + * which covers resume and jumps); drift from plan is exposed as + * elapsed - position * totalSeconds (positive means behind plan) + * + * Countdown: beginCountdown(seconds) holds the scroll and counts down, then + * flips to playing. The page renders the big digits from view().countdown. + * + * Phase B seam: voice-follow will drive targetPos via an easing setter + * instead of the constant-rate integration in frame(); the surface, position + * math, and word spans stay identical. + */ +(function () { + 'use strict'; + window.MC = window.MC || {}; + + var WPM_MIN = 10; + var WPM_MAX = 1200; + + function createEngine(opts) { + // opts: { surface: element, getWordCount: fn -> int, + // onFrame: fn(view), onChange: fn(reason), onFinish: fn } + var surface = opts.surface; + + var st = { + playing: false, + mode: 'manual', // 'manual' | 'timed' + wpm: 150, // manual-mode wpm + totalSeconds: null, // timed-mode plan length + elapsed: 0, // seconds of play time accumulated + countdownLeft: 0, + finished: false + }; + + var pos = 0; // float scroll position (scrollTop rounds) + var rafId = null; + var lastTs = null; + var selfScroll = false; // guards the scroll listener against our own writes + + function scrollable() { + return Math.max(surface.scrollHeight - surface.clientHeight, 1); + } + + function position() { + var p = pos / scrollable(); + return Math.min(Math.max(p, 0), 1); + } + + function remainingWords() { + var wc = opts.getWordCount() || 0; + return wc * (1 - position()); + } + + function currentWpm() { + if (st.mode === 'timed' && st.totalSeconds) { + var remS = Math.max(st.totalSeconds - st.elapsed, 1); + var w = remainingWords() / (remS / 60); + return Math.min(Math.max(w, WPM_MIN), WPM_MAX); + } + return st.wpm; + } + + function pxPerSecond() { + var wc = opts.getWordCount(); + if (!wc) return 0; + return currentWpm() / 60 * (scrollable() / wc); + } + + function remainingSeconds() { + if (st.mode === 'timed' && st.totalSeconds) { + return Math.max(st.totalSeconds - st.elapsed, 0); + } + var w = currentWpm(); + if (!w) return null; + return remainingWords() / w * 60; + } + + function driftSeconds() { + if (st.mode !== 'timed' || !st.totalSeconds) return null; + return st.elapsed - position() * st.totalSeconds; + } + + function view() { + return { + playing: st.playing, + position: position(), + wpm: Math.round(currentWpm()), + mode: st.mode, + totalSeconds: st.totalSeconds, + elapsed: st.elapsed, + remaining: remainingSeconds(), + countdown: st.countdownLeft > 0 ? st.countdownLeft : null, + drift: driftSeconds(), + finished: st.finished + }; + } + + function setScrollTop(v) { + selfScroll = true; + surface.scrollTop = v; + // The scroll event fires async; clear the guard on the next frame. + requestAnimationFrame(function () { selfScroll = false; }); + } + + function changed(reason) { + if (opts.onChange) opts.onChange(reason, view()); + } + + function frame(ts) { + rafId = requestAnimationFrame(frame); + if (lastTs === null) { lastTs = ts; return; } + var dt = Math.min((ts - lastTs) / 1000, 0.25); + lastTs = ts; + + if (st.countdownLeft > 0) { + st.countdownLeft = Math.max(st.countdownLeft - dt, 0); + if (st.countdownLeft === 0) { + st.playing = true; + changed('countdown-done'); + } + } else if (st.playing) { + st.elapsed += dt; + pos += pxPerSecond() * dt; + var max = scrollable(); + if (pos >= max) { + pos = max; + setScrollTop(pos); + st.playing = false; + st.finished = true; + changed('finished'); + if (opts.onFinish) opts.onFinish(); + } else { + setScrollTop(pos); + } + } + + if (opts.onFrame) opts.onFrame(view()); + + if (!st.playing && st.countdownLeft <= 0) stopLoop(); + } + + function startLoop() { + if (rafId === null) { + lastTs = null; + rafId = requestAnimationFrame(frame); + } + } + + function stopLoop() { + if (rafId !== null) { + cancelAnimationFrame(rafId); + rafId = null; + lastTs = null; + } + } + + var engine = { + view: view, + isPlaying: function () { return st.playing || st.countdownLeft > 0; }, + + play: function () { + if (st.playing) return; + st.finished = false; + st.countdownLeft = 0; + st.playing = true; + startLoop(); + changed('play'); + }, + + // Countdown, then play. seconds <= 0 plays immediately. + beginCountdown: function (seconds) { + var s = Number(seconds) || 0; + if (s <= 0) { engine.play(); return; } + st.finished = false; + st.playing = false; + st.countdownLeft = s; + startLoop(); + changed('countdown'); + }, + + pause: function () { + var was = st.playing || st.countdownLeft > 0; + st.playing = false; + st.countdownLeft = 0; + if (was) changed('pause'); + }, + + toggle: function (countdownSeconds) { + if (engine.isPlaying()) { + engine.pause(); + } else if (position() < 0.001 && countdownSeconds > 0) { + engine.beginCountdown(countdownSeconds); + } else { + engine.play(); + } + }, + + restart: function () { + pos = 0; + setScrollTop(0); + st.elapsed = 0; + st.playing = false; + st.countdownLeft = 0; + st.finished = false; + changed('restart'); + }, + + setWpm: function (n) { + st.wpm = Math.min(Math.max(Math.round(Number(n) || st.wpm), WPM_MIN), WPM_MAX); + changed('speed'); + }, + + deltaWpm: function (d) { + engine.setWpm(st.wpm + (Number(d) || 0)); + }, + + getManualWpm: function () { return st.wpm; }, + + // mode: 'manual' | 'timed'; totalMinutes required for timed. + setMode: function (mode, totalMinutes) { + if (mode === 'timed' && Number(totalMinutes) > 0) { + st.mode = 'timed'; + st.totalSeconds = Number(totalMinutes) * 60; + } else { + st.mode = 'manual'; + st.totalSeconds = null; + } + changed('mode'); + }, + + // Jump to an absolute pixel offset within the surface. + jumpToPx: function (px) { + pos = Math.min(Math.max(px, 0), scrollable()); + setScrollTop(pos); + st.finished = false; + changed('jump'); + }, + + // Jump to a 0..1 position ratio. + setPositionRatio: function (r) { + engine.jumpToPx((Number(r) || 0) * scrollable()); + }, + + getPositionRatio: position, + getScrollTop: function () { return pos; }, + + // Adopt an externally caused scrollTop (manual drag or touch while + // paused) so the next play resumes from where the user left the view. + adoptScrollTop: function () { + if (selfScroll || st.playing || st.countdownLeft > 0) return; + pos = surface.scrollTop; + st.finished = false; + changed('scroll'); + }, + + // Re-clamp after layout changes; the page keeps the position ratio + // stable across settings changes by capturing it before and restoring + // after (see prompt.js applyDisplay). + recalc: function () { + pos = Math.min(pos, scrollable()); + setScrollTop(pos); + }, + + // Seed from a state snapshot (welcome payload or leader promotion). + seed: function (snap) { + if (!snap) return; + if (typeof snap.wpm === 'number' && snap.mode !== 'timed') st.wpm = snap.wpm; + if (snap.mode === 'timed') { + var total = (Number(snap.elapsed) || 0) + (Number(snap.remaining) || 0); + st.mode = 'timed'; + st.totalSeconds = total > 0 ? total : null; + if (!st.totalSeconds) st.mode = 'manual'; + } else { + st.mode = 'manual'; + } + st.elapsed = Number(snap.elapsed) || 0; + if (typeof snap.position === 'number') { + pos = Math.min(Math.max(snap.position, 0), 1) * scrollable(); + setScrollTop(pos); + } + st.finished = false; + st.countdownLeft = 0; + st.playing = !!snap.playing; + if (st.playing) startLoop(); + changed('seed'); + }, + + destroy: stopLoop + }; + + return engine; + } + + window.MC.createEngine = createEngine; +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/settings.js b/skills/mc-prompter/scripts/server/static/js/settings.js new file mode 100644 index 0000000..079b3c2 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/settings.js @@ -0,0 +1,100 @@ +/* mc-prompter display settings (classic script, attaches to window.MC). + * + * Persisted per device in localStorage under the key "mc-prompter-display". + * Schema (kebab-case keys, all optional in storage, defaults below): + * mirror-h bool horizontal flip of the scroll surface + * mirror-v bool vertical flip (both may combine) + * font-family string CSS font-family (system stacks or free text) + * font-size number px + * text-color string CSS color + * background-color string CSS color + * margin-percent number horizontal margin, percent of surface width + * line-height number unitless multiplier + * eyeline-percent number marker position, percent from viewport top + * eyeline-style string "line" | "arrow" + * countdown-seconds number countdown before scroll starts (0 disables) + * hide-takes bool hide TAKE paragraphs entirely + * show-invented bool show the invented badge styling + * + * Server config defaults (from GET /api/state .config) may be passed to + * load() as overrides; stored per-device values still win over them. + */ +(function () { + 'use strict'; + window.MC = window.MC || {}; + + var KEY = 'mc-prompter-display'; + + var DEFAULTS = { + 'mirror-h': false, + 'mirror-v': false, + 'font-family': 'system-ui, "Segoe UI", Roboto, "Helvetica Neue", Arial, sans-serif', + 'font-size': 64, + 'text-color': '#f2f2f2', + 'background-color': '#000000', + 'margin-percent': 12, + 'line-height': 1.5, + 'eyeline-percent': 33, + 'eyeline-style': 'line', + 'countdown-seconds': 3, + 'hide-takes': false, + 'show-invented': true + }; + + // A few known-safe offline font stacks for the settings drawer select. + var FONT_STACKS = [ + { label: 'System sans', value: DEFAULTS['font-family'] }, + { label: 'Georgia serif', value: 'Georgia, "Times New Roman", serif' }, + { label: 'Verdana wide', value: 'Verdana, Geneva, Tahoma, sans-serif' }, + { label: 'Trebuchet', value: '"Trebuchet MS", "Segoe UI", sans-serif' }, + { label: 'Monospace', value: 'ui-monospace, "SF Mono", Menlo, Consolas, monospace' } + ]; + + function load(overrides) { + var out = {}; + var k; + for (k in DEFAULTS) { + if (Object.prototype.hasOwnProperty.call(DEFAULTS, k)) out[k] = DEFAULTS[k]; + } + if (overrides) { + for (k in overrides) { + if (Object.prototype.hasOwnProperty.call(DEFAULTS, k) && + overrides[k] !== null && overrides[k] !== undefined) { + out[k] = overrides[k]; + } + } + } + try { + var raw = localStorage.getItem(KEY); + if (raw) { + var saved = JSON.parse(raw); + for (k in saved) { + if (Object.prototype.hasOwnProperty.call(DEFAULTS, k)) out[k] = saved[k]; + } + } + } catch (e) { + // Corrupted storage: fall back to defaults silently. + } + return out; + } + + function save(settings) { + var out = {}; + for (var k in DEFAULTS) { + if (Object.prototype.hasOwnProperty.call(settings, k)) out[k] = settings[k]; + } + try { + localStorage.setItem(KEY, JSON.stringify(out)); + } catch (e) { + // Storage full or blocked: settings stay session-only. + } + } + + window.MC.settings = { + KEY: KEY, + DEFAULTS: DEFAULTS, + FONT_STACKS: FONT_STACKS, + load: load, + save: save + }; +})(); diff --git a/skills/mc-prompter/scripts/server/static/js/ws.js b/skills/mc-prompter/scripts/server/static/js/ws.js new file mode 100644 index 0000000..a14d0fb --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/js/ws.js @@ -0,0 +1,145 @@ +/* mc-prompter WebSocket client (classic script, attaches to window.MC). + * + * Protocol (Phase A contract): + * -> hello {type:"hello", role:"prompt"|"remote"|"overlay"|"home", token?} + * <- welcome {type:"welcome", session, leader, "doc-version", snapshot} + * <- role {type:"role", leader:true} leader promotion + * <> cmd {type:"cmd", cmd, value?, from?} relayed by server to all + * <> state {type:"state", position, section, playing, wpm, mode, + * elapsed, remaining, countdown} leader prompt -> everyone + * <- doc-updated {type:"doc-updated", "doc-version"} clients refetch /api/source + * <- error {type:"error", message} + * + * Reconnects with exponential backoff. Close code 4403 means the token was + * rejected; we stop retrying and flag state.rejected so pages can go read-only. + */ +(function () { + 'use strict'; + window.MC = window.MC || {}; + + function createWS(opts) { + // opts: { role: string, token: string|null } + var listeners = {}; // message type -> [fn] + var statusFns = []; + var ws = null; + var closedByUser = false; + var backoff = 500; + var BACKOFF_MAX = 8000; + + var state = { + connected: false, + leader: false, + session: null, + docVersion: null, + snapshot: null, // last state snapshot delivered in welcome + rejected: false // token rejected (close 4403); no further retries + }; + + function emitStatus() { + for (var i = 0; i < statusFns.length; i++) statusFns[i](state); + } + + function dispatch(msg) { + if (msg.type === 'welcome') { + state.session = msg.session || null; + state.leader = !!msg.leader; + state.docVersion = msg['doc-version']; + state.snapshot = msg.snapshot || null; + emitStatus(); + } else if (msg.type === 'role') { + state.leader = !!msg.leader; + emitStatus(); + } + var fns = (listeners[msg.type] || []).concat(listeners['*'] || []); + for (var i = 0; i < fns.length; i++) { + try { fns[i](msg); } catch (e) { /* one bad listener never kills dispatch */ } + } + } + + function connect() { + if (closedByUser || state.rejected) return; + var proto = location.protocol === 'https:' ? 'wss://' : 'ws://'; + var url = proto + location.host + '/ws'; + if (opts.token) url += '?token=' + encodeURIComponent(opts.token); + try { + ws = new WebSocket(url); + } catch (e) { + scheduleReconnect(); + return; + } + ws.onopen = function () { + backoff = 500; + state.connected = true; + var hello = { type: 'hello', role: opts.role }; + if (opts.token) hello.token = opts.token; + ws.send(JSON.stringify(hello)); + emitStatus(); + }; + ws.onmessage = function (ev) { + if (typeof ev.data !== 'string') return; // binary frames reserved for Phase B + var msg; + try { msg = JSON.parse(ev.data); } catch (e) { return; } + if (msg && msg.type) dispatch(msg); + }; + ws.onclose = function (ev) { + state.connected = false; + state.leader = false; + if (ev && ev.code === 4403) { + state.rejected = true; + emitStatus(); + return; + } + emitStatus(); + scheduleReconnect(); + }; + ws.onerror = function () { + try { ws.close(); } catch (e) { /* already closing */ } + }; + } + + function scheduleReconnect() { + if (closedByUser || state.rejected) return; + setTimeout(connect, backoff); + backoff = Math.min(backoff * 2, BACKOFF_MAX); + } + + connect(); + + return { + state: state, + + // on(type, fn): subscribe to a message type; '*' catches everything. + on: function (type, fn) { + (listeners[type] = listeners[type] || []).push(fn); + }, + + // onStatus(fn): connection or role changes; called once immediately. + onStatus: function (fn) { + statusFns.push(fn); + fn(state); + }, + + send: function (obj) { + if (ws && ws.readyState === WebSocket.OPEN) { + ws.send(JSON.stringify(obj)); + return true; + } + return false; + }, + + // cmd(name, value): send a command frame per the protocol. + cmd: function (name, value) { + var m = { type: 'cmd', cmd: name }; + if (value !== undefined && value !== null) m.value = value; + return this.send(m); + }, + + close: function () { + closedByUser = true; + if (ws) { try { ws.close(); } catch (e) { /* noop */ } } + } + }; + } + + window.MC.createWS = createWS; +})(); diff --git a/skills/mc-prompter/scripts/server/static/overlay.html b/skills/mc-prompter/scripts/server/static/overlay.html new file mode 100644 index 0000000..96c2e1c --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/overlay.html @@ -0,0 +1,21 @@ + + + + + + + mc-prompter | overlay + + + + + + +
mc-prompter session connected
+
+ + + + + diff --git a/skills/mc-prompter/scripts/server/static/prompt.html b/skills/mc-prompter/scripts/server/static/prompt.html new file mode 100644 index 0000000..5d7965f --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/prompt.html @@ -0,0 +1,191 @@ + + + + + + + mc-prompter | prompt + + + + + +
+
+
+
+
+
+ +
+
+ + + + +
+
+ 00:00 / --:-- + +
+
+ 150 wpm + manual + + + +
+
+ + + + + + + + + + + +
+ + + + + + + + diff --git a/skills/mc-prompter/scripts/server/static/remote.html b/skills/mc-prompter/scripts/server/static/remote.html new file mode 100644 index 0000000..38c7762 --- /dev/null +++ b/skills/mc-prompter/scripts/server/static/remote.html @@ -0,0 +1,55 @@ + + + + + + + mc-prompter | remote + + + + + + + + + + + + + + From 6348906b9a5abd8ce25a41d2a1a86f557990bc60 Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 19:16:43 -0500 Subject: [PATCH 5/6] Register mc-prompter skill: SKILL.md, customize.toml, help row, marketplace entry --- .claude-plugin/marketplace.json | 1 + skills/mc-prompter/SKILL.md | 33 +++++++++++++++++++++++++++++++ skills/mc-prompter/customize.toml | 26 ++++++++++++++++++++++++ skills/module-help.csv | 1 + 4 files changed, 61 insertions(+) create mode 100644 skills/mc-prompter/SKILL.md create mode 100644 skills/mc-prompter/customize.toml diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 29898a6..3c165f4 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -37,6 +37,7 @@ "./skills/mc-graphics", "./skills/mc-assets", "./skills/mc-audio", + "./skills/mc-prompter", "./skills/mc-package", "./skills/mc-stream-pack", "./skills/mc-retro" diff --git a/skills/mc-prompter/SKILL.md b/skills/mc-prompter/SKILL.md new file mode 100644 index 0000000..e90e7cc --- /dev/null +++ b/skills/mc-prompter/SKILL.md @@ -0,0 +1,33 @@ +--- +name: mc-prompter +description: Browser teleprompter for the record stage and standalone shows. A service skill like mc-audio, no stage, no gate, no project.json state. Launch a local prompter server, feed it the project script.md or any text, and the creator records at their own pace with a phone remote over LAN. Classic teleprompter only in this version; voice-follow and producer mode are planned tiers. +--- + +# mc-prompter + +The record stage is creator-owned; this skill hands the creator a teleprompter for it. It is a service skill: it owns no stage, stops at no gate, and writes no project state. It launches a local web server that serves a fullscreen prompter display, a phone remote, a home page for loading and editing the script, and an OBS overlay placeholder. Everything runs offline on the creator's machine; no models, no downloads, no external requests. + +## Steps + +1. Load this skill's own surface (`uv run {project-root}/_bmad/scripts/resolve_customization.py --skill {skill-root}`; run `{workflow.activation_steps_prepend}` now, `{workflow.activation_steps_append}` after this step, and hold `{workflow.persistent_facts}` as standing context). Take the default port from `[prompter] port`. If a studio config exists, also load it (`uv run {project-root}/_bmad/scripts/resolve_config.py --project-root {project-root} --key modules.manticore`) and take `[prompter] port` and `[owner] wpm` when present; a missing studio config is fine here, unlike the stage skills, because standalone shows need no studio. +2. Locate the script. Inside a pipeline project ("record with the teleprompter"), it is the project's `script.md` under the project folder. Standalone, it is any file path the creator names, markdown or plain text. No file at all is also valid: launch without `--script` and the creator pastes text on the home page. +3. Launch: `uv run {skill-root}/scripts/run_prompter.py --script ` with `--port ` when the config or the creator sets one and `--owner-wpm ` when `[owner] wpm` is known. Add `--lan` only when the creator wants the phone remote or a tablet display; it binds the LAN and on Windows triggers a firewall consent dialog. The launcher probes the port, prints the local URL, the remote URL with its session token, and the session file path, then keeps the server running until Ctrl-C. +4. Give the creator the URLs from the launcher output: the prompt page for the recording display, and the remote URL (token included) to open on a phone when `--lan` is on. Briefly explain the pages: `/` is home (load a file, paste text, edit in place, copy the remote URL), `/prompt` is the fullscreen scroller with keyboard controls and a settings drawer (press `?` there for all shortcuts), `/remote` is the phone controller, `/overlay` is an OBS browser-source placeholder for now. +5. Explain what the display does with pipeline markers: paragraphs carrying a `[TAKE ...]` marker render dimmed with a "have it already" badge because that line was already spoken well in the interview footage, and a toggle hides them entirely; sentences flagged `[INVENTED]` get a subtle badge, toggleable off; other bracketed text renders as a dimmed note and is never counted in timing. +6. If the creator edits the script from the home page and saves, the server first copies the current file to a timestamped backup under the temp session directory, then writes the edit back to the source file, so the prompted text and the pipeline artifact never silently diverge. "Session only" applies the edit without touching the file. +7. When the creator asks for voice-follow scrolling or producer mode, say plainly that those are planned tiers that have not landed yet and the classic prompter is what works today. Never pretend they work and never improvise a substitute. + +## Rules + +- The prompter never advances pipeline state; when a pipeline project is being recorded, the record stage remains the creator's, and mc-pipeline stays the source of truth. +- The server binds localhost by default; `--lan` is opt-in and the remote URL carries a per-session token so a random LAN device cannot drive the prompter mid-show. +- Loading files by path and saving edits work only from the server machine, never from a LAN device. +- Stop and relay the launcher's guidance when it exits nonzero: a missing or unreadable script path, an explicit port already held by another session, or a server that failed to start are all creator-facing problems, not things to retry silently. + +## Checklist + +- The port and wpm came from config when a studio config exists; defaults otherwise. +- The creator got both URLs (prompt page, remote with token) and knows the pages. +- Take and invented markers were explained if the script contains them. +- Any in-place save was backed up first (the server does this; confirm the backup path in its response). +- No planned tier was presented as working. diff --git a/skills/mc-prompter/customize.toml b/skills/mc-prompter/customize.toml new file mode 100644 index 0000000..45e4e9d --- /dev/null +++ b/skills/mc-prompter/customize.toml @@ -0,0 +1,26 @@ +# DO NOT EDIT -- overwritten on every update. +# +# Customization surface for mc-prompter. +# Override files (not edited here): +# {project-root}/_bmad/custom/mc-prompter.toml (team) +# {project-root}/_bmad/custom/mc-prompter.user.toml (personal) +# +# Studio-wide config (owner, paths, prompter defaults) lives in +# [modules.manticore] in {project-root}/_bmad/custom/config.toml, +# maintained by mc-setup. This file holds only this skill's own surface. + +[workflow] + +# Steps to run before / after the standard activation. +activation_steps_prepend = [] +activation_steps_append = [] + +# Persistent facts held for the whole run: literal sentences, or +# "file:{project-root}/..." paths whose contents are loaded as facts. +persistent_facts = [] + +[prompter] + +# Default port for the prompter server; the studio config's +# [prompter] port overrides this when set. +port = 8770 diff --git a/skills/module-help.csv b/skills/module-help.csv index 604048e..90aebbf 100644 --- a/skills/module-help.csv +++ b/skills/module-help.csv @@ -13,6 +13,7 @@ BMad Manticore,mc-graphics,Build Graphics,GX,"Execute the approved beat table in BMad Manticore,mc-assets,Farm Assets,FA,"Source and farm the stills and b-roll the beat table calls for through registered CLI tools (metered APIs opt-in), real verified imagery first.",,,3-graphics,mc-beats,mc-package,false,projects-path,*/assets/manifest.json BMad Manticore,mc-audio,Farm Sound,AU,"Service skill, no stage or gate: local-first TTS narration and two-host dialogue (Kokoro-82M), instrumental beds (MusicGen-small), SFX (AudioLDM2). Called from graphics, stream packs, and voiceover narration, or directly.",,,anytime,,,false,,*/manifest.json BMad Manticore,mc-ograf,OGraf Graphics,OG,"Service skill: editable broadcast graphics where the target supports them (DaVinci Resolve 21+ editor lane, OBS/SPX-GC live lane). Everyone else gets baked alpha.",,,anytime,,,false,, +BMad Manticore,mc-prompter,Teleprompter,TP,"Service skill, no stage or gate: browser teleprompter for the record stage and standalone shows. Fullscreen mirrored display, phone remote over LAN with a session token, script.md marker handling, edit in place with backups. Voice-follow and producer mode are planned tiers.",,,anytime,,,false,session URLs printed by the launcher, BMad Manticore,mc-package,Package,PK,"Titles, thumbnails verified at 120px, description, CTAs, dual-timeline chapters, series A/B pairs, live-event mode. May start any time after gate 1; offer it during dead time between stages.",,,4-package,mc-outline,mc-retro,false,projects-path,*/packaging/* BMad Manticore,mc-stream-pack,Stream Pack,LS,"A complete branded livestream asset pack for OBS (scenes, stinger, lower thirds) from brand tokens; the livestream-pack format lane.",,,anytime,,,false,projects-path,*/graphics/scenes/* BMad Manticore,mc-retro,Retro,RT,"After publishing: one round of notes edits the format profile, the bibles, and the brand files so the next video starts smarter, then the post-publish wrap.",,,5-wrap,mc-package,,false,brand-path,production-bible.md From e67c16c788e1ca742ccdadcab7f0271e72bc0af9 Mon Sep 17 00:00:00 2001 From: Brian Madison Date: Thu, 9 Jul 2026 20:28:45 -0500 Subject: [PATCH 6/6] Remove plan working document before PR --- PLAN-mc-prompter.md | 300 -------------------------------------------- 1 file changed, 300 deletions(-) delete mode 100644 PLAN-mc-prompter.md diff --git a/PLAN-mc-prompter.md b/PLAN-mc-prompter.md deleted file mode 100644 index ecfecc5..0000000 --- a/PLAN-mc-prompter.md +++ /dev/null @@ -1,300 +0,0 @@ -# mc-prompter: teleprompter and AI producer, implementation plan - -Working document for the feat-prompter branch. Not intended to merge; it guides the build and gets deleted (or distilled into skill references) before the PR. Written 2026-07-09 from a recon pass over the module and external research on real-time local ASR, teleprompter prior art, and live cueing UX. Revised the same day after an adversarial three-lens review (technical feasibility, module conventions, product); the review's fixes are folded in throughout and marked where they changed a decision. - -## Decisions record (2026-07-09, BMad) - -- LLM runtime for producer mode: Ollama is the default provider of a new `[llm]` config lane, with the standard provider-ladder pattern (local-first, other rungs opt-in later). -- Scope: all three tiers built on this one branch, in phases, with the classic teleprompter working end to end first. -- Skill shape: one service skill, `mc-prompter`, following the mc-audio pattern (no stage, no gate, no project.json state). The `record` stage stays creator-owned. -- Cue channels: visual-only in v1. The kokoro spoken tier is designed here but ships as a fast-follow behind a config flag. -- Landing (approved 2026-07-09): stacked PRs. PR 1 = Phase A plus minimal docs and the help row, PR 2 = Phase B, PR 3 = Phases C+D. All developed on this branch. - -## Product overview: three tiers - -Tier 1 is a classic teleprompter with the full standard feature set, no AI, no model downloads. Tier 2 is voice-follow: the scroll tracks the speaker through a known script using streaming local ASR plus a deterministic alignment algorithm, no LLM. Tier 3 is producer mode: a rundown file (duration, ordered talking points, intro, wrap), a rolling transcript, and a small local LLM that keeps the speaker on track with rate-limited visual cues. - -Each tier is progressive enhancement over the one below it. Tier 1 works on any machine with only the module installed. Tier 2 requires the prompter-lab workspace (ASR models, consent-gated download). Tier 3 additionally requires a running Ollama. - -Everything chosen is cross-platform (Windows, macOS, Linux): browser UI, browser mic capture, sherpa-onnx, silero VAD, kokoro-onnx, Ollama. This is the module's first fully cross-platform lane, and the sherpa-onnx dependency incidentally opens a path to the cross-platform transcription lane already on TODO.md. - -## Architecture - -### Skill shape - -New folder `skills/mc-prompter/`: - -``` -skills/mc-prompter/ - SKILL.md service skill: what it does, how to launch, tier gating - customize.toml [workflow] block + prompter defaults - references/ - cueing.md the cue design contract: escalation ladder, budget, vocabulary, - replan rules, coverage semantics, headphones/AEC dependency - rundown-spec.md the rundown.md file format specification (time math included) - scripts/ - ensure_workspace.py prompter-lab builder (mirrors mc-audio's, consent-gated) - run_prompter.py stdlib launcher: validates workspace, port probe, launches server - server/ the application package (both launch paths run python -m server.main) - __init__.py - main.py aiohttp app: HTTP + WebSocket + static UI; lazy tier-2/3 imports - asr.py sherpa-onnx streaming recognizer + silero VAD (workspace only) - align.py script-follow alignment engine (pure stdlib) - producer.py rundown state machine, replanner, cue engine, Ollama client - rundown.py rundown.md parser (pure stdlib) - script_ingest.py script.md / markdown / plain text ingestion (pure stdlib) - static/ vanilla HTML/JS/CSS, no build step - tests/ self-running test-*.py files (convention below) + fixtures/ -``` - -Conventions honored: SKILL.md frontmatter with exactly `name` and `description`; scripts invoked only via `uv run` with PEP 723 headers; scripts take explicit resolved arguments and do no config discovery; the skill reads only its own folder, `_bmad/scripts/`, and project files; nothing user-specific ships in the module; config keys kebab-case. - -### The prompter-lab workspace - -Heavy dependencies live in a persistent venv at `{engines-path}/prompter-lab`, exactly like audio-lab: - -``` -/.venv/ aiohttp (same pin as the PEP 723 header), numpy, - sherpa-onnx (pinned), soundfile -/models/ ASR models (below); later the kokoro pair -/out/ session artifacts (take logs, session transcripts, script backups) -``` - -`ensure_workspace.py --check` exits 0/4 like mc-audio's; the build asks consent before any download. Tier 1 does not require the workspace at all: `run_prompter.py` launches the server with ASR disabled when the workspace is absent, using `uv run` with aiohttp as a PEP 723 dependency. When the workspace exists, the launcher runs the venv interpreter instead, resolved portably (`.venv/bin/python` on POSIX, `.venv\Scripts\python.exe` on Windows). - -Two-launch-path discipline (review finding): - -- Both paths execute the server identically as `python -m server.main` from the scripts directory, so intra-package imports resolve the same way in both. A test executes main.py under a bare env with only aiohttp and asserts the tier-1 routes come up. -- main.py imports asr.py (and anything touching numpy/soundfile) lazily, only inside the workspace-present branch. align.py, rundown.py, and script_ingest.py stay pure stdlib. -- The aiohttp version is pinned identically in the PEP 723 header and the workspace venv, bumped together. - -Port handling (review finding): default port 8770 (mc-ograf's ephemeral verifier uses 8771). On launch, probe the port; if occupied, query `/health` (which reports session and script identity), then offer kill-and-replace or auto-increment to the next free port. The chosen port is printed and written to a session file the skill reads back. `--port` overrides. The server binds 127.0.0.1 by default; `--lan` opts into 0.0.0.0 for the phone remote and tablet displays, and the docs note the Windows Firewall consent dialog this triggers. The home page shows the LAN URL only when it is actually reachable. - -ASR models downloaded into `models/`: - -- Primary: sherpa-onnx export of nvidia nemotron-speech-streaming-en-0.6b, int8 (the streaming sibling of the parakeet family; NVIDIA Open Model License, commercial use permitted). Chunk setting 560 ms as the default latency point. -- Fallback for low-end hardware: a small streaming zipformer English export (Apache-2.0), selectable via config. -- VAD: silero VAD via sherpa-onnx's built-in VoiceActivityDetector (one dependency covers both). - -### Server and UI - -One aiohttp process serving four pages plus a WebSocket: - -- `/` home: pick a source (project script.md, rundown.md, pasted/loaded text), configure, preflight, launch. Shows the reconciled rundown plan before a show starts. -- `/prompt` the prompter display: fullscreen scroll surface, all tier-1 features, the producer rail when tier 3 is active. -- `/remote` phone-as-remote over LAN: play/pause, speed, jump to marker, next/prev section, and in producer mode the point-list override controls. Control only, no mic. The remote URL embeds a per-session token and is presented as a QR code on the home page, so a random LAN device cannot drive the prompter mid-show. -- `/overlay` OBS browser source: transparent background, renders only the ambient rail and cue cards for live shows. - -Concurrency model (review finding, binding): sherpa-onnx decode is CPU-bound and blocking, so it never runs on the event loop. A dedicated ASR thread consumes a bounded queue of PCM frames and publishes recognition events back via `call_soon_threadsafe`. Ollama ticks run as background tasks with hard timeouts and never hold shared state across an await. WebSocket fan-out never awaits decode or LLM work. The server smoke test asserts `/remote` command round-trip latency stays low while the replay harness saturates the ASR path. - -Audio capture and ownership (review findings, binding): - -- Mic capture happens on the server machine. getUserMedia requires a secure context, which http over LAN is not, so a tablet pointed at `/prompt` is display-only. The WS protocol separates the display role from the audio-producer role: the server grants a capture token to exactly one localhost connection; frames without the current token are rejected and the UI names the mic owner. Mic-on-remote-device would require shipping TLS and is explicitly deferred. -- getUserMedia constraints request `echoCancellation: false, noiseSuppression: false, autoGainControl: false, channelCount: 1`; browser speech processing measurably degrades ASR input. The preflight screen reads back `track.getSettings()` and surfaces what was actually applied, since browsers may ignore constraints. When the kokoro spoken tier ships, headphones-only output is what keeps AEC unnecessary; cueing.md records that dependency. -- The AudioContext is created with `{ sampleRate: 16000 }` so the browser resamples; if the browser refuses the rate, a small resampler in the worklet handles the conversion (naive decimation from 44.1 kHz aliases into the speech band). The preflight screen verifies `context.sampleRate`. The worklet ships ~120 ms PCM16 mono frames over the WebSocket. -- Backpressure has a policy at both ends: the browser checks `ws.bufferedAmount` and drops frames past a threshold; the server frame queue is bounded and drops oldest on overflow while raising an "ASR behind real-time" state on the rail. The replay harness asserts queue-depth behavior. - -State flows back to all connected pages over the same WebSocket (scroll position, VAD state, transcript tail, cue events), so the remote and overlay stay in sync with the prompter display. - -UI is vanilla HTML/JS/CSS with no build step. Display settings (mirror flips, font, colors, margins, eyeline position) persist in localStorage per device, with the `[prompter]` config values as defaults; a beam-splitter rig and the operator's browser keep independent settings without reconfiguration each launch. - -### Tier 1: the standard feature checklist - -Scroll and timing: - -- Smooth continuous scroll, speed as WPM with live +/- adjustment (keyboard, wheel, remote) -- Timed mode: give total duration, speed is continuously re-derived (remaining words over remaining time, recomputed on resume and after any jump), with the timer display showing drift from plan -- Pause/resume (spacebar), jump forward/back, jump to marker, restart -- Countdown before scroll starts -- Elapsed and remaining time, estimated read time from word count at the creator's measured wpm (`[owner] wpm` from the studio config when available) - -Display: - -- Mirror flip horizontal, vertical, and both (beam-splitter rigs) -- Font family/size, text and background colors, margins, line height -- Adjustable eyeline/cue marker (position, style) -- Fullscreen; works on a second monitor or a tablet pointed at the same URL (display-only on remote devices, see audio ownership above) -- Per-device settings persistence (localStorage), config defaults underneath - -Script handling: - -- Markdown and plain text; project `script.md` ingestion (below) -- Inline bracket notes render dimmed and are never matched by voice-follow -- Named markers/sections for jumping -- Edit-in-place from the home page between takes; edits write back to the source file with a timestamped backup copied to the workspace `out/` first, so the prompted text and the pipeline artifact never silently diverge - -Remote: - -- Keyboard shortcuts throughout; bluetooth presenters and USB foot pedals work as keyboard emulators for free -- `/remote` phone page over LAN, session-token URL via QR code - -### Script ingestion (pipeline tie-in) - -`script.md` from a project is directly consumable: plain spoken prose. Ingestion handles the two inline marker types: - -- `[INVENTED]` flags render as a subtle badge, toggleable off -- `[TAKE s-s]` lines render dimmed with a "have it already" badge, since those lines were already spoken well in the interview footage and may not need re-recording; a toggle hides them entirely - -The prompter takes a path argument; the skill resolves it from the project when launched inside the pipeline flow ("record with the teleprompter") or accepts any file standalone. - -### Tier 2: voice-follow alignment engine - -Deterministic, no LLM. The prior art (bounded-window Levenshtein prefix matching, the PromptSmart hold-and-re-anchor behavior) consumed Web-Speech-style utterance partials; a streaming transducer behaves differently, so the contract is adapted for sherpa-onnx output (review finding, binding): - -- Input contract: the engine consumes token deltas since the last partial, not whole hypotheses. The last K tokens (K around 3 to 5) of the hypothesis are held provisional because beam search can revise the tail between partials; the anchor commit lags the hypothesis head by K tokens and absorbs revisions. On endpoint detection (`is_final`), the anchor hard-commits and tail-tracking state resets for the next segment. BPE pieces merge to words during normalization before matching. -- Normalize both script tokens and ASR tokens: lowercase, strip punctuation, expand common number/abbreviation forms at index-build time (a small normalization table; "2026" also indexes as "twenty twenty six") -- Maintain a monotonic anchor (last committed script token). On each delta batch, take a lookahead window from the anchor (window size proportional to utterance length plus a constant, on the order of 2x + 10 tokens) and find the window prefix minimizing Levenshtein distance to the pending recognized words; the best prefix end becomes the provisional anchor -- Silence (VAD) produces no partials, so the scroll holds; ad-libs fail to match and the anchor holds until speech re-matches within the window -- Escape hatches: click/tap any word to re-anchor, arrow keys nudge the anchor, and a paragraph-skip gesture jumps the window when the creator deliberately skips content -- The scroll controller eases toward the anchor position rather than jumping; the anchor leads the eyeline by the measured end-to-end latency times the current speaking rate, so the eyeline sits where the speaker actually is, not where ASR last confirmed -- Match state is visible: matched text subtly tinted behind the eyeline, so trust in the tracker is inspectable - -Latency: the honest budget is end-to-end and includes terms the naive sum misses: capture framing (~120 ms) + ASR chunk emission (560 ms configured) + decode compute (hardware-dependent, grows under OBS load) + alignment (<10 ms) + scroll easing (a deliberate time constant). Realistic eyeline-follows-voice latency is 1 to 2 s depending on hardware. The replay harness measures capture-timestamp-to-anchor-update wall time on each target platform, and the measured number feeds the eyeline lead default. No fixed latency claim ships in docs. - -Preflight (review finding, this is the try-once-never-again defense): before any take, the preflight screen enumerates input devices with a picker persisted per machine, shows a live level meter, reports the applied audio constraints and sample rate, and runs a 10-second "read this sentence" tracking test that demonstrates the match-tint following before a real take starts. If recognition confidence is garbage, it fails loudly and names the device in use. - -The engine is a pure-stdlib module. Its fixture suite is built from recorded real partial sequences (the actual partial/final event stream captured from the model over the replay WAV), not hand-written final transcripts, covering: verbatim read, ad-lib excursion and return, skipped paragraph, number/abbreviation mismatch, repeated-phrase script traps, and tail-revision events. - -### Tier 3: producer mode - -Inputs: a rundown file, the rolling transcript from the same ASR stream, and the show clock. - -Show clock semantics (review finding, binding): the plan's arithmetic never keys off server or page start. An explicit GO LIVE control (on `/prompt` and `/remote`) starts the show clock after any pre-roll, and a plan-hold control freezes elapsed time and the state machine during BRB or technical trouble while VAD and the transcript keep running so context is not lost. Both are fixture-tested scenarios. - -The producer is two cooperating parts: - -- A deterministic state machine (code, not LLM). It tracks elapsed time against per-segment budgets, and it re-plans rather than merely flagging lateness: on every tick, remaining show time is redistributed across uncovered segments proportionally to their original budgets, with wrap-minutes protected as a hard reserve. Green/yellow/red state is always computed against the current re-plan, never the original rundown. When redistribution would push any segment below a feasibility floor, the state machine emits a card-tier CUT suggestion ("DROP: point 4, or 90s each"). This replan behavior is the core producer value (a countdown that only turns red is a nag, not a producer) and is specified in cueing.md as part of the binding contract. -- An LLM tick (Ollama): scheduled adaptively, and at VAD pause events, it receives a compact state block (rundown with per-point coverage, the current re-plan, elapsed vs plan, the last ~60 seconds of transcript) and returns structured JSON: proposed coverage transitions, current-topic guess, an optional suggested cue with tier and text, and a one-line reason. Cheap keyword/fuzzy matching runs continuously between ticks as a first-pass coverage signal the LLM confirms or overrides. - -Coverage semantics (review finding, binding): coverage is sticky and monotonic in the state machine; the LLM may only propose uncovered-to-covered transitions, never reversions, so the rail cannot flicker. "Next" is defined as the first uncovered point in rundown order, which stays well-defined when the creator covers points out of order. The human is the final authority: `/remote` gains producer controls, a tappable point list with mark-covered, skip, and make-current, so one tap mid-show rescues any model misjudgment. - -LLM tick budget (review finding, binding): the tick must never starve the ASR thread. Concretely: - -- Requests use Ollama structured outputs (`format` with a JSON schema), `think: false`, temperature 0, a hard `num_predict` cap, and `keep_alive` so the model stays resident -- The prompt keeps a stable prefix (system + rundown first, rolling transcript last) so Ollama prefix caching skips reprocessing -- Cadence is adaptive: the next tick is scheduled at `max(15 s, 3x last tick wall time)`, and ticks are skipped entirely while the ASR queue depth signals CPU pressure -- Default model is `qwen3:4b` where Ollama reports GPU/Metal offload; on CPU-only machines the producer startup check recommends and falls back to a sub-2B tag (`qwen3:1.7b`). Every request carries a hard timeout; a timed-out tick is dropped, not queued - -The cue engine (code) is the final authority on delivery: it applies the density setting, the one-active-cue rule, per-interval budget, and tier gating. The LLM proposes; the state machine disposes. A wrong suggestion costs nothing because rate limiting, tiering, and coverage stickiness are deterministic. - -Visual cue surface (v1, from the cueing research): - -- Ambient tier, always on: a rail showing current point, next point, and a green/yellow/red segment-time state computed against the re-plan (Toastmasters vocabulary), plus overall show progress. No motion, no reading required -- Card tier: a single quiet card ("NEXT: pricing demo", "STRETCH: 4 min left, 1 point to go", "DROP: point 4, or 90s each"), released at pauses, auto-expiring, never stacked -- Attention tier: the card flashes/enlarges for time-critical states ("WRAP", "2:00 OVER"), the one tier allowed to appear mid-sentence -- Vocabulary is the broadcast lexicon (standby, wrap, stretch, hard wrap, time remaining); `references/cueing.md` is the binding contract - -Free-talk support: a rundown with no script body per point is exactly the "5 ideas in this order plus intro and wrap" show. The intro and wrap can carry full scripted text (prompted via tier 1/2) while the middle segments run producer-only. The segment handoff is explicit UI behavior, not hand-waved: when a scripted segment ends (anchor reaches section end, or manual next-section), `/prompt` switches to a large-type rail view showing the current bullet set; entering the next scripted segment switches back to the scroll surface. - -Spoken tier (designed now, shipped later behind `spoken-cues = false`): short formulaic kokoro phrases only ("thirty seconds", "wrap"), synthesized by a persistent kokoro instance (~150 to 300 ms for a short cue on CPU), released only at pauses, hard requirement that output routes to headphones (this is also what keeps browser AEC unnecessary). Never speech-over-speech except a true emergency tier. The mix-minus principle from IFB practice: the speaker must never hear their own voice back. - -### The rundown artifact - -New file format, specified in `references/rundown-spec.md`. The starter template ships through mc-setup's `assets/` into the studio like tokens and format profiles, so Manny and any skill can read it as a project file without crossing skill-folder boundaries; mc-prompter owns the spec and does any template-based drafting, and Manny routes to it (review finding). - -```markdown ---- -show: "Why local models win" -duration-minutes: 30 -cue-density: normal # hands-off | minimal | normal | chatty -wrap-minutes: 3 ---- - -## Intro (3 min) - -Full scripted intro text here, prompted normally. - -## Point 1: The cost argument (5 min) - -- cloud bills compound, local is capex -- the 4090 anecdote - -## Point 2: Latency (5 min) -... - -## Wrap (3 min) - -Scripted wrap text. -``` - -Time math (review finding, in the spec): per-segment minutes are optional; unbudgeted segments split the remaining time evenly. If explicit minutes exceed duration-minutes, duration-minutes wins and the parser warns at load; the home page shows the reconciled plan before the show starts. The parser accepts exactly `(N min)` and `(Nm)` heading suffixes and rejects anything else with a line-numbered error, because hand-written and Manny-drafted rundowns will produce creative variants on day one. Frontmatter `cue-density` overrides the config value (most specific wins). - -Segments with prose bodies prompt as script; segments with only bullets run producer-only. For pipeline projects the file lives at `{projects-path}//rundown.md`; standalone shows pass any path. This artifact also fills the episode-plan gap the livestream formats already reference (the 1.0.x per-episode stream-pack fast-follow needs the same file), and the planned mc-research skill becomes its natural upstream. - -### Config: new studio sub-tables - -Two new tables in `[modules.manticore]`, seeded from mc-setup's `[defaults]`: - -```toml -[defaults.prompter] -workspace = "prompter-lab" # resolved {engines-path}/{prompter.workspace} -asr-provider = "nemotron-streaming" # nemotron-streaming (default) | zipformer-small | none -cue-density = "normal" # hands-off | minimal | normal | chatty -spoken-cues = false # kokoro tier, fast-follow -port = 8770 - -[defaults.llm] -provider = "ollama" # the only implemented rung; others planned, opt-in -model = "qwen3:4b" # any ollama tag; producer falls back to a sub-2B tag on CPU-only -endpoint = "http://localhost:11434" -api-key-env = "" # stays empty for local lanes, pattern-consistent -``` - -mc-setup changes: - -- Add both tables to `[defaults]`, a short optional interview step (offer the teleprompter, ask about producer mode and Ollama only if wanted), and both table names to the step 8 write list -- Migration (review finding, important): do NOT add these tables to the step 1a 0.x classifier list; that list defines what makes a studio 0.x, and adding 1.1 tables to it would mislabel every current 1.0.x studio as 0.x and run the full migration flow on it. Instead add a separate backfill rule alongside it: a config that has the 1.0 tables but is missing `[prompter]` or `[llm]` is a current studio predating the teleprompter; backfill both tables surgically from `[defaults]` and offer the optional prompter interview -- check_deps.py gains an optional `ollama` row (producer mode only) with install pointers -- PIPELINE.md's conventions line mentions the new sub-tables - -The `[llm]` lane follows the enforcement pattern: the producer accepts only `provider = "ollama"` and exits 3 for anything else, so no planned lane ever pretends to work. - -## Module integration checklist - -- `skills/module-help.csv`: one new row with the full 13-column schema (module, skill, display-name, menu-code, description, action, args, phase, preceded-by, followed-by, required, output-location, outputs); phase `anytime`, preceded-by and followed-by left empty per the mc-audio/mc-ograf service-skill precedent -- `.claude-plugin/marketplace.json`: add `./skills/mc-prompter` to the skills array (version bump at release, not in this PR) -- `skills/mc-pipeline/PIPELINE.md`: record stage row gains a sentence noting mc-prompter as the optional tool for the creator-owned record stage; no stage table change, no gate change -- mc-script SKILL.md step 6: the "now record" handoff names the teleprompter option (naming only, no cross-skill read) -- mc-agent (Manny) capabilities: route "teleprompter", "prompt me", "run my show", "producer mode" to mc-prompter; for rundown drafting Manny routes to mc-prompter or works from the studio-installed template, never reads mc-prompter's folder -- docs/user-guide.md: new section; README skill table row; TODO.md: add the spoken-cue fast-follow and the sherpa-onnx transcription-lane opportunity - -## Phasing - -- Phase A, classic teleprompter: workspace-less launch path, aiohttp server with the concurrency skeleton, `/prompt` with the full tier-1 checklist, `/remote` with session token, script.md ingestion with marker handling, home page, settings persistence, port handling. Verifiable end to end with zero models -- Phase B, voice-follow: ensure_workspace.py, browser audio path with the constraint/ownership/backpressure rules, sherpa-onnx + VAD integration on the ASR thread, align.py with the transducer contract and its recorded-partials fixture suite, preflight screen with device picker and tracking test, tracking UX (hold, re-anchor, click-to-anchor, match tinting, eyeline lead) -- Phase C, producer: rundown parser + spec, producer state machine with replanner and cue engine plus time-warped fixtures (running long, running short, out-of-order coverage, GO LIVE / hold), Ollama tick with the budget rules, ambient rail + cards on `/prompt` and `/overlay`, `/remote` producer controls, `[prompter]`/`[llm]` config plumbing, mc-setup + check_deps integration -- Phase D, finish: docs, help catalog row, marketplace entry, Manny routing, take-log hook, cross-platform smoke instructions, changelog - -Landing (approved 2026-07-09): develop everything on this branch, land as stacked PRs: PR 1 = Phase A plus minimal docs and the help row (a complete, useful teleprompter on its own), PR 2 = Phase B, PR 3 = Phases C+D. Each is independently green and valuable, and review feedback on the foundations arrives before the producer is built on top of them. - -## Testing and verification - -Test convention (review finding, matches the quality gate CI): the CI discovers `skills/*/scripts/tests/test-*.py` (hyphenated) and runs each file directly via `uv run`; there is no pytest. Every test file carries its own PEP 723 header, uses stdlib unittest with `unittest.main()`, and runs with no models, no network, no downloads. - -- `test-align.py`: the transducer-contract fixtures (recorded partial sequences committed as small JSON fixtures; the five failure-mode cases plus tail-revision events) -- `test-producer.py`: replanner and cue engine under time-warped scenarios (running long with replan and CUT suggestion, running short with STRETCH, point skipped, out-of-order coverage, GO LIVE and hold semantics, budget/ladder enforcement, coverage stickiness) -- `test-rundown.py`: parser, time-math reconciliation, heading-suffix rejection with line numbers -- `test-script_ingest.py`: marker handling, bracket-note exclusion -- `test-server.py`: aiohttp smoke via its own PEP 723 aiohttp dependency; routes up under the bare tier-1 env; WS state fan-out; remote-latency-under-ASR-load assertion lives here but auto-skips (exit 0 with a message) when the workspace is absent, so CI never needs models -- One small recorded fixture WAV (a few seconds, 16 kHz mono) committed under `scripts/tests/fixtures/`; nothing generates audio in-test -- The ASR replay harness (feed the fixture WAV through the same code path the WebSocket uses, measure capture-to-anchor wall time) is a documented manual check per platform, not a CI assertion; CI has no workspace and wall-clock assertions are flaky by construction -- Manual test matrix documented in the skill: macOS (reference), Windows, Linux; beam-splitter mirror check; phone remote; OBS overlay; device-picker preflight - -## Risks and mitigations - -- Nemotron streaming model quality/latency on low-end CPUs: zipformer-small fallback behind `asr-provider`, and tier 1 works with `none` -- ASR token timestamps are start-only in sherpa-onnx: alignment keys on token text order, not timestamps, so this costs nothing -- CPU contention between ASR, the LLM tick, and OBS on one machine: the tick budget rules above (adaptive cadence, pressure-skip, small-model fallback, timeouts) plus the bounded-queue drop policy keep voice-follow latency from compounding; the rail surfaces "ASR behind real-time" instead of silently lagging -- Ollama absent or model not pulled: producer mode degrades to the deterministic rail (timing and replan cues still work, coverage judgments off); the UI says exactly what is missing -- Repeated phrases in scripts confusing alignment: bounded window plus monotonic anchor limits damage; fixture-tested -- Browser mic pitfalls (wrong device, processing constraints ignored, wrong sample rate): the preflight screen with device picker, level meter, applied-settings readback, and the 10-second tracking test catches all of these before the first real take -- Scope creep: the kokoro spoken tier and the mc-cut take-log consumption are explicitly out of v1 - -## Explicitly out of scope (future hooks) - -- Spoken kokoro cue tier (designed, config key ships false; fast-follow) -- mc-cut consuming the take log (`out/take-log.json`, script positions and timestamps per take) to pre-anchor cut plans -- Chat/vision inputs to the producer (reading live chat is a natural producer input later) -- Cloud LLM rungs for `[llm]`, paid ASR rungs -- TLS for mic capture on remote devices (tablet-as-mic) -- Cross-platform batch transcription lane via sherpa-onnx parakeet-tdt offline export (separate TODO item this branch makes cheaper)