Parent: #421
Depends on: #427
Outcome
Add an offline-first speech-recognition provider that turns audio or an extracted audio stream into a versioned timed transcript, then allows ordinary transcript/subtitle derivatives.
Provider direction
Evaluate whisper.cpp or another maintained local-first open provider against:
- platform support;
- model licensing/distribution;
- deterministic/configuration-dependent behavior;
- timestamps;
- language handling;
- resource requirements;
- cancellation/progress;
- output formats;
- current maintenance.
Do not hardcode one provider into the domain model; the capability is provider-neutral.
Pipeline
audio/video
↓
audio stream selection/extraction
↓
speech.transcribe
↓
renderflow.timed-transcript/v1
├── SRT
├── WebVTT
├── JSON
└── text
Requirements
- Provider registry entry, version/model fingerprint, and explicit model artifact identity.
- Offline execution by default.
- Bounded CPU/memory/time/output behavior where practical.
- Progress/cancellation.
- Word/segment timestamp policy.
- Language selection/detection evidence.
- Transcript confidence/limitations.
- No claim that ASR transcript is human-approved copy.
- Preserve exact audio source digest and extraction lineage.
Acceptance criteria
Non-goals
Speaker diarization, translation, or advanced temporal editing unless separately scoped.
Candidate provider observation — 2026-09-25
Primary current candidate: https://github.com/ggml-org/whisper.cpp
The repository is public and not archived. Evaluate its current CLI/model/runtime/licensing surface during implementation; keep speech.transcribe provider-neutral so another local engine can satisfy the same contract later.
Parent: #421
Depends on: #427
Outcome
Add an offline-first speech-recognition provider that turns audio or an extracted audio stream into a versioned timed transcript, then allows ordinary transcript/subtitle derivatives.
Provider direction
Evaluate
whisper.cppor another maintained local-first open provider against:Do not hardcode one provider into the domain model; the capability is provider-neutral.
Pipeline
Requirements
Acceptance criteria
Non-goals
Speaker diarization, translation, or advanced temporal editing unless separately scoped.
Candidate provider observation — 2026-09-25
Primary current candidate: https://github.com/ggml-org/whisper.cpp
The repository is public and not archived. Evaluate its current CLI/model/runtime/licensing surface during implementation; keep
speech.transcribeprovider-neutral so another local engine can satisfy the same contract later.