Skip to content

[Artifact Forest 7/9] Add local speech-to-timed-transcript provider #428

Description

@szmyty

Parent: #421
Depends on: #427

Outcome

Add an offline-first speech-recognition provider that turns audio or an extracted audio stream into a versioned timed transcript, then allows ordinary transcript/subtitle derivatives.

Provider direction

Evaluate whisper.cpp or another maintained local-first open provider against:

  • platform support;
  • model licensing/distribution;
  • deterministic/configuration-dependent behavior;
  • timestamps;
  • language handling;
  • resource requirements;
  • cancellation/progress;
  • output formats;
  • current maintenance.

Do not hardcode one provider into the domain model; the capability is provider-neutral.

Pipeline

audio/video
   ↓
audio stream selection/extraction
   ↓
speech.transcribe
   ↓
renderflow.timed-transcript/v1
   ├── SRT
   ├── WebVTT
   ├── JSON
   └── text

Requirements

  • Provider registry entry, version/model fingerprint, and explicit model artifact identity.
  • Offline execution by default.
  • Bounded CPU/memory/time/output behavior where practical.
  • Progress/cancellation.
  • Word/segment timestamp policy.
  • Language selection/detection evidence.
  • Transcript confidence/limitations.
  • No claim that ASR transcript is human-approved copy.
  • Preserve exact audio source digest and extraction lineage.

Acceptance criteria

  • One redistribution-safe spoken-audio fixture produces deterministic-enough evidenced timed transcript output.
  • SRT/WebVTT/JSON/text derivatives are produced from one normalized transcript rather than independent ASR runs.
  • Model/provider identity participates in cache/checkpoint keys.
  • Missing provider/model/resource requirements fail clearly.
  • No network is required by the default provider path.
  • Source media remains unchanged.

Non-goals

Speaker diarization, translation, or advanced temporal editing unless separately scoped.

Candidate provider observation — 2026-09-25

Primary current candidate: https://github.com/ggml-org/whisper.cpp

The repository is public and not archived. Evaluate its current CLI/model/runtime/licensing surface during implementation; keep speech.transcribe provider-neutral so another local engine can satisfy the same contract later.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions