Skip to content

Latest commit

 

History

History
144 lines (115 loc) · 5.8 KB

File metadata and controls

144 lines (115 loc) · 5.8 KB
Dialogue Reader

Dialogue Reader

Reads on-screen game dialogue aloud, with a different voice per speaker.

Windows Python AutoHotkey Voices Last commit


Python, PySide6 (Qt), Windows

What it does

Capture OCR Speakers TTS
  • Pick rectangles on screen
  • Polled at 12 Hz for pixel changes
  • Window or screen capture, zoom-stable when possible
  • Auto-pauses while Magnifier is zoomed
  • WinOCR for clean UI text
  • EasyOCR for stylized game fonts
  • Per-region engine choice (dialogue vs speaker name)
  • Cosmetic-jitter dedup so the same line is not re-spoken
  • One region for the speaker name, one for dialogue
  • Each character gets their own voice from the pool
  • Mappings persist in speakers.json
  • Cycle voices live with a hotkey
  • Kokoro 28 English voices
  • Piper curated voice set
  • Sherpa VCTK (109), LibriTTS-R (904), MeloTTS
  • Mix engines in one pool, change speed live

Quick start

pip install -r requirements.txt

Install AutoHotkey v2, then double-click dialogue_reader.ahk. It launches main.py as a child process and binds your hotkeys. Closing the AHK script terminates the Python process.

In game, use:

Hotkey What it does
F1 Open/close the region manager: drag to add dialogue regions, drag an outline's edge to move it, its dots to resize, right-click to delete
Shift+F1 Same manager, adding speaker-name regions
Ctrl+F1 Clear all regions
End Pause or unpause
PgUp / PgDn TTS speed up / down
F2 / Ctrl+F2 Cycle the current speaker's voice forward / back

Bindings live in dialogue_reader.ini. Right-click the tray icon and pick "Reload Script" to apply changes.

Per-process regions and game profiles

Regions belong to the app they were drawn over: F1 manages only the focused app's boxes, tabbed-out apps go quiet, and closing a game removes its boxes. In the settings app you can save a game's layout as a profile (with a mini preview of its boxes), re-apply it any time, or flip on Auto to apply it whenever that game launches. Layouts are stored window-relative and scale with the window size.

Settings app

Double-click dialogue_reader_ui.pyw for a settings window: live controls (pause, speed, voice cycling), media pause, capture/OCR modes, the voice pool with per-voice previews, and hotkey editing. Changes hot-apply to the running reader (RELOAD_CONFIG over UDP); only hotkey changes need the restart button. Closing the window keeps it in the tray.


Configuration

dialogue_reader.ini is the single source of truth:

Section What you set
[Hotkeys] Which key triggers each command
[OCR] Engine for dialogue and speaker regions (winocr / easyocr)
[Capture] Capture mode (auto, screen, window)
[Speakers] Voice-assignment strategy (random, round_robin, inverse_round_robin)
[Magnifier] SkipWhenZoomed: pause polling while zoomed
[Media] PauseDuringSpeech: pause YouTube/Spotify etc. while the reader speaks; ResumeDelayMs: quiet period before resuming (default 1000)
[Voices] Default voice and the pool. Supports <engine>:all and sherpa:<model>:<a>-<b> ranges
[Polling] TextConfirmPolls: how many identical OCR polls before speaking

Layout

main.py             Main loop: poll regions, OCR, dedup, speak
capture.py          Region capture (mss / PrintWindow window-mode)
region_picker.py    Click-and-drag region selector (PySide6 overlay)
ocr.py              WinOCR and EasyOCR wrappers + worker thread
tts.py              TTS dispatcher (piper / kokoro / sherpa)
media_gate.py       Pauses other media (GSMTC) while speaking, resumes after
profiles.py         Game profiles: window-relative region snapshots (profiles.json)
ui/                 Settings app: pywebview window (api.py backend, index.html)
dialogue_reader_ui.pyw  Settings app launcher (close-to-tray, single instance)
kokoro_tts.py       Kokoro-ONNX backend
sherpa_tts.py       Sherpa-ONNX backend (VCTK, LibriTTS-R, MeloTTS)
speakers.py         Speaker to voice mapping with persistence
magnifier.py        Detects when Windows Magnifier is zoomed in
command_server.py   UDP server on port 7849 listening for AHK commands
dialogue_reader.ahk Hotkey script and Python child process supervisor
dialogue_reader.ini All user settings
speakers.json       Persistent speaker to voice assignments
docs/voices/        Voice catalogs (CSV) for each engine