Skip to content

Repository files navigation

Dialogue Reader

Dialogue Reader

Reads on-screen game dialogue aloud, with a different voice per speaker.

Windows Python AutoHotkey Voices Last commit


Python, PySide6 (Qt), Windows

What it does

Capture OCR Speakers TTS
  • Pick rectangles on screen
  • Polled at 12 Hz for pixel changes
  • Window or screen capture, zoom-stable when possible
  • Auto-pauses while Magnifier is zoomed
  • WinOCR for clean UI text
  • EasyOCR for stylized game fonts
  • Per-region engine choice (dialogue vs speaker name)
  • Cosmetic-jitter dedup so the same line is not re-spoken
  • One region for the speaker name, one for dialogue
  • Each character gets their own voice from the pool
  • Mappings persist in speakers.json
  • Cycle voices live with a hotkey
  • Kokoro 28 English voices
  • Piper curated voice set
  • Sherpa VCTK (109), LibriTTS-R (904), MeloTTS
  • Mix engines in one pool, change speed live

Quick start

pip install -r requirements.txt

Install AutoHotkey v2, then double-click dialogue_reader.ahk. It launches main.py as a child process and binds your hotkeys. Closing the AHK script terminates the Python process.

In game, use:

Hotkey What it does
F1 Open/close the region manager: drag to add dialogue regions, drag an outline's edge to move it, its dots to resize, right-click to delete
Shift+F1 Same manager, adding speaker-name regions
Ctrl+F1 Clear all regions
End Pause or unpause
PgUp / PgDn TTS speed up / down
F2 / Ctrl+F2 Cycle the current speaker's voice forward / back

Bindings live in dialogue_reader.ini. Right-click the tray icon and pick "Reload Script" to apply changes.

Per-process regions and game profiles

Regions belong to the app they were drawn over: F1 manages only the focused app's boxes, tabbed-out apps go quiet, and closing a game removes its boxes. In the settings app you can save a game's layout as a profile (with a mini preview of its boxes), re-apply it any time, or flip on Auto to apply it whenever that game launches. Layouts are stored window-relative and scale with the window size.

Settings app

Double-click dialogue_reader_ui.pyw for a settings window: live controls (pause, speed, voice cycling), media pause, capture/OCR modes, the voice pool with per-voice previews, and hotkey editing. Changes hot-apply to the running reader (RELOAD_CONFIG over UDP); only hotkey changes need the restart button. Closing the window keeps it in the tray.


Configuration

dialogue_reader.ini is the single source of truth:

Section What you set
[Hotkeys] Which key triggers each command
[OCR] Engine for dialogue and speaker regions (winocr / easyocr)
[Capture] Capture mode (auto, screen, window)
[Speakers] Voice-assignment strategy (random, round_robin, inverse_round_robin)
[Magnifier] SkipWhenZoomed: pause polling while zoomed
[Media] PauseDuringSpeech: pause YouTube/Spotify etc. while the reader speaks; ResumeDelayMs: quiet period before resuming (default 1000)
[Voices] Default voice and the pool. Supports <engine>:all and sherpa:<model>:<a>-<b> ranges
[Polling] TextConfirmPolls: how many identical OCR polls before speaking

Layout

main.py             Main loop: poll regions, OCR, dedup, speak
capture.py          Region capture (mss / PrintWindow window-mode)
region_picker.py    Click-and-drag region selector (PySide6 overlay)
ocr.py              WinOCR and EasyOCR wrappers + worker thread
tts.py              TTS dispatcher (piper / kokoro / sherpa)
media_gate.py       Pauses other media (GSMTC) while speaking, resumes after
profiles.py         Game profiles: window-relative region snapshots (profiles.json)
ui/                 Settings app: pywebview window (api.py backend, index.html)
dialogue_reader_ui.pyw  Settings app launcher (close-to-tray, single instance)
kokoro_tts.py       Kokoro-ONNX backend
sherpa_tts.py       Sherpa-ONNX backend (VCTK, LibriTTS-R, MeloTTS)
speakers.py         Speaker to voice mapping with persistence
magnifier.py        Detects when Windows Magnifier is zoomed in
command_server.py   UDP server on port 7849 listening for AHK commands
dialogue_reader.ahk Hotkey script and Python child process supervisor
dialogue_reader.ini All user settings
speakers.json       Persistent speaker to voice assignments
docs/voices/        Voice catalogs (CSV) for each engine

About

OCR tool to ahndle dialogue and text in games. Automatically updates as new dialogue is written and assigns individual voices per character.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages