| Capture | OCR | Speakers | TTS |
|---|---|---|---|
|
|
|
|
pip install -r requirements.txtInstall AutoHotkey v2, then double-click dialogue_reader.ahk. It launches main.py as a child process and binds your hotkeys. Closing the AHK script terminates the Python process.
In game, use:
| Hotkey | What it does |
|---|---|
F1 |
Open/close the region manager: drag to add dialogue regions, drag an outline's edge to move it, its dots to resize, right-click to delete |
Shift+F1 |
Same manager, adding speaker-name regions |
Ctrl+F1 |
Clear all regions |
End |
Pause or unpause |
PgUp / PgDn |
TTS speed up / down |
F2 / Ctrl+F2 |
Cycle the current speaker's voice forward / back |
Bindings live in dialogue_reader.ini. Right-click the tray icon and pick "Reload Script" to apply changes.
Regions belong to the app they were drawn over: F1 manages only the focused app's boxes, tabbed-out apps go quiet, and closing a game removes its boxes. In the settings app you can save a game's layout as a profile (with a mini preview of its boxes), re-apply it any time, or flip on Auto to apply it whenever that game launches. Layouts are stored window-relative and scale with the window size.
Double-click dialogue_reader_ui.pyw for a settings window: live controls (pause, speed, voice cycling), media pause, capture/OCR modes, the voice pool with per-voice previews, and hotkey editing. Changes hot-apply to the running reader (RELOAD_CONFIG over UDP); only hotkey changes need the restart button. Closing the window keeps it in the tray.
dialogue_reader.ini is the single source of truth:
| Section | What you set |
|---|---|
[Hotkeys] |
Which key triggers each command |
[OCR] |
Engine for dialogue and speaker regions (winocr / easyocr) |
[Capture] |
Capture mode (auto, screen, window) |
[Speakers] |
Voice-assignment strategy (random, round_robin, inverse_round_robin) |
[Magnifier] |
SkipWhenZoomed: pause polling while zoomed |
[Media] |
PauseDuringSpeech: pause YouTube/Spotify etc. while the reader speaks; ResumeDelayMs: quiet period before resuming (default 1000) |
[Voices] |
Default voice and the pool. Supports <engine>:all and sherpa:<model>:<a>-<b> ranges |
[Polling] |
TextConfirmPolls: how many identical OCR polls before speaking |
main.py Main loop: poll regions, OCR, dedup, speak
capture.py Region capture (mss / PrintWindow window-mode)
region_picker.py Click-and-drag region selector (PySide6 overlay)
ocr.py WinOCR and EasyOCR wrappers + worker thread
tts.py TTS dispatcher (piper / kokoro / sherpa)
media_gate.py Pauses other media (GSMTC) while speaking, resumes after
profiles.py Game profiles: window-relative region snapshots (profiles.json)
ui/ Settings app: pywebview window (api.py backend, index.html)
dialogue_reader_ui.pyw Settings app launcher (close-to-tray, single instance)
kokoro_tts.py Kokoro-ONNX backend
sherpa_tts.py Sherpa-ONNX backend (VCTK, LibriTTS-R, MeloTTS)
speakers.py Speaker to voice mapping with persistence
magnifier.py Detects when Windows Magnifier is zoomed in
command_server.py UDP server on port 7849 listening for AHK commands
dialogue_reader.ahk Hotkey script and Python child process supervisor
dialogue_reader.ini All user settings
speakers.json Persistent speaker to voice assignments
docs/voices/ Voice catalogs (CSV) for each engine
