Skip to content

CLI: transcript-only output, --process, and Hermes Agent docs - #65

Merged
dinooo13 merged 2 commits into
mainfrom
cli-stt-output
Sep 27, 2026
Merged

dinooo13 merged 2 commits into
mainfrom
cli-stt-output

Conversation

@dinooo13

Copy link
Copy Markdown
Owner

pladder-cli <audio file> now prints only the transcript, so another program can use it as its speech-to-text command. The first user is Hermes Agent, which runs any shell command declared under stt.providers and takes the transcript from stdout.

Changes

  • Default output is the text alone. Before, stdout also carried model ready in …, the timing line and a TEXT: prefix, and Hermes would have taken all three lines as the transcript. --verbose keeps the old output.
  • Download progress goes to stderr (in loadEngine, which bench also uses), so a first-run download never ends up in a transcript.
  • Clean failures. A file Core Audio cannot open, such as WebM, exits with status 1 and a message on stderr. Before, it trapped with "Error raised at top level".
  • --process runs the app's processors with the user's dictionary and processor toggles. With --verbose it also prints the RAW: text.
  • The processor list moves from AppModel to StandardProcessors in PladderSystem, so the app and the CLI build from the same list. TranscriptLanguage lives in PladderSystem, which is why the list goes there and not into Core. The list and its order are unchanged. CLAUDE.md now points processor additions at the new file.
  • The CLI never writes the settings file. It decodes it read-only instead of calling SettingsStore.load(), because load() moves a file it cannot decode aside. A CLI built from another branch must never move the live configuration. It honours PLADDER_SETTINGS_PATH.
  • Docs: new docs/HERMES.md, plus a note in INSTALL.md and a README FAQ entry.

Release-to-paste path

AppModel now reads StandardProcessors.factories in place of a local literal. The processors, their order and the pipeline code are the same, so I didn't run the benchmark. Say if you want the before-and-after numbers anyway.

Testing

  • swift build and swift test pass, 391 tests.
  • Tested by hand with the release build on an M1:
    • An Ogg Opus voice note (the Telegram format) decodes natively.
    • Default output is one line of text ending in a newline.
    • --verbose prints the old output.
    • --process removes fillers and turns spoken "new paragraph" into a paragraph break.
    • A WebM file and a missing file both exit 1 with a message on stderr.
    • A broken settings file prints a warning and falls back to the defaults. The live settings.json was byte-identical after the runs.
    • The ffmpeg command in docs/HERMES.md for WebM input works through sh.
  • Not tested: a real Hermes install. The config follows Hermes's tools/transcription_command.py: it runs the command with shell=True, reads {output_path} if it's non-empty and otherwise uses the stripped stdout, and passes command providers the audio file unconverted.

🤖 Generated with Claude Code

dinooo13 and others added 2 commits September 27, 2026 20:11
`pladder-cli <audio file>` now prints the transcript and nothing else, so
another program can read stdout as its speech-to-text. The load and
processing times move behind --verbose, and download progress goes to
stderr, bench included. A file Core Audio cannot open exits 1 with a
message instead of trapping.

--process runs the app's processors with the app's dictionary and toggles.
The processor list moves from AppModel to StandardProcessors in
PladderSystem so the app and the CLI build from the same one; the list and
its order are unchanged. The CLI decodes the settings file itself rather
than through SettingsStore.load(), which moves an undecodable file aside.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dinooo13
dinooo13 merged commit b400d20 into main Sep 27, 2026
1 check passed
@dinooo13
dinooo13 deleted the cli-stt-output branch September 27, 2026 18:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant