Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Before and after any change that touches the code between `recordingStopped` and
swift build # debug build of everything
swift test # unit tests, well under a second
./scripts/bundle.sh [--run] [--install] # release build → dist/Pladder.app, signed
swift run -c release pladder-cli <audio file> # transcribe one file, print timing
swift run -c release pladder-cli <audio file> [--process] [--verbose] # transcribe one file to stdout; --process runs the app's processors, --verbose adds timing
./scripts/make-fixtures.sh # synthesise benchmark fixtures into bench/fixtures (gitignored)
swift run -c release pladder-cli bench bench/fixtures # whole-buffer benchmark
swift run -c release pladder-cli bench bench/fixtures --paced # feed at real time, time endUtterance, check identity
Expand Down Expand Up @@ -62,7 +62,7 @@ swift run -c release pladder-cli polish-set docs/polish-set.json --model s1-mini

- `PladderCore` imports Foundation only. It never imports FluidAudio, AVFoundation or AppKit, so tests compile fast and engines are truly swappable.
- Adding an engine: implement `TranscriptionEngine` in its own file under `PladderEngines`, register it in the `EngineRegistry` built in `AppModel`. One file plus one registry line; the settings picker reads the registry. The engine lifecycle — building, loading, status polling and swapping — lives in `EngineLoader`.
- Adding a processor: implement `TextProcessor` in its own file, append a factory to `processorFactories` in `AppModel`. The pipeline is rebuilt when settings change, never per dictation. A processor sits on the critical path, so the benchmark rule applies.
- Adding a processor: implement `TextProcessor` in its own file, append a factory to `StandardProcessors.factories` in `PladderSystem`, which the app and `pladder-cli --process` both build from. The pipeline is rebuilt when settings change, never per dictation. A processor sits on the critical path, so the benchmark rule applies.
- Adding a polish model: a `PolishModel` case in `Settings`, a `ModelFile` pinned to a commit with its SHA-256 if it downloads, a `TranscriptRefiner` in `PladderRefine`, and a branch in `AppModel.applyPolishModel`; `PolishRouter` hands it the calls, the picker and `polish-set` read the enum. Judge it with `pladder-cli polish-set` before it goes in.
- Adding a prompt: build an `OnDeviceLanguageModel(instructions:)` in `PladderRefine` and call `respond(to:)` or `respond(to:generating:)`; availability, prewarm, timeout and the drain of an abandoned call come with it. The coordinator only ever sees `TranscriptRefiner`.
- The correction learner's two seams are protocols in `PladderCore`, `PastedTextObserver` and `CorrectionReviewer`, with fakes in the tests; the Accessibility and Foundation Models implementations live in `PladderSystem` (`AXPasteObserver`) and `PladderRefine` (`FoundationModelsCorrectionReviewer`).
Expand Down
5 changes: 4 additions & 1 deletion INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,12 +75,15 @@ A free Apple ID is enough for an Apple Development certificate: sign in to Xcode

### Command-line transcriber

`pladder-cli` loads the same engine and prints the transcript for any audio file. It is the quickest way to check the model without the GUI, and prints timing so you can see the real-time factor on your machine.
`pladder-cli` loads the same engine and prints the transcript for any audio file, and nothing else, so a script can read it. It is the quickest way to check the model without the GUI. `--verbose` adds timing, so you can see the real-time factor on your machine, and `--process` runs the app's processors over the text with your dictionary.

```sh
swift run -c release pladder-cli recording.wav
swift run -c release pladder-cli recording.wav --process --verbose
```

It reads WAV, M4A, MP3, FLAC, CAF and Ogg Opus. To use it as the speech-to-text of another program, see [docs/HERMES.md](docs/HERMES.md).

`pladder-cli bench <dir>` runs the benchmark over synthetic fixtures. See [docs/BENCHMARKS.md](docs/BENCHMARKS.md) for the procedure and the M1 baseline.

### Icon and screenshots
Expand Down
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,6 +157,9 @@ Turn on Polish dictations in Settings, Processing, and pick a model. Apple Intel
**Can I toggle instead of holding?**
Yes. Record a toggle key in Settings: one tap starts a recording, the next tap inserts it. Give it the same combination as the push-to-talk key and that key does both: tap to start and tap again to insert, or hold and release as before. Escape discards a recording either way.

**Can another program use it for speech-to-text?**
Yes. `pladder-cli` transcribes an audio file and prints only the text, so anything that runs a command for speech-to-text can use it. [docs/HERMES.md](docs/HERMES.md) sets it up for Hermes Agent's voice messages.

**What about long dictations?**
Recordings stop at 10 minutes, so a lost key-up never leaves the microphone on. Audio longer than 15 seconds is transcribed in overlapping windows.

Expand Down
17 changes: 3 additions & 14 deletions Sources/Pladder/AppModel.swift
Original file line number Diff line number Diff line change
Expand Up @@ -212,20 +212,9 @@ final class AppModel {
initial.engineID = fallback.id
}

// Processors, in pipeline order: fillers go first so the dictionary sees
// cleaned text, the fuzzy custom-word corrector runs after the exact
// replacer so it only sees what the replacer could not fix, whitespace
// is tidied next, and spoken punctuation comes last, because the
// whitespace step would fold its paragraph breaks back into spaces.
// Each entry is a factory so a processor that needs settings builds
// itself from them; nothing here knows which processor that is.
let processorFactories: [@Sendable (Settings) -> any TextProcessor] = [
{ _ in FillerRemover(languageHint: { TranscriptLanguage.hint(for: $0) }) },
{ DictionaryReplacer(entries: $0.dictionary) },
{ CustomWordCorrector(entries: $0.dictionary) },
{ _ in WhitespaceNormalizer() },
{ _ in SpokenPunctuation(languageHint: { TranscriptLanguage.hint(for: $0) }) },
]
// Processors, in pipeline order; the list and its reasons are in
// `StandardProcessors`, which the CLI shares.
let processorFactories = StandardProcessors.factories
self.processors = processorFactories.map { $0(initial) }

let router = PolishRouter(applePolisher)
Expand Down
88 changes: 75 additions & 13 deletions Sources/PladderCLI/main.swift
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,13 @@ import PladderSystem

// Developer tool.
//
// pladder-cli <audio file> load Parakeet, print the transcript and timing
// pladder-cli <audio file> load Parakeet, print the transcript and nothing else,
// so a script or another program's STT hook can read
// stdout. Errors go to stderr with exit status 1.
// [--process] run the app's processors over it with the app's
// dictionary and toggles, read from its settings file
// (or PLADDER_SETTINGS_PATH), which is never written.
// [--verbose] also print the load and processing times.
// pladder-cli bench <fixtures dir> run the benchmark (see docs/BENCHMARKS.md)
// [--runs N] runs per fixture, default 6; the first is discarded.
// Use 11 to settle a result near the noise line.
Expand Down Expand Up @@ -50,7 +56,7 @@ import PladderSystem

func usage() -> Never {
FileHandle.standardError.write(Data("""
usage: pladder-cli <audio file>
usage: pladder-cli <audio file> [--process] [--verbose]
pladder-cli bench <fixtures dir> [--runs N] [--pause S]
pladder-cli bench <fixtures dir> --paced [--runs N] [--pause S] [--all] [--live]
pladder-cli polish <text file | -> [--model apple|s1-mini|s1-mini-8bit | --gguf <file>] [--control <line>] [--instructions <file>]
Expand Down Expand Up @@ -95,7 +101,11 @@ func loadEngine(_ engine: any TranscriptionEngine) async throws -> Duration {
while !Task.isCancelled {
if case .downloading(let p) = await engine.status, let p {
let pct = Int(p * 100)
if pct != lastPrinted { print("downloading \(pct)%"); lastPrinted = pct }
// stderr, so a first run's download never lands in a transcript.
if pct != lastPrinted {
FileHandle.standardError.write(Data("downloading \(pct)%\n".utf8))
lastPrinted = pct
}
}
try? await Task.sleep(for: .milliseconds(500))
}
Expand Down Expand Up @@ -171,15 +181,52 @@ func thermalTag() -> String {

// MARK: - Transcribe one file

func transcribeFile(_ path: String) async throws {
let engine = FluidAudioIncrementalEngine()
let loadTime = try await loadEngine(engine)
print(String(format: "model ready in %.1fs", seconds(loadTime)))
/// The app's settings, for `--process`. Decoded here rather than through
/// `SettingsStore.load()`, which moves a file it cannot decode aside: a CLI
/// built from another branch must never touch the live configuration.
func appSettings() -> Settings {
let url = ProcessInfo.processInfo.environment["PLADDER_SETTINGS_PATH"].flatMap { $0.isEmpty ? nil : URL(filePath: $0) }
?? FileManager.default.homeDirectoryForCurrentUser
.appending(path: "Library/Application Support/Pladder/settings.json")
let fallback = Settings(engineID: FluidAudioIncrementalEngine.engineID)
guard let data = try? Data(contentsOf: url) else { return fallback }
do {
return try JSONDecoder().decode(Settings.self, from: data)
} catch {
FileHandle.standardError.write(Data("pladder-cli: \(url.path): \(error); processing with the defaults\n".utf8))
return fallback
}
}

let samples = try loadSamples(URL(fileURLWithPath: path))
let transcript = try await engine.transcribe(samples)
print(String(format: "audio %.2fs, processed in %.3fs (%.0fx realtime)", transcript.audioDuration, transcript.processingTime, transcript.realtimeFactor))
print("TEXT: \(transcript.text)")
func transcribeFile(_ path: String, process: Bool, verbose: Bool) async {
do {
let engine = FluidAudioIncrementalEngine()
let loadTime = try await loadEngine(engine)
if verbose { print(String(format: "model ready in %.1fs", seconds(loadTime))) }

let samples = try loadSamples(URL(fileURLWithPath: path))
let transcript = try await engine.transcribe(samples)
var text = transcript.text
if process {
let settings = appSettings()
let pipeline = ProcessorPipeline(StandardProcessors.factories.map { $0(settings) }, onFailure: { id, error in
FileHandle.standardError.write(Data("pladder-cli: processor \(id) failed: \(error)\n".utf8))
})
text = await pipeline.run(text, disabled: settings.disabledProcessors)
}
if verbose {
print(String(format: "audio %.2fs, processed in %.3fs (%.0fx realtime)", transcript.audioDuration, transcript.processingTime, transcript.realtimeFactor))
if process { print("RAW: \(transcript.text)") }
print("TEXT: \(text)")
} else {
print(text)
}
} catch {
// A file Core Audio cannot open (WebM, say) is the likely failure; a
// message and a status rather than a trap, for whatever called us.
FileHandle.standardError.write(Data("pladder-cli: \(path): \(error)\n".utf8))
exit(1)
}
}

// MARK: - Benchmark
Expand Down Expand Up @@ -701,6 +748,21 @@ case "polish", "polish-set":
} else {
try await runPolishSet(textPath, model: model, options: options)
}
case let path?:
try await transcribeFile(path)
default:
var process = false
var verbose = false
var path: String?
for arg in arguments {
if arg == "--process" {
process = true
} else if arg == "--verbose" {
verbose = true
} else if path == nil {
path = arg
} else {
usage()
}
}
guard let path else { usage() }
await transcribeFile(path, process: process, verbose: verbose)
}
22 changes: 22 additions & 0 deletions Sources/PladderSystem/StandardProcessors.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
import Foundation
import PladderCore

/// The app's text processors, shared with `pladder-cli --process` so the two
/// cannot drift apart.
///
/// Pipeline order: fillers go first so the dictionary sees cleaned text, the
/// fuzzy custom-word corrector runs after the exact replacer so it only sees
/// what the replacer could not fix, whitespace is tidied next, and spoken
/// punctuation comes last, because the whitespace step would fold its
/// paragraph breaks back into spaces. Each entry is a factory so a processor
/// that needs settings builds itself from them; nothing here knows which
/// processor that is.
public enum StandardProcessors {
public static let factories: [@Sendable (Settings) -> any TextProcessor] = [
{ _ in FillerRemover(languageHint: { TranscriptLanguage.hint(for: $0) }) },
{ DictionaryReplacer(entries: $0.dictionary) },
{ CustomWordCorrector(entries: $0.dictionary) },
{ _ in WhitespaceNormalizer() },
{ _ in SpokenPunctuation(languageHint: { TranscriptLanguage.hint(for: $0) }) },
]
}
64 changes: 64 additions & 0 deletions docs/HERMES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Pladder as Hermes Agent's speech-to-text

[Hermes Agent](https://github.com/NousResearch/hermes-agent) transcribes voice
messages, from Telegram or its own voice mode, with a speech-to-text provider.
Besides its built-in ones it runs any shell command declared under
`stt.providers`, and `pladder-cli` is such a command: it takes an audio file and
prints the transcript to stdout, nothing else. Voice messages are then
transcribed by Parakeet on the Neural Engine, on the Mac Hermes runs on, with no
API key and no network.

Hermes has to run on the Mac: the CLI needs Apple Silicon and macOS 26.

## Install the CLI

Build it once and copy it somewhere stable; the binary under `.build` is
replaced by every build.

```sh
swift build -c release --product pladder-cli
sudo install .build/release/pladder-cli /usr/local/bin/
```

The first run downloads Parakeet (about 460 MB) into the same cache the app
uses, so with the app installed there is nothing to download.

```sh
pladder-cli some-voice-note.ogg
```

## Configure Hermes

In `~/.hermes/config.yaml`:

```yaml
stt:
provider: pladder
providers:
pladder:
type: command
command: "/usr/local/bin/pladder-cli {input_path} --process"
timeout: 120
```

`--process` runs the app's processors over the text: fillers removed, your
dictionary applied, spoken punctuation. It reads the app's settings file for the
dictionary and the processor toggles and never writes it. Leave the flag out
for Parakeet's text as it came.

## What to expect

- **Speed.** Hermes starts a process per message. With the model already
compiled, loading takes 0.1 to 0.3 s and a short voice note is transcribed in
about 0.2 s on an M1.
- **Formats.** Hermes hands command providers the file as it arrived. The CLI
reads WAV, M4A, MP3, FLAC, CAF and Ogg Opus, which covers Telegram voice
notes. It cannot open WebM; for a source that sends it, convert first:
`command: "ffmpeg -loglevel error -i {input_path} -f wav {output_dir}/in.wav && /usr/local/bin/pladder-cli {output_dir}/in.wav --process"`.
- **Language.** Parakeet detects the language itself, one of the 25 European
languages it knows, so Hermes's `stt.language` and `{language}` are ignored.
- **Errors.** A file the CLI cannot read, or a model that fails to load, exits
with status 1 and a message on stderr, which Hermes reports as the
provider's error.
- **No polish.** The polish models, Apple Intelligence and S1-mini, are not
applied; the CLI stops after the processors.
Loading