Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Push-to-talk dictation for macOS 26 on Apple Silicon. Hold a key, speak, release

1. **Speed of the release-to-paste path.** This is the metric. The path runs from the hotkey release to the paste: capture stop, engine, processors, output. In code it is everything between the coordinator's `recordingStopped` and `inserted` events. Nothing goes on that path without a before-and-after benchmark.
2. **Minimal UI.** No new settings, windows or overlay elements unless a feature cannot work without them. Prefer removing to adding.
3. **On-device processing and privacy.** Audio and text never leave the Mac. No network calls, no telemetry, no accounts. The one-time model download from Hugging Face is the only exception and stays the only one.
3. **On-device processing and privacy.** Audio and text never leave the Mac. No network calls, no telemetry, no accounts. One-time model downloads from Hugging Face are the only exception and stay the only one: the speech model at first launch, and an S1-mini polish model only when the user picks it. Each is pinned to a commit and checked against its SHA-256.

## Critical-path rule

Expand All @@ -26,7 +26,8 @@ swift run -c release pladder-cli <audio file> # transcribe one file, print timin
./scripts/make-fixtures.sh # synthesise benchmark fixtures into bench/fixtures (gitignored)
swift run -c release pladder-cli bench bench/fixtures # whole-buffer benchmark
swift run -c release pladder-cli bench bench/fixtures --paced # feed at real time, time endUtterance, check identity
swift run -c release pladder-cli polish <text file> # run the polish prompt over a transcript, print both timings
swift run -c release pladder-cli polish <text file> [--model apple|s1-mini|s1-mini-8bit] # one transcript, cold and warm
swift run -c release pladder-cli polish-set docs/polish-set.json --model s1-mini-8bit # a polish model over the test set: error rate, exact matches, timings
./scripts/overlay-demo.sh [dir] [--style …] [--speed …] # play every overlay path with stand-ins, record the screen, cut contact sheets
```

Expand All @@ -50,7 +51,8 @@ swift run -c release pladder-cli polish <text file> # run the polish prompt ov
| Toggle key | Another recordable chord, off by default. A chord of its own latches at release however long the press; equal to the push-to-talk chord it makes that key hybrid: a tap under 400 ms latches, a longer hold stops at release. The next press of any chord, Escape or the 10 min cap ends a latched recording. The overlay is the same as for a held recording, in and out; only the menu's status line says which key stops it | Two-minute dictations should not need a key held for two minutes. Handy and VoiceInk default to hybrid on one key; here it is opt-in, because a stray tap would otherwise leave the microphone open until the cap pastes two minutes of room noise. Hold, toggle and hybrid are decided in `HotkeyGestureTracker`, a clockless value type timed by the instant each monitor stamps on its events, so both monitors behave alike and a press that waits for the microphone cannot make the next release look longer. A same-chord press within 50 ms of its release is a bounce (some Bluetooth keyboards do this mid-hold): it never acts, and the first one seen turns on a 50 ms settle before every stopping release for the rest of the run, so only a keyboard that needs it pays for it and the release path is otherwise untouched. No separate press debounce: both monitors already report alternating presses and releases. Without Accessibility the toggle chord registers with Carbon like the key; one Carbon cannot register has no stand-in, except that a hybrid chord follows the key's. With a lone modifier as a hybrid key, the Command of a later Cmd+C ends a latched recording |
| Escape | Discards a recording without transcribing and plays the stop sound; taken only while a recording is on | Never taken globally, so Escape keeps closing dialogs. On the tap `HotkeyChordSet` catches it before the chord trackers, so the interrupted-press rule never sees it, and Escape with the chord's own modifiers held still counts; Carbon registers the bare key around each recording, from the main queue so the release path never waits on it. Under Secure Event Input a modifier-only chord stays on the tap, where no key-down arrives, so Escape cannot cancel there |
| Output | Clipboard + simulated Cmd+V; the old clipboard is restored off the critical path | Universal, fast |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, in that order | No latency, no network. An earlier Apple Intelligence step was removed from this path unmeasured; the model is back behind the experimental polish toggle only |
| Polish model | A picker under the toggle: Apple Intelligence (default), S1-mini by Superwhisper at full precision (f16, 1.5 GB) or 8-bit (Q8_0, 805 MB), the S1-mini files run by llama.cpp on the GPU. `PolishRouter` is the coordinator's one refiner and forwards to the chosen model | On the polish set (docs/BENCHMARKS.md) S1-mini is more accurate than Apple's model and three to five times faster, and translated nothing where Apple's model translated two dictations. It is Qwen3-0.6B fine-tuned for dictation cleanup, trained on English, and needs its own fixed prompt. llama.cpp comes as its prebuilt XCFramework, pinned by release and checksum, because MLX Swift needs Xcode to build its Metal shaders; `bundle.sh` embeds and signs the framework. The GPU leaves the Neural Engine to Parakeet. A file downloads only when polish is on and the model picked, is loaded at the first key-down with the fixed prompt prefix decoded once, and is freed on a switch or with polish off. The licence asks for the name as "S1-mini" by "Superwhisper" |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, spoken punctuation, in that order | No latency, no network. Spoken punctuation takes only phrases that are never ordinary words ("period", "Punkt", "punto" stay words), runs last because the whitespace step folds line breaks, and needs Spanish evidence for "coma"; Parakeet already writes numbers as digits. An earlier Apple Intelligence step was removed from this path unmeasured; the model is back behind the experimental polish toggle only |
| Mute while dictating | Off by default; `kAudioDevicePropertyMute` on the default output device 200 ms into a recording, restored off the release path | Music or a call otherwise goes into the microphone. The delay means a tap-and-release never toggles anything; a device the user had already muted is left alone, and the device that was muted is the one unmuted even if the default changed meanwhile |
| UI language | Follows the macOS system language; no setting | String Catalogs (`Localizable.xcstrings` in the app, `KeyNames.xcstrings` in `PladderSystem`) are compiled by `swift build`; `bundle.sh` merges their `.lproj` folders into `Pladder.app/Contents/Resources`, so `Bundle.main` serves them and no code names a bundle. `swift run` shows English. Core and Engines emit enum cases; the app turns them into text. German first; more languages are catalog contributions |
| Recording cap | 10 min | Keeps the microphone from staying on when a key-up is lost. The cap ends a latched recording the same way |
Expand All @@ -61,6 +63,7 @@ swift run -c release pladder-cli polish <text file> # run the polish prompt ov
- `PladderCore` imports Foundation only. It never imports FluidAudio, AVFoundation or AppKit, so tests compile fast and engines are truly swappable.
- Adding an engine: implement `TranscriptionEngine` in its own file under `PladderEngines`, register it in the `EngineRegistry` built in `AppModel`. One file plus one registry line; the settings picker reads the registry. The engine lifecycle — building, loading, status polling and swapping — lives in `EngineLoader`.
- Adding a processor: implement `TextProcessor` in its own file, append a factory to `processorFactories` in `AppModel`. The pipeline is rebuilt when settings change, never per dictation. A processor sits on the critical path, so the benchmark rule applies.
- Adding a polish model: a `PolishModel` case in `Settings`, a `ModelFile` pinned to a commit with its SHA-256 if it downloads, a `TranscriptRefiner` in `PladderRefine`, and a branch in `AppModel.applyPolishModel`; `PolishRouter` hands it the calls, the picker and `polish-set` read the enum. Judge it with `pladder-cli polish-set` before it goes in.
- Adding a prompt: build an `OnDeviceLanguageModel(instructions:)` in `PladderRefine` and call `respond(to:)` or `respond(to:generating:)`; availability, prewarm, timeout and the drain of an abandoned call come with it. The coordinator only ever sees `TranscriptRefiner`.
- The correction learner's two seams are protocols in `PladderCore`, `PastedTextObserver` and `CorrectionReviewer`, with fakes in the tests; the Accessibility and Foundation Models implementations live in `PladderSystem` (`AXPasteObserver`) and `PladderRefine` (`FoundationModelsCorrectionReviewer`).
- Adding a hotkey role: a case in `HotkeyRole`, a `Hotkey` field in `Settings` with `[]` meaning off, an entry in the chords `startHotkey` hands the monitor and a mode in `startGesture`, a `HotkeyRecorderField` row. The monitors and trackers need nothing.
Expand All @@ -73,7 +76,8 @@ swift run -c release pladder-cli polish <text file> # run the polish prompt ov
- **First launch.** About 700 MB of CoreML models download from Hugging Face and compile on first load. The menu shows progress and the hotkey is disabled until the engine is ready.
- **Cold latency.** The engine loads at launch and stays resident. The audio engine is prepared at launch and runs only while the key is held, so the system microphone indicator is off when idle. A cold encoder pass costs about 110 ms more than a warm one, which is more than every other stage together, so the coordinator warms the Neural Engine every two seconds while the key is held. A release that lands inside a warm pass waits for it: the signature is `engine` well above `engine-time`.
- **Long recordings.** FluidAudio's encoder window is 15 s. Longer audio is split into windows and stitched, and seams can drop or duplicate words. Those windows now run while the key is held rather than at release, so the wait is flat with length, but the seam risk is unchanged: it is the same layout and the same merge. The paced benchmark guards it by requiring the text to be byte-identical to transcribing the whole recording at once, and the 30 s to 10 min fixtures watch the word error rate.
- **Polish latency and quality.** The model takes about 1.5 to 2 seconds warm on an M1 for a typical dictation and is capped at eight; past the cap the text is pasted as dictated. A small model can still rewrite rather than clean; the guided `cleanedText` field, greedy sampling, inline examples and naming the transcript's language keep that rare, and `pladder-cli polish` is where prompt changes are judged.
- **Polish latency and quality.** Apple's model takes about 1.5 to 2 seconds warm on an M1 for a typical dictation, S1-mini 0.3 to 0.5, and both are capped at eight; past the cap the text is pasted as dictated. A small model can still rewrite rather than clean; for Apple's the guided `cleanedText` field, greedy sampling, inline examples and naming the transcript's language keep that rare. S1-mini's known slips are in docs/BENCHMARKS.md, the worst a German number word read wrong, which real dictation rarely feeds it since Parakeet writes digits. `pladder-cli polish-set` is where model and prompt changes are judged.
- **Polish model memory.** S1-mini stays loaded while chosen and polish is on: about 0.8 or 1.5 GB of weights plus a 235 MB key-value cache for its 2,048-token context. Long dictations are polished in chunks of about 250 words to fit.
- **Clipboard clobbering.** Output saves the pasteboard, pastes, and restores it after a short delay.
- **Permissions.** Accessibility and Microphone grants are keyed to the code signature. `bundle.sh` signs with an Apple Development or Developer ID certificate when one is in the keychain; an ad-hoc signature changes on every build and resets both grants.
- **Lost key-up.** When the 10 min watchdog fires, the coordinator transcribes and pastes as if the user had released. Discarding is probably the right behaviour; tracked separately.
Expand Down
17 changes: 13 additions & 4 deletions Package.swift
Original file line number Diff line number Diff line change
Expand Up @@ -45,9 +45,18 @@ let package = Package(
]
),

// Apple's on-device model behind the polish hotkey. The only target
// that imports FoundationModels.
.target(name: "PladderRefine", dependencies: ["PladderCore"]),
// The polish models: Apple's on-device model, the only target that
// imports FoundationModels, and S1-mini through llama.cpp.
.target(name: "PladderRefine", dependencies: ["PladderCore", "llama"]),

// llama.cpp's own prebuilt release, Metal included: a dynamic
// framework that scripts/bundle.sh embeds in the app. Pinned by
// release and checksum; bump both together.
.binaryTarget(
name: "llama",
url: "https://github.com/ggml-org/llama.cpp/releases/download/b11191/llama-b11191-xcframework.zip",
checksum: "c8f9af07555a15b00a87334e13a21320596178c58bbafa2d7c4915c574e7086e"
),

// The menu bar app.
.executableTarget(
Expand All @@ -73,7 +82,7 @@ let package = Package(
// engines, or run the benchmark (see docs/BENCHMARKS.md).
.executableTarget(
name: "PladderCLI",
dependencies: ["PladderCore", "PladderEngines", "PladderAudio", "PladderBench", "PladderRefine"]
dependencies: ["PladderCore", "PladderEngines", "PladderAudio", "PladderBench", "PladderRefine", "PladderSystem"]
),

.testTarget(name: "PladderCoreTests", dependencies: ["PladderCore"]),
Expand Down
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,8 @@ You type long prompts all day. Speech is three to four times faster than typing,
- **Nothing leaves your Mac.** The speech model runs on the Neural Engine. There is no account, no server, no telemetry, and the app makes no network requests after the one-time model download.
- **It knows your words.** A dictionary turns what the model hears into what you meant. "clode code" becomes "Claude Code", "get hub" becomes "GitHub", every time, at zero cost in latency. List a word on its own and near misses of it are repaired too, so "Chat G P T" comes out as "ChatGPT" without you predicting every way the model might mangle it. Correct a word by hand after a dictation and Pladder offers to remember it — checked on device, added only when you say so.
- **It skips over the ums.** Hesitation sounds — "uh", "um", German "äh"/"ähm", Spanish "eh" — are dropped before the text is pasted. Still no cost in latency. Only English, German and Spanish fillers are covered for now; other languages pass through unchanged.
- **Polish it when you want to.** Hold a second key instead and Apple's on-device model cleans the transcript before it is pasted: "wait, no, Friday" becomes "Friday", spoken numbers become digits, "first… second…" becomes a list. A second or two, still on your Mac, and only when you ask.
- **Say "comma" and get one.** Spoken punctuation — "comma", "question mark", "new paragraph", German "Fragezeichen", Spanish "signo de interrogación" — becomes the mark itself, with no cost in latency.
- **Polish it if you like.** Turn on the experimental polish and a small model cleans the transcript on your Mac before it is pasted: "wait, no, Friday" becomes "Friday", "first… second…" becomes a list. Apple's on-device model takes a second or two; S1-mini by Superwhisper, downloaded once when you pick it, about half a second.
- **It behaves like part of macOS.** A menu bar app with a Liquid Glass status pill, a standard settings window, and nothing in the Dock.
- **Twenty-five languages, detected automatically.** English, German, Spanish, French and the rest of Europe in the same session, with no setting to flip.
- **Free and MIT.** The source is here. Read it, build it, change it.
Expand Down Expand Up @@ -94,7 +95,7 @@ Privacy here is not a policy, it is how the thing is built.

- **Audio never leaves the Mac.** The microphone is open only while the key is held, and macOS shows the orange indicator only then. Audio goes from the microphone to the Neural Engine and is discarded.
- **Text never leaves the Mac.** The transcript exists long enough to be pasted. Your previous clipboard is put back afterwards. After a paste, Pladder watches the field it pasted into for up to a minute through Accessibility, to notice when you fix a word; it reads only the pasted words and a little context, keeps nothing, and asks before adding anything to the dictionary.
- **No network.** The only request Pladder ever makes is the one-time download of the speech model from Hugging Face, about 700 MB, on first launch. After that it works with Wi-Fi off. There is no update check, no crash reporter, no analytics.
- **No network.** The only request Pladder ever makes is the one-time download of the speech model from Hugging Face, about 700 MB, on first launch, and, only if you pick S1-mini for the experimental polish, the one-time download of that model from Hugging Face too. After that it works with Wi-Fi off. There is no update check, no crash reporter, no analytics.
- **No account.** Nothing to sign up for, nothing to log in to, nothing to cancel.
- **Auditable.** The app is about five thousand lines of Swift under the MIT license, and none of them open a network connection. The model download is FluidAudio's, and it runs once.

Expand Down Expand Up @@ -151,7 +152,7 @@ Option+Space is Alfred's default hotkey and a common Raycast choice, and it type
Yes. Press the send key, Right Option by default, at any point while you hold the push-to-talk key, and Return is pressed after the text is pasted. That sends a chat message or runs a terminal command without touching the keyboard again. The send key can be changed in Settings, like the push-to-talk key.

**Can it clean up what I said?**
Record a key for Dictate and polish in Settings and hold that instead. The dictation goes through Apple Intelligence on your Mac before it is pasted, which takes a second or two. It needs Apple Intelligence turned on in System Settings; without it that key pastes the text as dictated.
Turn on Polish dictations in Settings, Processing, and pick a model. Apple Intelligence needs nothing downloaded but takes a second or two and needs Apple Intelligence turned on in System Settings. S1-mini by Superwhisper is downloaded once from Hugging Face (1.5 GB, or 805 MB at 8-bit) and takes about half a second; it is trained on English and also handles German and Spanish. Anything a model cannot fix is pasted as dictated.

**Can I toggle instead of holding?**
Yes. Record a toggle key in Settings: one tap starts a recording, the next tap inserts it. Give it the same combination as the push-to-talk key and that key does both: tap to start and tap again to insert, or hold and release as before. Escape discards a recording either way.
Expand Down
Loading
Loading