Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ swift run -c release pladder-cli <audio file> # transcribe one file, print timin
./scripts/make-fixtures.sh # synthesise benchmark fixtures into bench/fixtures (gitignored)
swift run -c release pladder-cli bench bench/fixtures # whole-buffer benchmark
swift run -c release pladder-cli bench bench/fixtures --paced # feed at real time, time endUtterance, check identity
swift run -c release pladder-cli polish <text file> # run the polish prompt over a transcript, print both timings
```

## Decisions
Expand All @@ -43,8 +44,9 @@ swift run -c release pladder-cli bench bench/fixtures --paced # feed at real
| Secure Event Input | `IsSecureEventInputEnabled()` polled with the grant; sustained 3 s and a chord Carbon can register → Carbon monitor until it clears | A password field or Terminal's Secure Keyboard Entry stops taps receiving key events; modifier-only chords are unaffected and stay on the tap |
| Without Accessibility | Carbon `RegisterEventHotKey` plus clipboard-only output | A standard account cannot grant Accessibility without an admin. Carbon needs no permission but wants exactly one regular key and collapses left and right, so modifier-only chords are refused in the recorder; the transcript is left on the clipboard and the overlay says "press ⌘V". A stored chord Carbon cannot register, a lone Right Command say, is stood in for by the default Option+Space and the menu names it; the stored chord returns with the grant. `CopySymbolicHotKeys` only feeds the warning that an enabled macOS shortcut owns the recorded chord. `AppModel` polls the grant every two seconds and swaps the monitor in both directions |
| Send key | Press Right Option (configurable) while the hotkey is held and Return is posted 50 ms after Cmd+V | Sends a chat message or runs a command without a second trip to the keyboard; the Return is posted from a detached task so it stays off the release-to-paste path |
| Polish hotkey | "Dictate and polish": a second recordable chord, off by default. A dictation started with it runs the usual pipeline, then Apple's on-device model (FoundationModels, `PladderRefine`) with a fixed cleanup prompt, then pastes | Self-corrections, spoken punctuation, number words and lists are beyond the deterministic processors, and the model runs on device with nothing to download. It costs one to three seconds, so it never touches the normal hotkey's path: the branch is one Bool read; the session is created and prewarmed at key-down; transcripts under four words skip it; anything the model cannot do (Apple Intelligence off, refusal, the 8 s timeout) pastes the text as dictated. Logged as its own `polished release-to-paste` line |
| Output | Clipboard + simulated Cmd+V; the old clipboard is restored off the critical path | Universal, fast |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, in that order | No latency, no network. The Apple Intelligence cleanup step was removed because it sat on the critical path without anyone measuring what it cost |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, in that order | No latency, no network. An earlier Apple Intelligence step was removed from this path unmeasured; the model is back behind the polish hotkey only |
| Mute while dictating | Off by default; `kAudioDevicePropertyMute` on the default output device 200 ms into a recording, restored off the release path | Music or a call otherwise goes into the microphone. The delay means a tap-and-release never toggles anything; a device the user had already muted is left alone, and the device that was muted is the one unmuted even if the default changed meanwhile |
| UI language | Follows the macOS system language; no setting | String Catalogs (`Localizable.xcstrings` in the app, `KeyNames.xcstrings` in `PladderSystem`) are compiled by `swift build`; `bundle.sh` merges their `.lproj` folders into `Pladder.app/Contents/Resources`, so `Bundle.main` serves them and no code names a bundle. `swift run` shows English. Core and Engines emit enum cases; the app turns them into text. German first; more languages are catalog contributions |
| Recording cap | 120 s | Keeps the microphone from staying on when a key-up is lost |
Expand All @@ -55,6 +57,7 @@ swift run -c release pladder-cli bench bench/fixtures --paced # feed at real
- `PladderCore` imports Foundation only. It never imports FluidAudio, AVFoundation or AppKit, so tests compile fast and engines are truly swappable.
- Adding an engine: implement `TranscriptionEngine` in its own file under `PladderEngines`, register it in the `EngineRegistry` built in `AppModel`. One file plus one registry line; the settings picker reads the registry. The engine lifecycle — building, loading, status polling and swapping — lives in `EngineLoader`.
- Adding a processor: implement `TextProcessor` in its own file, append a factory to `processorFactories` in `AppModel`. The pipeline is rebuilt when settings change, never per dictation. A processor sits on the critical path, so the benchmark rule applies.
- Adding a prompt: build an `OnDeviceLanguageModel(instructions:)` in `PladderRefine` and call `respond(to:)` or `respond(to:generating:)`; availability, prewarm, timeout and the drain of an abandoned call come with it. The coordinator only ever sees `TranscriptRefiner`.
- Adding a language: add a `<code>` localization to both catalogs; nothing else. Adding a *string*: the key is the exact English text, and `PladderCore` never holds one — it emits an enum case and `Sources/Pladder/StatusText.swift` words it.
- Engine and capture are actors. The coordinator is `@MainActor` because it drives UI. It owns the state machine and nothing else; every dependency is injected, so tests run it with in-memory fakes.

Expand All @@ -64,6 +67,7 @@ swift run -c release pladder-cli bench bench/fixtures --paced # feed at real
- **First launch.** About 700 MB of CoreML models download from Hugging Face and compile on first load. The menu shows progress and the hotkey is disabled until the engine is ready.
- **Cold latency.** The engine loads at launch and stays resident. The audio engine is prepared at launch and runs only while the key is held, so the system microphone indicator is off when idle. A cold encoder pass costs about 110 ms more than a warm one, which is more than every other stage together, so the coordinator warms the Neural Engine every two seconds while the key is held. A release that lands inside a warm pass waits for it: the signature is `engine` well above `engine-time`.
- **Long recordings.** FluidAudio's encoder window is 15 s. Longer audio is split into windows and stitched, and seams can drop or duplicate words. Those windows now run while the key is held rather than at release, so the wait is flat with length, but the seam risk is unchanged: it is the same layout and the same merge. The paced benchmark guards it by requiring the text to be byte-identical to transcribing the whole recording at once, and the 30 s to 10 min fixtures watch the word error rate.
- **Polish latency and quality.** The model takes about 1.5 to 2 seconds warm on an M1 for a typical dictation and is capped at eight; past the cap the text is pasted as dictated. A small model can still rewrite rather than clean; the guided `cleanedText` field, greedy sampling, inline examples and naming the transcript's language keep that rare, and `pladder-cli polish` is where prompt changes are judged.
- **Clipboard clobbering.** Output saves the pasteboard, pastes, and restores it after a short delay.
- **Permissions.** Accessibility and Microphone grants are keyed to the code signature. `bundle.sh` signs with an Apple Development or Developer ID certificate when one is in the keychain; an ad-hoc signature changes on every build and resets both grants.
- **Lost key-up.** When the 120 s watchdog fires, the coordinator transcribes and pastes as if the user had released. Discarding is probably the right behaviour; tracked separately.
Expand Down
2 changes: 2 additions & 0 deletions INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,8 @@ This compiles a release build, wraps it into `Pladder.app`, signs it, copies it
2. **Grant Accessibility** when prompted. System Settings opens on the Accessibility list; switch Pladder on. That is what lets Pladder see the push-to-talk key in other apps and paste the result. It does not need Input Monitoring.

On a standard (non-administrator) account, ticking that box asks for an administrator password, so ask an admin to do it once — the grant is keyed to the app's code signature and survives updates. Without it Pladder still works in a reduced form: Option+Space works the same, any combination with a regular key can be recorded in Settings (a modifier-only key such as Right Command needs Accessibility, and Option+Space stands in for it until then), and the transcript is left on the clipboard for you to paste with ⌘V. Managed Macs can pre-approve Accessibility for Pladder with an MDM Privacy Preferences Policy Control (PPPC) profile, which needs no prompt at all.

Apple Intelligence, if it is on, powers the optional Dictate and polish key; nothing else needs it.
3. **Wait for the model.** The first run downloads the Parakeet TDT v3 CoreML models from Hugging Face into `~/Library/Application Support/FluidAudio/Models` and compiles them. The menu bar icon shows progress, and the push-to-talk key is disabled until the engine is ready. This happens once; later launches load in well under a second.

Then click into any text field, hold **Option+Space**, say something, and let go.
Expand Down
12 changes: 8 additions & 4 deletions Package.swift
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,7 @@ let package = Package(
// Microphone capture and resampling.
.target(name: "PladderAudio", dependencies: ["PladderCore"]),

// Hotkey, pasteboard output, permissions, optional Foundation Models
// processor. AppKit lives here.
// Hotkey, pasteboard output, permissions. AppKit lives here.
.target(
name: "PladderSystem",
dependencies: ["PladderCore"],
Expand All @@ -46,10 +45,14 @@ let package = Package(
]
),

// Apple's on-device model behind the polish hotkey. The only target
// that imports FoundationModels.
.target(name: "PladderRefine", dependencies: ["PladderCore"]),

// The menu bar app.
.executableTarget(
name: "Pladder",
dependencies: ["PladderCore", "PladderAudio", "PladderSystem", "PladderEngines"],
dependencies: ["PladderCore", "PladderAudio", "PladderSystem", "PladderEngines", "PladderRefine"],
// Info.plist is copied into the .app by scripts/bundle.sh; SwiftPM
// refuses to treat it as a resource, so keep it out of the bundle.
exclude: ["Resources/Info.plist"],
Expand All @@ -70,12 +73,13 @@ let package = Package(
// engines, or run the benchmark (see docs/BENCHMARKS.md).
.executableTarget(
name: "PladderCLI",
dependencies: ["PladderCore", "PladderEngines", "PladderAudio", "PladderBench"]
dependencies: ["PladderCore", "PladderEngines", "PladderAudio", "PladderBench", "PladderRefine"]
),

.testTarget(name: "PladderCoreTests", dependencies: ["PladderCore"]),
.testTarget(name: "PladderAudioTests", dependencies: ["PladderAudio"]),
.testTarget(name: "PladderBenchTests", dependencies: ["PladderBench"]),
.testTarget(name: "PladderSystemTests", dependencies: ["PladderSystem"]),
.testTarget(name: "PladderRefineTests", dependencies: ["PladderRefine"]),
]
)
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@ You type long prompts all day. Speech is three to four times faster than typing,
- **Nothing leaves your Mac.** The speech model runs on the Neural Engine. There is no account, no server, no telemetry, and the app makes no network requests after the one-time model download.
- **It knows your words.** A dictionary turns what the model hears into what you meant. "clode code" becomes "Claude Code", "get hub" becomes "GitHub", every time, at zero cost in latency. List a word on its own and near misses of it are repaired too, so "Chat G P T" comes out as "ChatGPT" without you predicting every way the model might mangle it.
- **It skips over the ums.** Hesitation sounds — "uh", "um", German "äh"/"ähm", Spanish "eh" — are dropped before the text is pasted. Still no cost in latency. Only English, German and Spanish fillers are covered for now; other languages pass through unchanged.
- **Polish it when you want to.** Hold a second key instead and Apple's on-device model cleans the transcript before it is pasted: "wait, no, Friday" becomes "Friday", spoken numbers become digits, "first… second…" becomes a list. A second or two, still on your Mac, and only when you ask.
- **It behaves like part of macOS.** A menu bar app with a Liquid Glass status pill, a standard settings window, and nothing in the Dock.
- **Twenty-five languages, detected automatically.** English, German, Spanish, French and the rest of Europe in the same session, with no setting to flip.
- **Free and MIT.** The source is here. Read it, build it, change it.
Expand Down Expand Up @@ -149,6 +150,9 @@ Option+Space is Alfred's default hotkey and a common Raycast choice, and it type
**Can it press Return for me?**
Yes. Press the send key, Right Option by default, at any point while you hold the push-to-talk key, and Return is pressed after the text is pasted. That sends a chat message or runs a terminal command without touching the keyboard again. The send key can be changed in Settings, like the push-to-talk key.

**Can it clean up what I said?**
Record a key for Dictate and polish in Settings and hold that instead. The dictation goes through Apple Intelligence on your Mac before it is pasted, which takes a second or two. It needs Apple Intelligence turned on in System Settings; without it that key pastes the text as dictated.

**What about long dictations?**
Recordings stop at 120 seconds, so a lost key-up never leaves the microphone on. Audio longer than 15 seconds is transcribed in overlapping windows.

Expand Down
20 changes: 18 additions & 2 deletions Sources/Pladder/AppModel.swift
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import os
import PladderSystem
import PladderAudio
import PladderEngines
import PladderRefine

/// Composition root. Builds the engine registry, the settings store and the
/// coordinator, owns the overlay, and exposes everything the UI needs.
Expand Down Expand Up @@ -43,6 +44,13 @@ final class AppModel {
/// someone visits System Settings. Never on a key press.
private(set) var systemShortcuts: Set<Hotkey> = []

/// Whether Apple Intelligence can polish right now, for the polish key's
/// row. Polled with the permissions: it can be switched on or off in
/// System Settings while the app runs, and the read is cheap. The
/// coordinator never asks; an unavailable model makes the polish key a
/// plain dictation on its own.
private(set) var polishAvailability: OnDeviceModelAvailability = TranscriptPolisher.availability

/// The default chord, standing in for a stored chord Carbon cannot
/// register while Accessibility is missing. The stored chord is never
/// rewritten and returns with the grant.
Expand Down Expand Up @@ -195,6 +203,7 @@ final class AppModel {
outputMuter: OutputMuteController(
control: CoreAudioOutputMute(),
log: { Self.muteLog.info("\($0, privacy: .public)") }),
refiner: TranscriptPolisher(),
hotkeyMonitor: trusted ? tapHotkey : carbonHotkey,
makePipeline: { s in
ProcessorPipeline(processorFactories.map { $0(s) }, onFailure: { id, error in
Expand Down Expand Up @@ -263,11 +272,15 @@ final class AppModel {
guard let released = releaseInstant else { return }
releaseInstant = nil
let total = Self.seconds(instant - released)
// A polish cycle gets its own line, so the plain one stays the
// number the benchmark rule watches; one predicate finds both.
let polish = timing.polish.map { "polish \(fmt($0)), " } ?? ""
let label = timing.polish == nil ? "release-to-paste" : "polished release-to-paste"
let stages = "stop \(fmt(timing.captureStop)), engine \(fmt(timing.engine)), " +
"process \(fmt(timing.processing)), paste \(fmt(timing.insert))"
"process \(fmt(timing.processing)), \(polish)paste \(fmt(timing.insert))"
Self.timing.log(
"""
release-to-paste \(total, format: .fixed(precision: 3), privacy: .public) s: \(stages, privacy: .public); \
\(label, privacy: .public) \(total, format: .fixed(precision: 3), privacy: .public) s: \(stages, privacy: .public); \
audio \(transcript.audioDuration, format: .fixed(precision: 1), privacy: .public) s, \
engine-time \(transcript.processingTime, format: .fixed(precision: 3), privacy: .public) s
"""
Expand All @@ -287,6 +300,8 @@ final class AppModel {
func refreshPermissions() {
accessibilityTrusted = Permissions.isAccessibilityTrusted
microphoneStatus = Permissions.microphoneStatus
let polish = TranscriptPolisher.availability
if polish != polishAvailability { polishAvailability = polish }
let sustained = secureInput.observe(SecureInput.isEnabled)
if sustained != secureInputSustained { secureInputSustained = sustained }
// Granting Accessibility upgrades the hotkey to the tap; revoking it
Expand Down Expand Up @@ -411,6 +426,7 @@ final class AppModel {
switch coordinator.state {
case .recording: return String(localized: "Recording…")
case .transcribing: return String(localized: "Transcribing…")
case .polishing: return String(localized: "Polishing…")
case .inserting: return String(localized: "Inserting…")
case .error(let failure): return String(localized: "Error: \(failure.text)")
case .copied: return String(localized: "Copied — press ⌘V")
Expand Down
4 changes: 2 additions & 2 deletions Sources/Pladder/MenuBarIcon.swift
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ enum MenuBarIcon {
/// Bars follow the input level, quantised to `levelSteps` so the
/// image cache stays bounded.
case recording(step: Int)
/// Transcribing or inserting: dimmed to read as "busy".
/// Transcribing, polishing or inserting: dimmed to read as "busy".
case busy
/// Engine unavailable or an error: slashed like `mic.slash`.
case off
Expand All @@ -32,7 +32,7 @@ enum MenuBarIcon {
// input level means.
let amp = CGFloat(WaveformMeter.amplitude(for: level))
self = .recording(step: Int((amp * CGFloat(Self.levelSteps - 1)).rounded()))
case .transcribing, .inserting: self = .busy
case .transcribing, .polishing, .inserting: self = .busy
case .unavailable, .error: self = .off
}
}
Expand Down
Loading
Loading