Skip to content
Merged
2 changes: 2 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@ swift run -c release pladder-cli polish <text file> # run the polish prompt ov
| Without Accessibility | Carbon `RegisterEventHotKey` plus clipboard-only output | A standard account cannot grant Accessibility without an admin. Carbon needs no permission but wants exactly one regular key and collapses left and right, so modifier-only chords are refused in the recorder; the transcript is left on the clipboard and the overlay says "press ⌘V". A stored chord Carbon cannot register, a lone Right Command say, is stood in for by the default Option+Space and the menu names it; the stored chord returns with the grant. `CopySymbolicHotKeys` only feeds the warning that an enabled macOS shortcut owns the recorded chord. `AppModel` polls the grant every two seconds and swaps the monitor in both directions |
| Send key | Press Right Option (configurable) while the hotkey is held and Return is posted 50 ms after Cmd+V | Sends a chat message or runs a command without a second trip to the keyboard; the Return is posted from a detached task so it stays off the release-to-paste path |
| Polish hotkey | "Dictate and polish": a second recordable chord, off by default. A dictation started with it runs the usual pipeline, then Apple's on-device model (FoundationModels, `PladderRefine`) with a fixed cleanup prompt, then pastes | Self-corrections, spoken punctuation, number words and lists are beyond the deterministic processors, and the model runs on device with nothing to download. It costs one to three seconds, so it never touches the normal hotkey's path: the branch is one Bool read; the session is created and prewarmed at key-down; transcripts under four words skip it; anything the model cannot do (Apple Intelligence off, refusal, the 8 s timeout) pastes the text as dictated. Logged as its own `polished release-to-paste` line |
| Learned corrections | After a paste the field is watched through Accessibility for up to 60 s; a word the user corrects that passes a token diff, a phonetic gate (Soundex or edit distance ≤ 2) and a yes/no review by the on-device model becomes one menu line, "Learned “x” → “y”? Add / Dismiss" | Nothing runs before Cmd+V is posted: the hook is in `AppModel.handle(.inserted)`, the watcher lives on its own thread and reads only the pasted range plus a margin, the review runs on a detached task. Present only with Accessibility and Apple Intelligence, absent otherwise, no setting, no change to the menu bar glyph. Dismissed pairs go to `dismissed-corrections.json`, not settings, so a bug there can never cost the dictionary. Pure case changes are never proposed. Terminals and TUIs expose a screen buffer, not a field, so nothing is learned there |
| Output | Clipboard + simulated Cmd+V; the old clipboard is restored off the critical path | Universal, fast |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, in that order | No latency, no network. An earlier Apple Intelligence step was removed from this path unmeasured; the model is back behind the polish hotkey only |
| Mute while dictating | Off by default; `kAudioDevicePropertyMute` on the default output device 200 ms into a recording, restored off the release path | Music or a call otherwise goes into the microphone. The delay means a tap-and-release never toggles anything; a device the user had already muted is left alone, and the device that was muted is the one unmuted even if the default changed meanwhile |
Expand All @@ -58,6 +59,7 @@ swift run -c release pladder-cli polish <text file> # run the polish prompt ov
- Adding an engine: implement `TranscriptionEngine` in its own file under `PladderEngines`, register it in the `EngineRegistry` built in `AppModel`. One file plus one registry line; the settings picker reads the registry. The engine lifecycle — building, loading, status polling and swapping — lives in `EngineLoader`.
- Adding a processor: implement `TextProcessor` in its own file, append a factory to `processorFactories` in `AppModel`. The pipeline is rebuilt when settings change, never per dictation. A processor sits on the critical path, so the benchmark rule applies.
- Adding a prompt: build an `OnDeviceLanguageModel(instructions:)` in `PladderRefine` and call `respond(to:)` or `respond(to:generating:)`; availability, prewarm, timeout and the drain of an abandoned call come with it. The coordinator only ever sees `TranscriptRefiner`.
- The correction learner's two seams are protocols in `PladderCore`, `PastedTextObserver` and `CorrectionReviewer`, with fakes in the tests; the Accessibility and Foundation Models implementations live in `PladderSystem` (`AXPasteObserver`) and `PladderRefine` (`FoundationModelsCorrectionReviewer`).
- Adding a language: add a `<code>` localization to both catalogs; nothing else. Adding a *string*: the key is the exact English text, and `PladderCore` never holds one — it emits an enum case and `Sources/Pladder/StatusText.swift` words it.
- Engine and capture are actors. The coordinator is `@MainActor` because it drives UI. It owns the state machine and nothing else; every dependency is injected, so tests run it with in-memory fakes.

Expand Down
7 changes: 5 additions & 2 deletions INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ This compiles a release build, wraps it into `Pladder.app`, signs it, copies it
## First launch

1. **Grant Microphone** when macOS asks. That is what records your voice.
2. **Grant Accessibility** when prompted. System Settings opens on the Accessibility list; switch Pladder on. That is what lets Pladder see the push-to-talk key in other apps and paste the result. It does not need Input Monitoring.
2. **Grant Accessibility** when prompted. System Settings opens on the Accessibility list; switch Pladder on. That is what lets Pladder see the push-to-talk key in other apps, paste the result, and notice when you correct a word it got wrong. It does not need Input Monitoring.

On a standard (non-administrator) account, ticking that box asks for an administrator password, so ask an admin to do it once — the grant is keyed to the app's code signature and survives updates. Without it Pladder still works in a reduced form: Option+Space works the same, any combination with a regular key can be recorded in Settings (a modifier-only key such as Right Command needs Accessibility, and Option+Space stands in for it until then), and the transcript is left on the clipboard for you to paste with ⌘V. Managed Macs can pre-approve Accessibility for Pladder with an MDM Privacy Preferences Policy Control (PPPC) profile, which needs no prompt at all.

Expand Down Expand Up @@ -112,7 +112,10 @@ Pladder pastes into whatever has keyboard focus when the key is released. Click
A push-to-talk key without a modifier is swallowed system wide while Pladder runs, so a bare letter or Space would become untypeable. Settings warns about this; use a modifier or a chord.

**Where are the settings?**
`~/Library/Application Support/Pladder/settings.json`, plain JSON. The dictionary can also be imported and exported from the Dictionary tab.
`~/Library/Application Support/Pladder/settings.json`, plain JSON. The dictionary can also be imported and exported from the Dictionary tab. Corrections you answered Dismiss to are kept beside it in `dismissed-corrections.json`; delete that file to be asked about them again.

**A word I corrected is never proposed.**
Proposals need Accessibility and Apple Intelligence, and appear in the menu bar menu up to a minute after the dictation, or as soon as you click away from the field. Only a word or two changed into something that sounds alike is proposed; a change of case, a rewording or an edit next to the dictation is not. Terminals, and apps such as Claude Code that run in one, show a screen rather than a text field, so nothing is learned there.

**How do I see the release-to-paste time?**
Every dictation logs one line:
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ You type long prompts all day. Speech is three to four times faster than typing,
- **One key, everywhere.** Hold Option+Space, or any key or chord you record. Speak. Release. The words land at the cursor in any app that takes text. No window to open, no button to click, no mode to leave.
- **Two keys, and it is sent.** Tap Right Option while you speak and the dictation goes out with Return the moment you let go. A prompt to an agent, a chat message, a shell command, without touching the keyboard again.
- **Nothing leaves your Mac.** The speech model runs on the Neural Engine. There is no account, no server, no telemetry, and the app makes no network requests after the one-time model download.
- **It knows your words.** A dictionary turns what the model hears into what you meant. "clode code" becomes "Claude Code", "get hub" becomes "GitHub", every time, at zero cost in latency. List a word on its own and near misses of it are repaired too, so "Chat G P T" comes out as "ChatGPT" without you predicting every way the model might mangle it.
- **It knows your words.** A dictionary turns what the model hears into what you meant. "clode code" becomes "Claude Code", "get hub" becomes "GitHub", every time, at zero cost in latency. List a word on its own and near misses of it are repaired too, so "Chat G P T" comes out as "ChatGPT" without you predicting every way the model might mangle it. Correct a word by hand after a dictation and Pladder offers to remember it — checked on device, added only when you say so.
- **It skips over the ums.** Hesitation sounds — "uh", "um", German "äh"/"ähm", Spanish "eh" — are dropped before the text is pasted. Still no cost in latency. Only English, German and Spanish fillers are covered for now; other languages pass through unchanged.
- **Polish it when you want to.** Hold a second key instead and Apple's on-device model cleans the transcript before it is pasted: "wait, no, Friday" becomes "Friday", spoken numbers become digits, "first… second…" becomes a list. A second or two, still on your Mac, and only when you ask.
- **It behaves like part of macOS.** A menu bar app with a Liquid Glass status pill, a standard settings window, and nothing in the Dock.
Expand Down Expand Up @@ -93,7 +93,7 @@ Why it is fast is written up in [docs/PERFORMANCE.md](docs/PERFORMANCE.md). The
Privacy here is not a policy, it is how the thing is built.

- **Audio never leaves the Mac.** The microphone is open only while the key is held, and macOS shows the orange indicator only then. Audio goes from the microphone to the Neural Engine and is discarded.
- **Text never leaves the Mac.** The transcript exists long enough to be pasted. Your previous clipboard is put back afterwards.
- **Text never leaves the Mac.** The transcript exists long enough to be pasted. Your previous clipboard is put back afterwards. After a paste, Pladder watches the field it pasted into for up to a minute through Accessibility, to notice when you fix a word; it reads only the pasted words and a little context, keeps nothing, and asks before adding anything to the dictionary.
- **No network.** The only request Pladder ever makes is the one-time download of the speech model from Hugging Face, about 700 MB, on first launch. After that it works with Wi-Fi off. There is no update check, no crash reporter, no analytics.
- **No account.** Nothing to sign up for, nothing to log in to, nothing to cancel.
- **Auditable.** The app is about five thousand lines of Swift under the MIT license, and none of them open a network connection. The model download is FluidAudio's, and it runs once.
Expand Down
95 changes: 94 additions & 1 deletion Sources/Pladder/AppModel.swift
Original file line number Diff line number Diff line change
Expand Up @@ -91,6 +91,18 @@ final class AppModel {
/// the app quitting mid-recording, a device vanishing — is silent and
/// baffling otherwise, so both ends of it are logged `.public`.
private nonisolated static let muteLog = Logger(subsystem: "de.dinooo13.pladder", category: "mute")
/// The correction learner's stages. They quote the user's words, so the
/// text is `.private`: `log show` prints it only with private data on.
private nonisolated static let learningLog = Logger(subsystem: "de.dinooo13.pladder", category: "learning")

/// Watches a field after the paste and proposes what the user corrected.
private let learner: CorrectionLearner
private let dismissedCorrections: DismissedCorrections
private let proposalRelay = ProposalRelay()
/// Corrections the model agreed with, newest first, waiting in the menu
/// for Add or Dismiss. Kept for the app's life or until answered.
private(set) var proposals: [CorrectionProposal] = []
static let maximumProposals = 3

/// Settings live in the coordinator (it reacts to hotkey/engine changes);
/// this forwards and persists. Applying the appearance covers every
Expand Down Expand Up @@ -195,7 +207,7 @@ final class AppModel {
let events = self.events
let trusted = Permissions.isAccessibilityTrusted
hotkeyUsesTap = trusted
coordinator = DictationCoordinator(
let coordinator = DictationCoordinator(
settings: initial,
registry: registry,
capture: AVAudioEngineCapture(),
Expand All @@ -212,9 +224,26 @@ final class AppModel {
},
onEvent: { [events] event in events.send(event) }
)
self.coordinator = coordinator

// Nothing here runs before the paste: `handle(.inserted)` hands the
// pasted text over, and the learner watches and reviews on its own
// thread and task.
let dismissed = DismissedCorrections(url: Self.dismissedCorrectionsURL)
dismissedCorrections = dismissed
let proposalRelay = self.proposalRelay
learner = CorrectionLearner(
observer: AXPasteObserver(),
reviewer: FoundationModelsCorrectionReviewer(),
dismissed: dismissed,
dictionary: { await coordinator.settings.dictionary },
log: { Self.learningLog.info("\($0, privacy: .private)") },
onProposal: { proposalRelay.send($0) }
)

overlay = OverlayController(coordinator: coordinator)
events.handler = { [weak self] event, at in self?.handle(event, at: at) }
proposalRelay.handler = { [weak self] proposal in self?.propose(proposal) }
}

static var settingsURL: URL {
Expand All @@ -223,6 +252,12 @@ final class AppModel {
.appending(path: "Library/Application Support/Pladder/settings.json")
}

/// The corrections the user dismissed, beside the settings but not in
/// them (see `DismissedCorrections`).
static var dismissedCorrectionsURL: URL {
settingsURL.deletingLastPathComponent().appending(path: "dismissed-corrections.json")
}

/// One-time migration from the pre-rename location. The dictionary and
/// hotkey settings were kept in `~/Library/Application Support/SpeakUp/`
/// before the app was called Pladder; copy them across exactly once, only
Expand Down Expand Up @@ -269,6 +304,13 @@ final class AppModel {
releaseInstant = instant
if settings.playSounds { SoundPlayer.playStop() }
case .inserted(let transcript, let timing):
defer {
// Strictly after the paste and off the measured window, which
// ended when the coordinator emitted this event: the learner
// spawns its own task and reads the field on its own thread.
// Without the grant nothing was pasted, only copied.
if accessibilityTrusted { learner.pasted(transcript.text) }
}
guard let released = releaseInstant else { return }
releaseInstant = nil
let total = Self.seconds(instant - released)
Expand All @@ -290,6 +332,46 @@ final class AppModel {
}
}

// MARK: Learned corrections

private func propose(_ proposal: CorrectionProposal) {
let key = proposal.pair.key
guard !proposals.contains(where: { $0.pair.key == key }),
!Self.dictionary(settings.dictionary, has: proposal.pair.heard) else { return }
proposals = Array(([proposal] + proposals).prefix(Self.maximumProposals))
}

/// Adds `heard → corrected` to the dictionary, overwriting a rule with the
/// same `from` the way the Dictionary tab's import does. Through the
/// settings setter, so it is saved and the next dictation uses it.
func acceptProposal(_ proposal: CorrectionProposal) {
proposals.removeAll { $0.id == proposal.id }
let from = proposal.pair.heard.lowercased()
var dictionary = settings.dictionary
let entry = DictionaryEntry(from: proposal.pair.heard, to: proposal.pair.corrected)
if let index = dictionary.firstIndex(where: {
$0.from.trimmingCharacters(in: .whitespaces).lowercased() == from
}) {
dictionary[index].to = entry.to
dictionary[index].matchCase = false
} else {
dictionary.append(entry)
}
settings.dictionary = dictionary
}

/// Drops the line and remembers the pair so it is never proposed again.
func dismissProposal(_ proposal: CorrectionProposal) {
proposals.removeAll { $0.id == proposal.id }
let dismissed = dismissedCorrections
Task { await dismissed.dismiss(proposal.pair) }
}

private static func dictionary(_ entries: [DictionaryEntry], has heard: String) -> Bool {
let key = heard.lowercased()
return entries.contains { $0.from.trimmingCharacters(in: .whitespaces).lowercased() == key }
}

private static func seconds(_ duration: Duration) -> Double {
let parts = duration.components
return Double(parts.seconds) + Double(parts.attoseconds) / 1e18
Expand Down Expand Up @@ -493,6 +575,17 @@ final class AppModel {
}
}

/// Bridges the learner's proposals onto the main actor, and lets the learner
/// be built before `self` exists, as `EventRelay` does for the coordinator.
@MainActor
final class ProposalRelay {
var handler: ((CorrectionProposal) -> Void)?

nonisolated func send(_ proposal: CorrectionProposal) {
Task { @MainActor in self.handler?(proposal) }
}
}

/// Bridges the coordinator's nonisolated `onEvent` callback back onto the main
/// actor, and lets us hand the coordinator a callback before `self` exists.
///
Expand Down
14 changes: 14 additions & 0 deletions Sources/Pladder/PladderApp.swift
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,20 @@ private struct MenuContent: View {
}
}

// A correction the user made by hand after a paste, which the
// on-device model agreed is a reusable spelling. A menu line rather
// than an overlay toast: the pill is click-through by construction,
// and a proposal arriving a minute later should wait, not interrupt.
if !model.proposals.isEmpty {
Divider()
ForEach(model.proposals) { proposal in
Menu("Learned “\(proposal.pair.heard)” → “\(proposal.pair.corrected)”?") {
Button("Add") { model.acceptProposal(proposal) }
Button("Dismiss") { model.dismissProposal(proposal) }
}
}
}

Divider()

Button("Settings…") {
Expand Down
30 changes: 30 additions & 0 deletions Sources/Pladder/Resources/Localizable.xcstrings
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,16 @@
}
}
},
"Add": {
"localizations": {
"de": {
"stringUnit": {
"state": "translated",
"value": "Hinzufügen"
}
}
}
},
"Add a rule": {
"localizations": {
"de": {
Expand Down Expand Up @@ -281,6 +291,16 @@
}
}
},
"Dismiss": {
"localizations": {
"de": {
"stringUnit": {
"state": "translated",
"value": "Verwerfen"
}
}
}
},
"Download failed: %@ (Retry resumes it)": {
"localizations": {
"de": {
Expand Down Expand Up @@ -501,6 +521,16 @@
}
}
},
"Learned “%@” → “%@”?": {
"localizations": {
"de": {
"stringUnit": {
"state": "translated",
"value": "„%1$@“ → „%2$@“ ins Wörterbuch?"
}
}
}
},
"Leave Heard as empty to list a word that should be repaired when it comes out nearly right.": {
"localizations": {
"de": {
Expand Down
Loading
Loading