Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ swift run -c release pladder-cli polish-set docs/polish-set.json --model s1-mini
| Learned corrections | After a paste the field is watched through Accessibility for up to 60 s; a word the user corrects that passes a token diff, a phonetic gate (Soundex or edit distance ≤ 2) and a yes/no review by the on-device model becomes one menu line, "Learned “x” → “y”? Add / Dismiss" | Nothing runs before Cmd+V is posted: the hook is in `AppModel.handle(.inserted)`, the watcher lives on its own thread and reads only the pasted range plus a margin, the review runs on a detached task. Present only with Accessibility and Apple Intelligence, absent otherwise, no setting, no change to the menu bar glyph. Dismissed pairs go to `dismissed-corrections.json`, not settings, so a bug there can never cost the dictionary. Pure case changes are never proposed. Terminals and TUIs expose a screen buffer, not a field, so nothing is learned there |
| Toggle key | Another recordable chord, off by default. A chord of its own latches at release however long the press; equal to the push-to-talk chord it makes that key hybrid: a tap under 400 ms latches, a longer hold stops at release. The next press of any chord, Escape or the 10 min cap ends a latched recording. The overlay is the same as for a held recording, in and out; only the menu's status line says which key stops it | Two-minute dictations should not need a key held for two minutes. Handy and VoiceInk default to hybrid on one key; here it is opt-in, because a stray tap would otherwise leave the microphone open until the cap pastes two minutes of room noise. Hold, toggle and hybrid are decided in `HotkeyGestureTracker`, a clockless value type timed by the instant each monitor stamps on its events, so both monitors behave alike and a press that waits for the microphone cannot make the next release look longer. A same-chord press within 50 ms of its release is a bounce (some Bluetooth keyboards do this mid-hold): it never acts, and the first one seen turns on a 50 ms settle before every stopping release for the rest of the run, so only a keyboard that needs it pays for it and the release path is otherwise untouched. No separate press debounce: both monitors already report alternating presses and releases. Without Accessibility the toggle chord registers with Carbon like the key; one Carbon cannot register has no stand-in, except that a hybrid chord follows the key's. With a lone modifier as a hybrid key, the Command of a later Cmd+C ends a latched recording |
| Escape | Discards a recording without transcribing and plays the stop sound; taken only while a recording is on | Never taken globally, so Escape keeps closing dialogs. On the tap `HotkeyChordSet` catches it before the chord trackers, so the interrupted-press rule never sees it, and Escape with the chord's own modifiers held still counts; Carbon registers the bare key around each recording, from the main queue so the release path never waits on it. Under Secure Event Input a modifier-only chord stays on the tap, where no key-down arrives, so Escape cannot cancel there |
| Output | Clipboard + simulated Cmd+V; the old clipboard is restored off the critical path | Universal, fast |
| Output | Clipboard + simulated Cmd+V; the transcript is a pasteboard promise and the old clipboard comes back off the critical path, 400 ms after Cmd+V if the target app has read it by then, otherwise 200 ms after it does, or after 8 s if nothing does | Universal, fast. Waiting for the read as well as the 400 ms: an app busy when Cmd+V arrives reads late, and the timer alone handed it the old clipboard (issue #40). A read never shortens the 400 ms, because Chromium sometimes reads once before the real paste and the pasteboard serves the second read itself. The promise is served on the main thread, so a busy Pladder main thread delays the target's read; the `paste` log line records when it came. It relies on the nspasteboard.org markers: without them Universal Clipboard reads every write within ~15 ms, which would pass for the paste |
| Polish model | A picker under the toggle: Apple Intelligence (default), S1-mini by Superwhisper at full precision (f16, 1.5 GB) or 8-bit (Q8_0, 805 MB), the S1-mini files run by llama.cpp on the GPU. `PolishRouter` is the coordinator's one refiner and forwards to the chosen model | On the polish set (docs/BENCHMARKS.md) S1-mini is more accurate than Apple's model and three to five times faster, and translated nothing where Apple's model translated two dictations. It is Qwen3-0.6B fine-tuned for dictation cleanup, trained on English, and needs its own fixed prompt. llama.cpp comes as its prebuilt XCFramework, pinned by release and checksum, because MLX Swift needs Xcode to build its Metal shaders; `bundle.sh` embeds and signs the framework. The GPU leaves the Neural Engine to Parakeet. A file downloads only when polish is on and the model picked, is loaded at the first key-down with the fixed prompt prefix decoded once, and is freed on a switch or with polish off. The licence asks for the name as "S1-mini" by "Superwhisper" |
| Post-processing | Filler remover, dictionary replacer, fuzzy custom-word corrector, whitespace normaliser, spoken punctuation, in that order | No latency, no network. Spoken punctuation takes only phrases that are never ordinary words ("period", "Punkt", "punto" stay words), runs last because the whitespace step folds line breaks, and needs Spanish evidence for "coma"; Parakeet already writes numbers as digits. An earlier Apple Intelligence step was removed from this path unmeasured; the model is back behind the experimental polish toggle only |
| Mute while dictating | Off by default; `kAudioDevicePropertyMute` on the default output device 200 ms into a recording, restored off the release path | Music or a call otherwise goes into the microphone. The delay means a tap-and-release never toggles anything; a device the user had already muted is left alone, and the device that was muted is the one unmuted even if the default changed meanwhile |
Expand Down Expand Up @@ -78,7 +78,7 @@ swift run -c release pladder-cli polish-set docs/polish-set.json --model s1-mini
- **Long recordings.** FluidAudio's encoder window is 15 s. Longer audio is split into windows and stitched, and seams can drop or duplicate words. Those windows now run while the key is held rather than at release, so the wait is flat with length, but the seam risk is unchanged: it is the same layout and the same merge. The paced benchmark guards it by requiring the text to be byte-identical to transcribing the whole recording at once, and the 30 s to 10 min fixtures watch the word error rate.
- **Polish latency and quality.** Apple's model takes about 1.5 to 2 seconds warm on an M1 for a typical dictation, S1-mini 0.3 to 0.5, and both are capped at eight; past the cap the text is pasted as dictated. A small model can still rewrite rather than clean; for Apple's the guided `cleanedText` field, greedy sampling, inline examples and naming the transcript's language keep that rare. S1-mini's known slips are in docs/BENCHMARKS.md, the worst a German number word read wrong, which real dictation rarely feeds it since Parakeet writes digits. `pladder-cli polish-set` is where model and prompt changes are judged.
- **Polish model memory.** S1-mini stays loaded while chosen and polish is on: about 0.8 or 1.5 GB of weights plus a 235 MB key-value cache for its 2,048-token context. Long dictations are polished in chunks of about 250 words to fit.
- **Clipboard clobbering.** Output saves the pasteboard, pastes, and restores it after a short delay.
- **Clipboard clobbering.** Output saves the pasteboard, pastes, and restores it once the target app has read the transcript. Nothing read within 8 s leaves the transcript on the clipboard that long, never stale text in the target.
- **Permissions.** Accessibility and Microphone grants are keyed to the code signature. `bundle.sh` signs with an Apple Development or Developer ID certificate when one is in the keychain; an ad-hoc signature changes on every build and resets both grants.
- **Lost key-up.** When the 10 min watchdog fires, the coordinator transcribes and pastes as if the user had released. Discarding is probably the right behaviour; tracked separately.

Expand Down
Loading
Loading