Polish: model picker (Apple Intelligence, S1-mini), spoken punctuation, filler fix - #64
Merged
Merged
Conversation
"Ich habe einen Test geschrieben. Ähm, beim Timeout …" came out as "… geschrieben. beim Timeout …": only a filler that opened the whole transcript handed its capital on. Any filler after a full stop, question or exclamation mark now does too; an ellipsis trails off and does not. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Parakeet already writes numbers as digits and takes most spoken marks,
but "question mark", "Fragezeichen" and the like reach the paste as
words, usually with the model's own mark stuck to them ("verschieben?
Fragezeichen."). A deterministic step now turns the unambiguous ones
into the mark and absorbs the stray punctuation: comma, question and
exclamation marks, full stop, colon, semicolon and new paragraph.
"period", "Punkt" and "punto" stay words, and Spanish "coma" needs
Spanish language evidence. Between digits a spoken comma is the decimal
comma.
It runs last, after the whitespace step, which would otherwise fold its
paragraph breaks away. On the fixtures' text through the whole pipeline
(M1, median of 31): a typical dictation costs the same as on main up to
2 min and 1 ms more at 10 min; a spoken mark in every sentence of a
10 min dictation, 2.7 ms more.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Model picker under "Polish dictations": Apple Intelligence as before, or S1-mini by Superwhisper at full precision (f16, 1.5 GB) or 8-bit (Q8_0, 805 MB). S1-mini is Qwen3-0.6B fine-tuned for dictation cleanup and runs through llama.cpp on the GPU, leaving the Neural Engine to Parakeet. On the polish set it is more accurate than Apple's model and three to five times faster: 0.49 s and 0.34 s median against 1.66 s. - PolishRouter is the coordinator's one refiner and forwards to the chosen model, so switching needs no new coordinator. - ModelFiles downloads a file once from Hugging Face, only while polish is on and that model is picked, pinned to a commit and moved into place only after its SHA-256 matched. The picker shows progress, and a failed download offers Try Again. - S1MiniPolisher uses the prompt S1-mini was trained on, decodes its fixed prefix once at load, chunks long dictations to its 2,048-token context, and keeps the 8 s cap and the paste-as-dictated fallback. It loads at the first key-down and is freed on a switch or with polish off; llama.cpp's Metal setup is warmed as soon as S1-mini is chosen, since its first run after an install compiles shaders for seconds. - llama.cpp is its prebuilt XCFramework, pinned by release and checksum. bundle.sh embeds it in Contents/Frameworks, drops the Intel slice and signs it with the app's identity; an ad-hoc build goes without the hardened runtime, which would otherwise refuse the framework. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
pladder-cli polish takes --model apple|s1-mini|s1-mini-8bit and fetches an S1-mini file into the app's own model directory if needed. polish-set runs a model over docs/polish-set.json, 42 dictations in English, German and Spanish with the text that should be pasted, after the app's processors, warm, and prints every answer, the word error rate per language, exact matches and the polish time. It is how the models in the picker were chosen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
BENCHMARKS gains the polish-set results for all three models and the processor timings. CLAUDE.md records the picker, llama.cpp and the widened download exception; the README's polish entries still described the old second hotkey. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
polish and polish-set take --gguf <file> for a model file that is not in the picker and --control <line> for another style than the app's, so a fine-tune is judged the same way as the picker's models before it goes in. Used on Joni000000000/s1-mini-de-v3, a German fine-tune of S1-mini: no better than S1-mini 8-bit on the set's German, a self-correction missed, so it stays out. The results are in BENCHMARKS.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A model picker for the experimental polish, the two text bugs the model evaluation turned up, and the parts of the cleanup that code does better than a model.
PolishRouteris the coordinator's one refiner and forwards to the chosen model.ModelFilesdownloads a file once from Hugging Face, only while polish is on and that model is picked. Each file is pinned to a commit and moved into place only after its SHA-256 matches. The picker shows progress, and a failed download offers Try Again.S1MiniPolisheruses the prompt the model was trained on and decodes the fixed prefix once at load. It chunks long dictations to fit its 2,048-token context and keeps the 8 s cap and the paste-as-dictated fallback. The model loads at the first key-down and is freed on a switch or when polish goes off.bundle.shembeds the framework, drops the Intel slice and signs it with the app's identity.pladder-cli polish --model …, andpladder-cli polish-set docs/polish-set.json --model …, which runs a model over 42 dictations in English, German and Spanish and prints each answer, word error per language, exact matches and timings.Polish models (M1 MacBook Air 16 GB, macOS 27.0)
From
pladder-cli polish-set, after the app's processors, warm, greedy:A cold S1-mini start (load plus polish) is 1.1 s at 8-bit and 1.7 s at full precision. The load starts at key-down, while the user speaks.
What the numbers hide:
LFM2.5-1.2B and Gemma 3 1B were tried with Pladder's prompt and dropped; details in BENCHMARKS.md.
Critical path
The engine and the coordinator's release path are unchanged, so the engine benchmark would not see this; I didn't rerun it. What does change is the processor pipeline (a new processor and the filler fix). Measured on the fixtures' text through the whole pipeline, median of 31, M1. "Typical" has a filler in every sentence; "worst" adds a spoken question mark to every sentence:
At most 3 ms at 10 minutes, against a ~230 ms release-to-paste median. The 10 s difference is noise; the ~1.5 ms floor is the language recogniser the filler remover already calls. With polish on, the polish stage itself drops from about 1.7 s to 0.3–0.5 s with S1-mini.
Testing
swift test: all targets pass (338 + 21 + 16 + 10 + 6). New tests cover spoken punctuation, the filler fix, S1-mini's prompt format and chunking, model-file mapping and pinning, the download's checksum and failure paths (from a file URL), the polisher without its file, and the setting's persistence and fallback.pladder-cli polish-seton all three models, which exercises the download, verification, load and generation end to end.bundle.shwith a certificate and ad-hoc. Checked that an ad-hoc hardened-runtime binary refuses the framework ("different Team IDs"), hence no hardened runtime for ad-hoc builds.dist/, showed the new processor row and the picker. I did not dictate through the app with S1-mini selected; that end-to-end check is yours.Notes
Joni000000000/s1-mini-de-v3. It could be judged withpolish-setbefore adding it as a picker entry.🤖 Generated with Claude Code