Skip to content

Polish: model picker (Apple Intelligence, S1-mini), spoken punctuation, filler fix - #64

Merged
dinooo13 merged 6 commits into
mainfrom
polish-model-picker
Sep 27, 2026
Merged

dinooo13 merged 6 commits into
mainfrom
polish-model-picker

Conversation

@dinooo13

Copy link
Copy Markdown
Owner

Summary

A model picker for the experimental polish, the two text bugs the model evaluation turned up, and the parts of the cleanup that code does better than a model.

  • Model picker under "Polish dictations" (Processing → Experimental): Apple Intelligence (default, unchanged), S1-mini by Superwhisper at full precision (f16, 1.5 GB) or 8-bit (Q8_0, 805 MB). S1-mini is Qwen3-0.6B fine-tuned for dictation cleanup and runs through llama.cpp on the GPU, leaving the Neural Engine to Parakeet.
    • PolishRouter is the coordinator's one refiner and forwards to the chosen model.
    • ModelFiles downloads a file once from Hugging Face, only while polish is on and that model is picked. Each file is pinned to a commit and moved into place only after its SHA-256 matches. The picker shows progress, and a failed download offers Try Again.
    • S1MiniPolisher uses the prompt the model was trained on and decodes the fixed prefix once at load. It chunks long dictations to fit its 2,048-token context and keeps the 8 s cap and the paste-as-dictated fallback. The model loads at the first key-down and is freed on a switch or when polish goes off.
    • llama.cpp's Metal setup is warmed as soon as S1-mini is chosen. The first run after an install compiles shaders for about 7 s on an M1; macOS caches them after that.
    • llama.cpp is its prebuilt XCFramework, pinned by release and checksum. MLX Swift needs Xcode to build its Metal shaders, and this project has no Xcode project. bundle.sh embeds the framework, drops the Intel slice and signs it with the app's identity.
  • Spoken punctuation is a new processor, last in the pipeline. It turns "comma", "question mark", "new paragraph", "Fragezeichen", "signo de interrogación" and similar into the marks, absorbing the punctuation Parakeet adds around them ("verschieben? Fragezeichen." → "verschieben?").
    • Only phrases that are never ordinary words are taken: "period", "Punkt" and "punto" stay words, and Spanish "coma" needs Spanish language evidence. A spoken comma between digits is the decimal comma.
    • Parakeet already writes numbers as digits in English and German, so no number converter was needed.
  • Filler fix: a filler that opens a later sentence now hands on its capital. "…geschrieben. Ähm, beim Timeout…" used to paste as "…geschrieben. beim Timeout…".
  • CLI: pladder-cli polish --model …, and pladder-cli polish-set docs/polish-set.json --model …, which runs a model over 42 dictations in English, German and Spanish and prints each answer, word error per language, exact matches and timings.
  • Docs: CLAUDE.md, BENCHMARKS.md and the README. The README still described polish as a second hotkey.

Polish models (M1 MacBook Air 16 GB, macOS 27.0)

From pladder-cli polish-set, after the app's processors, warm, greedy:

Model Word error (en / de / es) Exact of 42 Polish median p90
Apple Intelligence 0.108 (0.037 / 0.152 / 0.135) 18 1.66 s 2.07 s
S1-mini, full precision 0.071 (0.021 / 0.115 / 0.078) 22 0.49 s 0.86 s
S1-mini, 8-bit 0.091 (0.008 / 0.115 / 0.149) 21 0.34 s 0.61 s

A cold S1-mini start (load plus polish) is 1.1 s at 8-bit and 1.7 s at full precision. The load starts at key-down, while the user speaks.

What the numbers hide:

  • Apple's model: translated two dictations into English, turned "halb acht" into "8:00", and never wrote a multi-line list.
  • S1-mini: translated nothing and resolved the German and Spanish self-corrections, but:
    • it misreads the German number word "zweihundertfünfzig" as 150 (real dictation arrives as digits from Parakeet);
    • it garbles one German sentence around a code identifier;
    • it ignores "Nächster Punkt" and "Siguiente punto" as list cues.

LFM2.5-1.2B and Gemma 3 1B were tried with Pladder's prompt and dropped; details in BENCHMARKS.md.

Critical path

The engine and the coordinator's release path are unchanged, so the engine benchmark would not see this; I didn't rerun it. What does change is the processor pipeline (a new processor and the filler fix). Measured on the fixtures' text through the whole pipeline, median of 31, M1. "Typical" has a filler in every sentence; "worst" adds a spoken question mark to every sentence:

Fixture main, typical this branch, typical main, worst this branch, worst
10s 1.53 ms 1.97 ms 1.51 ms 1.62 ms
60s 1.71 ms 1.81 ms 1.73 ms 2.04 ms
2m 1.96 ms 2.14 ms 2.00 ms 2.61 ms
10m 3.77 ms 4.83 ms 4.13 ms 6.85 ms

At most 3 ms at 10 minutes, against a ~230 ms release-to-paste median. The 10 s difference is noise; the ~1.5 ms floor is the language recogniser the filler remover already calls. With polish on, the polish stage itself drops from about 1.7 s to 0.3–0.5 s with S1-mini.

Testing

  • swift test: all targets pass (338 + 21 + 16 + 10 + 6). New tests cover spoken punctuation, the filler fix, S1-mini's prompt format and chunking, model-file mapping and pinning, the download's checksum and failure paths (from a file URL), the polisher without its file, and the setting's persistence and fallback.
  • pladder-cli polish-set on all three models, which exercises the download, verification, load and generation end to end.
  • bundle.sh with a certificate and ad-hoc. Checked that an ad-hoc hardened-runtime binary refuses the framework ("different Team IDs"), hence no hardened runtime for ad-hoc builds.
  • A test copy with its own settings, launched from dist/, showed the new processor row and the picker. I did not dictate through the app with S1-mini selected; that end-to-end check is yours.

Notes

  • License: S1-mini is Apache 2.0 with a naming clause, so it appears as "S1-mini" by "Superwhisper". The 8-bit file is mradermacher's quantisation of the same release, because Superwhisper publishes only f16 and 4-bit.
  • Privacy: the Hugging Face exception now covers a polish model the user explicitly picks. Nothing downloads unless polish is on and S1-mini is chosen.
  • Possible follow-up: someone has published a German fine-tune, Joni000000000/s1-mini-de-v3. It could be judged with polish-set before adding it as a picker entry.

🤖 Generated with Claude Code

dinooo13 and others added 6 commits September 26, 2026 01:11
"Ich habe einen Test geschrieben. Ähm, beim Timeout …" came out as
"… geschrieben. beim Timeout …": only a filler that opened the whole
transcript handed its capital on. Any filler after a full stop, question
or exclamation mark now does too; an ellipsis trails off and does not.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Parakeet already writes numbers as digits and takes most spoken marks,
but "question mark", "Fragezeichen" and the like reach the paste as
words, usually with the model's own mark stuck to them ("verschieben?
Fragezeichen."). A deterministic step now turns the unambiguous ones
into the mark and absorbs the stray punctuation: comma, question and
exclamation marks, full stop, colon, semicolon and new paragraph.
"period", "Punkt" and "punto" stay words, and Spanish "coma" needs
Spanish language evidence. Between digits a spoken comma is the decimal
comma.

It runs last, after the whitespace step, which would otherwise fold its
paragraph breaks away. On the fixtures' text through the whole pipeline
(M1, median of 31): a typical dictation costs the same as on main up to
2 min and 1 ms more at 10 min; a spoken mark in every sentence of a
10 min dictation, 2.7 ms more.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Model picker under "Polish dictations": Apple Intelligence as before,
or S1-mini by Superwhisper at full precision (f16, 1.5 GB) or 8-bit
(Q8_0, 805 MB). S1-mini is Qwen3-0.6B fine-tuned for dictation cleanup
and runs through llama.cpp on the GPU, leaving the Neural Engine to
Parakeet. On the polish set it is more accurate than Apple's model and
three to five times faster: 0.49 s and 0.34 s median against 1.66 s.

- PolishRouter is the coordinator's one refiner and forwards to the
  chosen model, so switching needs no new coordinator.
- ModelFiles downloads a file once from Hugging Face, only while polish
  is on and that model is picked, pinned to a commit and moved into place
  only after its SHA-256 matched. The picker shows progress, and a failed
  download offers Try Again.
- S1MiniPolisher uses the prompt S1-mini was trained on, decodes its
  fixed prefix once at load, chunks long dictations to its 2,048-token
  context, and keeps the 8 s cap and the paste-as-dictated fallback. It
  loads at the first key-down and is freed on a switch or with polish
  off; llama.cpp's Metal setup is warmed as soon as S1-mini is chosen,
  since its first run after an install compiles shaders for seconds.
- llama.cpp is its prebuilt XCFramework, pinned by release and checksum.
  bundle.sh embeds it in Contents/Frameworks, drops the Intel slice and
  signs it with the app's identity; an ad-hoc build goes without the
  hardened runtime, which would otherwise refuse the framework.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
pladder-cli polish takes --model apple|s1-mini|s1-mini-8bit and fetches
an S1-mini file into the app's own model directory if needed. polish-set
runs a model over docs/polish-set.json, 42 dictations in English, German
and Spanish with the text that should be pasted, after the app's
processors, warm, and prints every answer, the word error rate per
language, exact matches and the polish time. It is how the models in the
picker were chosen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
BENCHMARKS gains the polish-set results for all three models and the
processor timings. CLAUDE.md records the picker, llama.cpp and the
widened download exception; the README's polish entries still described
the old second hotkey.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
polish and polish-set take --gguf <file> for a model file that is not in
the picker and --control <line> for another style than the app's, so a
fine-tune is judged the same way as the picker's models before it goes
in. Used on Joni000000000/s1-mini-de-v3, a German fine-tune of S1-mini:
no better than S1-mini 8-bit on the set's German, a self-correction
missed, so it stays out. The results are in BENCHMARKS.md.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dinooo13
dinooo13 merged commit f1995e0 into main Sep 27, 2026
1 check passed
@dinooo13
dinooo13 deleted the polish-model-picker branch September 28, 2026 18:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant