Skip to content

Release: podcaster features (long episodes, audio, publishing, speakers, Shorts) - #38

Merged
amide-init merged 43 commits into
mainfrom
dev
Sep 24, 2026
Merged

amide-init merged 43 commits into
mainfrom
dev

Conversation

@amide-init

Copy link
Copy Markdown
Owner

Merges everything on dev into main: 43 commits from PRs #30–#37, each already reviewed and merged into dev.

What's in it

PR Feature
#30 Long episodes: uploads up to 10 GB with progress, background transcription in ~10-minute chunks split at pauses, 720p editing proxy, waveform from a small audio track
#31 Podcast audio: -16/-14 LUFS loudness (measured gain plus a true-peak limiter), noise reduction, speaker leveling, rumble filter, 15s before/after preview
#32 Publishing: MP3/WAV export, AI chapters built into MP4/MP3 exports and copyable for YouTube, editable show notes and titles, TXT/Markdown transcripts
#33 Speakers: detection (gpt-4o-transcribe-diarize) kept consistent across chunks, rename everywhere, cut/restore everything one person said
#34 Clips / Shorts: AI highlights (gpt-5.6-luna), clips from a transcript selection, 9:16/1:1/16:9 reframing with an on-player crop preview, Shorts captions
#35 Fix: words silently dropped at Whisper segment boundaries (0.9% of words in a conversation)
#36 Mac app fixes: existing databases get new migrations on startup, and the app now passes FFPROBE_PATH
#37 Docs: Podcast Workflow guide, README and docs site refresh

Every AI feature is an explicit button. Models return text and sentence picks, never timestamps or edits, and every response is validated before use.

Pre-release security check

Run over all 43 commits (full history, not just the final diff, so a key added and later removed would still be caught):

  • No secrets: no OpenAI sk- keys, AWS/GitHub/Slack/Google keys, private keys, or hardcoded key/token/password values in any commit. No commit adds or removes an sk- string.
  • No sensitive files: no .env, database, settings.json, data/, media, build output, generated code or binaries. The only env-related change is a comment for MAX_UPLOAD_MB in .env.example. .gitignore still covers .env, databases, settings, server-bundle/, Tauri target/ and the generated Prisma client.
  • No personal info: no local machine paths, usernames or emails in code, docs or commit messages. App-data paths in the docs use generic ~ paths. The commit author email is the same one already on all 160 commits on main.
  • Code:
    • No shell execution added; ffmpeg is still argument-list only.
    • Every filter value comes from constants, validated enums or numbers.
    • The only raw SQL is the startup migrator, which runs the repo's own migration files with bound parameters.
    • File paths are built on the server and confined to DATA_DIR.
    • Clip, job and render IDs are checked to belong to the project in the URL.
    • Chapter titles are escaped in FFMETADATA (with an injection test).
    • AI output is rendered as text only.
    • Download filenames are sanitized.
  • No dependency changes.
  • Data sent to OpenAI: audio and transcript text for the AI buttons, which are all explicit. This is documented in the README and docs.

Known low-severity note (not a blocker): raising the upload limit to 10 GB also raised Bun's request-size limit for JSON endpoints, so an oversized JSON request could use a lot of memory. There's no auth and the app is local-only by design (claude.md §18), so anyone who can reach the server can already do more. A per-route body limit is a reasonable follow-up.

Verified on dev just now

  • Server: typecheck clean, 245 tests pass.
  • Client: production build OK, 100 tests pass, lint clean.
  • Each feature was tested end to end in Chrome during its PR. The Mac app was rebuilt and tested on a real existing install: the database migrated from 6 to 11 migrations with projects intact.

After merging

  • The Mac app needs a rebuild or a v*.*.* tag for CI to publish it. Existing installs upgrade their database automatically on first launch.
  • The docs site redeploys from main.

🤖 Generated with Claude Code

amide-init and others added 30 commits September 25, 2026 01:06
Long podcast episodes take minutes to transcribe, too long to hold an API
request open (claude.md section 14). A job row mirroring RenderJob lets the
work run in the background while the client polls status and per-chunk
progress. The finished transcript still lives in Transcript; this row only
tracks the run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…proxy

Deterministic argv builders for the long-episode pipeline: extract a mono
16kHz 32kbps speech track (~14MB/hour, well under Whisper's 25MB cap per
chunk), detect silences to choose chunk boundaries, slice chunks with an
exact re-encoded seek so merged timestamps stay precise, and build a 720p
editing proxy (claude.md section 16).

runFfmpegCapturingStderr exists because silencedetect reports its results
only as log lines, and probeDuration is needed to plan chunks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast episodes run 30-120 minutes and several GB, which broke both the
500MB upload cap and Whisper's 25MB file limit, and held the transcribe
request open for the whole run.

Transcription is now a background job (lib/ai/transcription-job.ts):
extract a small speech track, split it into ~10-minute chunks at detected
silences so no word is cut in half, transcribe 3 chunks at a time, then
shift and merge them onto the source timeline. The merge also drops
Whisper's end-of-clip repeats: a copy of the previous sentence made of
zero-length words, which chunking would otherwise produce at every chunk
end. Genuine repeated phrases with real timing are kept.

POST /transcribe now returns 202 + jobId, and the upload screen polls it
and shows chunk progress. Uploads use XHR so they can report progress,
and the upload cap is 10GB by default, configurable with MAX_UPLOAD_MB.
Jobs still running at startup were interrupted by a restart, so they are
marked failed; otherwise they would block retries forever.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Multi-GB 1080p/4K podcast originals stutter when the browser plays them,
and the timeline waveform fetched and decoded the entire video. For a
2-hour episode that is over 1GB of decoded audio at 48kHz.

After upload, a 720p proxy is built in the background when the source is
above 720p or over 300MB (claude.md section 16). It is served at a
separate ?variant=proxy URL rather than swapped in behind /video, so a
proxy that finishes mid-session can't change the bytes under an in-flight
Range request. Proxy creation is best-effort: on failure the editor plays
the original. Export always uses the original.

The waveform now decodes the small extracted speech track (GET /audio)
at 8kHz, and falls back to the video for projects transcribed before the
track existed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Support long podcast episodes: chunked background transcription, 720p proxy, 10GB uploads
Adds AudioSettings: a loudness target (off / -16 / -14 LUFS), noise
reduction (off / light / strong), speaker leveling, and an 80Hz rumble
filter. Stored as Project.audioSettingsJson and validated with Zod on
PATCH.

Every field is an enum or boolean, never a raw number, so nothing the
client sends can reach an ffmpeg filter string. A stored value that no
longer validates falls back to all-off. Defaults are all off, so existing
projects export exactly as before.

The migration also adds RenderJob.kind, which separates full exports
from the short audio previews added later in this series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cleanup chain runs after the cuts in this order: high-pass, afftdn
denoise, dynaudnorm leveling, then loudness. Each stage maps to a fixed
filter string (lib/ffmpeg/audio-filters.ts).

Loudness uses measured gain plus a peak limiter, not loudnorm's two-pass
linear mode. Speech is peaky, so the gain needed usually pushes peaks
past the ceiling, and loudnorm then silently falls back to dynamic mode
and undershoots: a -14 target measured -16.3 LUFS on a test clip.
resolveLoudnessGain measures the edited audio, applies the gain through
the real chain, and corrects with secant steps, since the limiter lowers
loudness more as it works harder. Corrections may add at most 6dB over
the first estimate; beyond that, AAC encoding pushed peaks to clipping
(+0.7 dBTP) on a very quiet source.

The ceiling is -2 dBTP rather than -1.5 because AAC encoding overshoots
the limiter; measured, -1.5 came out at -1.1 dBTP. Measured with ebur128
on real renders:
- -16 target: -15.9 to -16.2 LUFS
- -14 target: -14.1 LUFS

The audio edit graph (atrim + acrossfade) is now shared by the export,
the loudness measurement and an audio-only render, so all three process
exactly the same audio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The browser can't run ffmpeg's audio filters live, so there was no way
to hear the cleanup without a full export. POST /render/audio-preview
renders the same 15s of the edited program twice, once untouched and
once through the cleanup chain, starting from the playhead. Cuts inside
that window are respected.

The request carries the panel's current settings rather than reading
them from the project row, so the preview matches what's on screen even
while the settings PATCH is still in flight. Loudness is measured over
the sample itself, because a whole-episode pass would make a 15s preview
as slow as an export. Only the latest preview per project is kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The new Audio tab in the editor sidebar has:
- "Make podcast-ready", which applies -16 LUFS, light denoise, speaker
  leveling and the rumble filter in one click.
- Individual controls for each setting.
- "Preview 15s from playhead", which plays the before and after samples
  side by side.

Changing a setting clears the previous preview, since it no longer
matches what will export.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast audio cleanup: loudness targets, denoise, speaker leveling, before/after preview
PublishingMeta holds a project's chapters and show notes, both
AI-generated and user-editable. Chapters are stored in source time, so
cuts made after generation don't strand them on the wrong moment; they
are mapped onto the edited timeline whenever they're shown or exported.

RenderJob.format ("mp4" | "mp3" | "wav") makes room for audio-only
exports. Podcast feeds are audio, not video.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pure functions, unit-tested:
- buildEditedSentences: the transcript as it will sound after cuts, with
  both source and edited timing and speaker labels.
- chaptersFromPicks: turns the model's sentence picks into chapters.
  Starts come from the sentences themselves, never from a timestamp the
  model wrote. The first chapter is forced to the start, and picks closer
  than minChapterSeconds are merged. That minimum scales with episode
  length because, on a transcript that mentions a new topic every 20s,
  unconstrained picks put three chapters in the first 27s of an
  18-minute episode.
- resolveChapters: places stored chapters on the edited timeline. A
  chapter whose moment was cut moves to the next surviving moment.
- toYoutubeChapters / toFfmetadata: "00:00 Title" text, and the ffmpeg
  metadata file used to embed chapters in exports. Values are escaped,
  so a title can't inject keys or [CHAPTER] sections.
- formatTranscript: TXT/Markdown transcript grouped by speaker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- POST /publishing/chapters and POST /publishing/show-notes call
  gpt-4o-mini with JSON output. The model returns only text and
  sentence indices, never timestamps or edit operations, and every
  response is Zod-validated before it is stored (claude.md 22-23).
- The chapter prompt states the episode length and a minimum chapter
  length, so chapters span the episode instead of clustering at the
  start.
- The show-notes prompt forbids turning a mention of a topic into
  "tips" or "insights". Without that rule, a transcript that only named
  topics came back with invented advice.
- PUT endpoints save user edits. GET /transcript/export downloads the
  edited episode as TXT or Markdown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Audio-only exports reuse the video export's cut graph and audio
  cleanup chain, and skip captions, logo and color.
- MP3 is 192k at 44.1kHz with ID3v2.3, the version podcast apps read
  most reliably. The sample rate is fixed because at a 22.05kHz source
  rate MP3 can't reach 192k; that test source came out at 160k.
- Every MP4 and MP3 export now carries the episode title and chapter
  markers from a generated FFMETADATA input. Checked with ffprobe:
  chapters land at their edited-timeline times, and a title containing
  = ; # survives intact. WAV can't hold chapters, so it skips them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Publish tab has:
- Chapters: generate, rename, remove, add one at the playhead, click a
  time to jump there, and copy "00:00 Title" lines for YouTube.
- Show notes: summary, key points and title ideas, all editable, plus
  "Copy all" as Markdown with the chapters appended.
- TXT/Markdown transcript downloads.

Chapter times are re-fetched whenever the edit changes, since the server
places them on the edited timeline. The Export button now has a format
picker: MP4 video, MP3 audio or WAV audio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast publishing: AI chapters and show notes, MP3/WAV export, transcript export
Speaker detection needs the same speech track and silence-aligned chunks
as transcription, and projects transcribed before the track existed must
be able to create it. extractSpeechAudio, planSpeechChunks and
mapWithConcurrency move out of runTranscriptionJob's body into exported
helpers. No behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Speaker detection is a second kind of long transcript job. It reuses the
same row type for status and chunk progress, and the same startup cleanup
for runs interrupted by a restart. The transcribe status endpoint now only
reports "transcribe" jobs, and its 409 message covers either kind, since
the two must never rewrite the transcript at the same time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gpt-4o-transcribe-diarize returns speaker-labeled spans but no word
timings, so Whisper stays the source of truth for words and diarization
only decides who said each word. The model is behind lib/ai/diarize.ts
so the provider can be swapped.

lib/ai/speakers.ts (pure, unit-tested):
- assignSpeakers labels each word by span overlap, falling back to the
  nearest span in gaps. It splits segments where the speaker changes,
  keeping punctuation by splitting text by word position (the same rule
  sentences.ts uses) and keeping word ids, and absorbs one-word boundary
  blips. Speakers become "Speaker 1..." in order of first appearance.
- mergeDiarizedChunks makes labels consistent across chunks, which a
  test run showed they are not: a guest labeled "B" in one chunk came
  back as "C" in the next. Labels matching the first chunk's reference
  voices are kept. Any other label is matched to the speaker it
  co-occurs with in a stretch the two chunks share. Relying on
  references alone failed: when the first chunk merged two similar
  voices, the reference clip was the wrong one, and the host came back
  as a new speaker for the rest of the episode.
- pickReferenceSpans picks one 2-8s clip per speaker, up to the API's
  limit of 4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- POST /speakers/detect runs a background job, because diarization runs
  at roughly half real time (108s of audio took 50s). The first chunk is
  diarized alone and gives the reference clips. Later chunks run in
  parallel with those clips, each re-covering the last 45s of the
  previous chunk so labels can be matched across the boundary.
- The transcript is re-read just before writing, so edits made during
  the slow run are labeled rather than lost.
- POST /speakers/rename renames or clears a speaker on every segment at
  once.

Measured on a 14-minute, 3-voice synthetic episode spanning two chunks:
99.9% of words got the right label, 100% in every loop across the chunk
boundary. The one miss: two similar female voices merged into one
speaker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Above the transcript: "Detect speakers" (with chunk progress), and one
colored chip per speaker. Click a chip's name to rename that speaker
everywhere. The scissors cut everything that speaker said as ordinary cut
operations, so render, captions and chapters all follow; the restore
button removes those cuts again. Speaker bubbles in the transcript use
the same colors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Speaker detection: detect, rename, and cut speakers
A clip is a source-time range (like chapters, so later edits don't move
it) with an output aspect (9:16, 1:1, 16:9), a horizontal crop position
and a captions toggle. RenderJob.clipId marks a render of just that clip.
The relation uses SetNull, so deleting a clip keeps its finished renders.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
libass scales PlayResX and PlayResY independently, so the hardcoded
1920x1080 script would stretch every glyph on a 9:16 clip. toAssKaraoke
now takes the output frame, plus optional margin and outline overrides.
Defaults are unchanged, so normal exports render exactly as before.

splitCuesForShorts re-cuts sentence cues into even 3-4 word groups that
hold until the next group starts. At Shorts font sizes on a narrow
frame, a whole sentence wraps into a wall of text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
buildRenderArgs takes an optional reframe: crop to the target aspect,
then scale to the exact output size, before captions and the logo so
both are laid out on the new frame. When the source is wider than the
target, the crop takes the full height and slides horizontally by cropX;
when it's taller, it takes the full width, centered. Sizes are validated
integers, cropX is clamped, and the if() expressions are quoted so their
commas stay inside the filter option.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- clipAsCuts expresses a clip as two extra cuts, everything before and
  after it. Added to the project's own cuts, the existing render,
  caption and loudness pipeline produces just the clip, with the
  project's edits still applied inside it.
- highlightsFromPicks turns the model's sentence-index ranges into
  clips. Bounds come from word timings with a little breathing room,
  never from times the model wrote. Picks over 90s are shortened a
  sentence at a time, under 15s are dropped, and overlaps and bad
  indices are dropped. At most 8, in timeline order.

EditedSentence gains sourceEnd, which clips need.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
amide-init and others added 13 commits September 25, 2026 02:43
- POST /clips/highlights asks gpt-5.6-luna for the episode's best
  standalone moments. Judging what works as a clip across a whole
  episode is the "selecting important sections" work claude.md
  section 4 routes to Luna rather than 4o-mini. New AI picks replace
  earlier ones; hand-made clips are kept.
- Clips can be created, updated and deleted. POST /render with a
  clipId renders the clip: the project plus the clip's two extra cuts,
  reframed to the clip's aspect, with 3-4 word word-highlight captions
  in the project's font and colors.
- Captions are sized for phones: 5.5% of height, a heavier outline,
  and raised 18% clear of the app UI. On a first 9:16 render the
  normal 2px outline was too thin to read over busy footage.
- A clip render is 1080x1920 at 9:16. The episode's chapters and title
  are left out, and the download is named after the clip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Clips tab has:
- "Find highlights" (AI picks show their reason) and "Clip from
  transcript selection".
- Per clip: name, aspect, crop slider, captions toggle, play from
  start, render, and download.

The focused clip's crop is outlined on the player, with the cut-away
area dimmed, using the same geometry as the export. The overlay sizes
itself to the video's on-screen rectangle with container query units,
because the player box letterboxes the video. Measured: positioning
against the player box put the frame about 20% too wide and shifted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clips / Shorts: AI highlights, 9:16 reframing, Shorts captions
A word was assigned to a segment only if it fit entirely inside that
segment's time range. Whisper computes word and segment timestamps
separately, and they routinely disagree by a fraction of a second at the
edges: in a real response, "Thanks" was timed 6.40-6.70s while its
segment starts at 6.68s. Such a word matched no segment and silently
disappeared from the editable transcript, captions, exports and AI
features. It also shifted sentence boundaries in the segment, because
sentences.ts maps text to words by count.

Every word now goes to exactly one segment: the one it overlaps most,
or the nearest one for a word in a gap. Assignments never go backwards,
so order is preserved, and a segment's bounds grow to cover its words.

On a 14-minute, 3-speaker conversation, missing words went from 22
(0.9%, 22 of 222 segments) to 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fix words silently dropped at Whisper segment boundaries
The packaged Mac app copies its pre-migrated app.db.template only on
first launch, and the Prisma CLI isn't in the bundle, so an existing
install never received new tables. Its database stopped at the first 6
migrations, and the updated server would crash at startup on the
TranscriptionJob cleanup query.

On startup the server now applies any migration not yet recorded in
Prisma's own _prisma_migrations table, and records it in Prisma's format
(sha256 checksum, millisecond timestamps). It is a no-op for databases
that are already current, including dev databases managed with
`prisma migrate dev`.

Tested on a copy of a real older app database: the 5 missing migrations
applied, a second run did nothing, existing projects were kept, the
schema is identical to one built by `prisma migrate deploy`, and
`prisma migrate status` reports it up to date.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Apps launched from Finder don't get Homebrew on their PATH, so a bare
"ffprobe" isn't found. The app already passes an absolute FFMPEG_PATH
for the same reason. Since the long-episode work, every transcription
probes the audio's duration, so without this, transcription failed in
the packaged app. The editing proxy and logo/caption sizing probe media
too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mac app: migrate existing databases on startup, pass FFPROBE_PATH
A new page walks through the podcast features in the order a podcaster
uses them: long uploads, transcript editing, speaker labels, the audio
preset, publishing (chapters, show notes, transcript, MP3/WAV) and
Shorts. It includes measured figures (an 18-minute episode transcribes in
about a minute; speaker detection was 99.9% accurate on a synthetic
3-voice test) and the model behind each AI button. It doesn't quote
costs that weren't measured; only Whisper's published per-minute price
is given.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Home and Status: the podcast features, and an updated AI card.
- AI-Assisted Editing: the new AI actions, a model table, and the rule
  that models return sentence picks rather than timestamps, with
  everything validated before use.
- Architecture: corrected data layout (render/, not renders/; the proxy
  and audio track now actually exist), background jobs, and database
  upgrades on startup.
- Getting Started: ffprobe requirement, MAX_UPLOAD_MB, and a tip about
  the bun --watch EBADF error in long dev sessions.
- Download: updating keeps your projects, and troubleshooting for the
  404 caused by a dev server holding port 3001.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The intro now describes the podcast-first product, the status list
covers long episodes, speakers, audio, publishing and Shorts, and there
is a note on how AI output is constrained. Setup mentions ffprobe, the
port 3001 clash between the dev server and the Mac app, and the
bun --watch EBADF error. It also notes that installing a newer app keeps
existing projects.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Docs: podcast features, Podcast Workflow guide, README refresh
@amide-init
amide-init merged commit 815be05 into main Sep 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant