Repository navigation
Release: podcaster features (long episodes, audio, publishing, speakers, Shorts) - #38
Merged
Merged
Conversation
Long podcast episodes take minutes to transcribe, too long to hold an API request open (claude.md section 14). A job row mirroring RenderJob lets the work run in the background while the client polls status and per-chunk progress. The finished transcript still lives in Transcript; this row only tracks the run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…proxy Deterministic argv builders for the long-episode pipeline: extract a mono 16kHz 32kbps speech track (~14MB/hour, well under Whisper's 25MB cap per chunk), detect silences to choose chunk boundaries, slice chunks with an exact re-encoded seek so merged timestamps stay precise, and build a 720p editing proxy (claude.md section 16). runFfmpegCapturingStderr exists because silencedetect reports its results only as log lines, and probeDuration is needed to plan chunks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast episodes run 30-120 minutes and several GB, which broke both the 500MB upload cap and Whisper's 25MB file limit, and held the transcribe request open for the whole run. Transcription is now a background job (lib/ai/transcription-job.ts): extract a small speech track, split it into ~10-minute chunks at detected silences so no word is cut in half, transcribe 3 chunks at a time, then shift and merge them onto the source timeline. The merge also drops Whisper's end-of-clip repeats: a copy of the previous sentence made of zero-length words, which chunking would otherwise produce at every chunk end. Genuine repeated phrases with real timing are kept. POST /transcribe now returns 202 + jobId, and the upload screen polls it and shows chunk progress. Uploads use XHR so they can report progress, and the upload cap is 10GB by default, configurable with MAX_UPLOAD_MB. Jobs still running at startup were interrupted by a restart, so they are marked failed; otherwise they would block retries forever. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Multi-GB 1080p/4K podcast originals stutter when the browser plays them, and the timeline waveform fetched and decoded the entire video. For a 2-hour episode that is over 1GB of decoded audio at 48kHz. After upload, a 720p proxy is built in the background when the source is above 720p or over 300MB (claude.md section 16). It is served at a separate ?variant=proxy URL rather than swapped in behind /video, so a proxy that finishes mid-session can't change the bytes under an in-flight Range request. Proxy creation is best-effort: on failure the editor plays the original. Export always uses the original. The waveform now decodes the small extracted speech track (GET /audio) at 8kHz, and falls back to the video for projects transcribed before the track existed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Support long podcast episodes: chunked background transcription, 720p proxy, 10GB uploads
Adds AudioSettings: a loudness target (off / -16 / -14 LUFS), noise reduction (off / light / strong), speaker leveling, and an 80Hz rumble filter. Stored as Project.audioSettingsJson and validated with Zod on PATCH. Every field is an enum or boolean, never a raw number, so nothing the client sends can reach an ffmpeg filter string. A stored value that no longer validates falls back to all-off. Defaults are all off, so existing projects export exactly as before. The migration also adds RenderJob.kind, which separates full exports from the short audio previews added later in this series. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cleanup chain runs after the cuts in this order: high-pass, afftdn denoise, dynaudnorm leveling, then loudness. Each stage maps to a fixed filter string (lib/ffmpeg/audio-filters.ts). Loudness uses measured gain plus a peak limiter, not loudnorm's two-pass linear mode. Speech is peaky, so the gain needed usually pushes peaks past the ceiling, and loudnorm then silently falls back to dynamic mode and undershoots: a -14 target measured -16.3 LUFS on a test clip. resolveLoudnessGain measures the edited audio, applies the gain through the real chain, and corrects with secant steps, since the limiter lowers loudness more as it works harder. Corrections may add at most 6dB over the first estimate; beyond that, AAC encoding pushed peaks to clipping (+0.7 dBTP) on a very quiet source. The ceiling is -2 dBTP rather than -1.5 because AAC encoding overshoots the limiter; measured, -1.5 came out at -1.1 dBTP. Measured with ebur128 on real renders: - -16 target: -15.9 to -16.2 LUFS - -14 target: -14.1 LUFS The audio edit graph (atrim + acrossfade) is now shared by the export, the loudness measurement and an audio-only render, so all three process exactly the same audio. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The browser can't run ffmpeg's audio filters live, so there was no way to hear the cleanup without a full export. POST /render/audio-preview renders the same 15s of the edited program twice, once untouched and once through the cleanup chain, starting from the playhead. Cuts inside that window are respected. The request carries the panel's current settings rather than reading them from the project row, so the preview matches what's on screen even while the settings PATCH is still in flight. Loudness is measured over the sample itself, because a whole-episode pass would make a 15s preview as slow as an export. Only the latest preview per project is kept. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The new Audio tab in the editor sidebar has: - "Make podcast-ready", which applies -16 LUFS, light denoise, speaker leveling and the rumble filter in one click. - Individual controls for each setting. - "Preview 15s from playhead", which plays the before and after samples side by side. Changing a setting clears the previous preview, since it no longer matches what will export. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast audio cleanup: loudness targets, denoise, speaker leveling, before/after preview
PublishingMeta holds a project's chapters and show notes, both
AI-generated and user-editable. Chapters are stored in source time, so
cuts made after generation don't strand them on the wrong moment; they
are mapped onto the edited timeline whenever they're shown or exported.
RenderJob.format ("mp4" | "mp3" | "wav") makes room for audio-only
exports. Podcast feeds are audio, not video.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pure functions, unit-tested: - buildEditedSentences: the transcript as it will sound after cuts, with both source and edited timing and speaker labels. - chaptersFromPicks: turns the model's sentence picks into chapters. Starts come from the sentences themselves, never from a timestamp the model wrote. The first chapter is forced to the start, and picks closer than minChapterSeconds are merged. That minimum scales with episode length because, on a transcript that mentions a new topic every 20s, unconstrained picks put three chapters in the first 27s of an 18-minute episode. - resolveChapters: places stored chapters on the edited timeline. A chapter whose moment was cut moves to the next surviving moment. - toYoutubeChapters / toFfmetadata: "00:00 Title" text, and the ffmpeg metadata file used to embed chapters in exports. Values are escaped, so a title can't inject keys or [CHAPTER] sections. - formatTranscript: TXT/Markdown transcript grouped by speaker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- POST /publishing/chapters and POST /publishing/show-notes call gpt-4o-mini with JSON output. The model returns only text and sentence indices, never timestamps or edit operations, and every response is Zod-validated before it is stored (claude.md 22-23). - The chapter prompt states the episode length and a minimum chapter length, so chapters span the episode instead of clustering at the start. - The show-notes prompt forbids turning a mention of a topic into "tips" or "insights". Without that rule, a transcript that only named topics came back with invented advice. - PUT endpoints save user edits. GET /transcript/export downloads the edited episode as TXT or Markdown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Audio-only exports reuse the video export's cut graph and audio cleanup chain, and skip captions, logo and color. - MP3 is 192k at 44.1kHz with ID3v2.3, the version podcast apps read most reliably. The sample rate is fixed because at a 22.05kHz source rate MP3 can't reach 192k; that test source came out at 160k. - Every MP4 and MP3 export now carries the episode title and chapter markers from a generated FFMETADATA input. Checked with ffprobe: chapters land at their edited-timeline times, and a title containing = ; # survives intact. WAV can't hold chapters, so it skips them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Publish tab has: - Chapters: generate, rename, remove, add one at the playhead, click a time to jump there, and copy "00:00 Title" lines for YouTube. - Show notes: summary, key points and title ideas, all editable, plus "Copy all" as Markdown with the chapters appended. - TXT/Markdown transcript downloads. Chapter times are re-fetched whenever the edit changes, since the server places them on the edited timeline. The Export button now has a format picker: MP4 video, MP3 audio or WAV audio. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Podcast publishing: AI chapters and show notes, MP3/WAV export, transcript export
Speaker detection needs the same speech track and silence-aligned chunks as transcription, and projects transcribed before the track existed must be able to create it. extractSpeechAudio, planSpeechChunks and mapWithConcurrency move out of runTranscriptionJob's body into exported helpers. No behavior change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Speaker detection is a second kind of long transcript job. It reuses the same row type for status and chunk progress, and the same startup cleanup for runs interrupted by a restart. The transcribe status endpoint now only reports "transcribe" jobs, and its 409 message covers either kind, since the two must never rewrite the transcript at the same time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gpt-4o-transcribe-diarize returns speaker-labeled spans but no word timings, so Whisper stays the source of truth for words and diarization only decides who said each word. The model is behind lib/ai/diarize.ts so the provider can be swapped. lib/ai/speakers.ts (pure, unit-tested): - assignSpeakers labels each word by span overlap, falling back to the nearest span in gaps. It splits segments where the speaker changes, keeping punctuation by splitting text by word position (the same rule sentences.ts uses) and keeping word ids, and absorbs one-word boundary blips. Speakers become "Speaker 1..." in order of first appearance. - mergeDiarizedChunks makes labels consistent across chunks, which a test run showed they are not: a guest labeled "B" in one chunk came back as "C" in the next. Labels matching the first chunk's reference voices are kept. Any other label is matched to the speaker it co-occurs with in a stretch the two chunks share. Relying on references alone failed: when the first chunk merged two similar voices, the reference clip was the wrong one, and the host came back as a new speaker for the rest of the episode. - pickReferenceSpans picks one 2-8s clip per speaker, up to the API's limit of 4. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- POST /speakers/detect runs a background job, because diarization runs at roughly half real time (108s of audio took 50s). The first chunk is diarized alone and gives the reference clips. Later chunks run in parallel with those clips, each re-covering the last 45s of the previous chunk so labels can be matched across the boundary. - The transcript is re-read just before writing, so edits made during the slow run are labeled rather than lost. - POST /speakers/rename renames or clears a speaker on every segment at once. Measured on a 14-minute, 3-voice synthetic episode spanning two chunks: 99.9% of words got the right label, 100% in every loop across the chunk boundary. The one miss: two similar female voices merged into one speaker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Above the transcript: "Detect speakers" (with chunk progress), and one colored chip per speaker. Click a chip's name to rename that speaker everywhere. The scissors cut everything that speaker said as ordinary cut operations, so render, captions and chapters all follow; the restore button removes those cuts again. Speaker bubbles in the transcript use the same colors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Speaker detection: detect, rename, and cut speakers
A clip is a source-time range (like chapters, so later edits don't move it) with an output aspect (9:16, 1:1, 16:9), a horizontal crop position and a captions toggle. RenderJob.clipId marks a render of just that clip. The relation uses SetNull, so deleting a clip keeps its finished renders. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
libass scales PlayResX and PlayResY independently, so the hardcoded 1920x1080 script would stretch every glyph on a 9:16 clip. toAssKaraoke now takes the output frame, plus optional margin and outline overrides. Defaults are unchanged, so normal exports render exactly as before. splitCuesForShorts re-cuts sentence cues into even 3-4 word groups that hold until the next group starts. At Shorts font sizes on a narrow frame, a whole sentence wraps into a wall of text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
buildRenderArgs takes an optional reframe: crop to the target aspect, then scale to the exact output size, before captions and the logo so both are laid out on the new frame. When the source is wider than the target, the crop takes the full height and slides horizontally by cropX; when it's taller, it takes the full width, centered. Sizes are validated integers, cropX is clamped, and the if() expressions are quoted so their commas stay inside the filter option. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- clipAsCuts expresses a clip as two extra cuts, everything before and after it. Added to the project's own cuts, the existing render, caption and loudness pipeline produces just the clip, with the project's edits still applied inside it. - highlightsFromPicks turns the model's sentence-index ranges into clips. Bounds come from word timings with a little breathing room, never from times the model wrote. Picks over 90s are shortened a sentence at a time, under 15s are dropped, and overlaps and bad indices are dropped. At most 8, in timeline order. EditedSentence gains sourceEnd, which clips need. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- POST /clips/highlights asks gpt-5.6-luna for the episode's best standalone moments. Judging what works as a clip across a whole episode is the "selecting important sections" work claude.md section 4 routes to Luna rather than 4o-mini. New AI picks replace earlier ones; hand-made clips are kept. - Clips can be created, updated and deleted. POST /render with a clipId renders the clip: the project plus the clip's two extra cuts, reframed to the clip's aspect, with 3-4 word word-highlight captions in the project's font and colors. - Captions are sized for phones: 5.5% of height, a heavier outline, and raised 18% clear of the app UI. On a first 9:16 render the normal 2px outline was too thin to read over busy footage. - A clip render is 1080x1920 at 9:16. The episode's chapters and title are left out, and the download is named after the clip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Clips tab has: - "Find highlights" (AI picks show their reason) and "Clip from transcript selection". - Per clip: name, aspect, crop slider, captions toggle, play from start, render, and download. The focused clip's crop is outlined on the player, with the cut-away area dimmed, using the same geometry as the export. The overlay sizes itself to the video's on-screen rectangle with container query units, because the player box letterboxes the video. Measured: positioning against the player box put the frame about 20% too wide and shifted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clips / Shorts: AI highlights, 9:16 reframing, Shorts captions
A word was assigned to a segment only if it fit entirely inside that segment's time range. Whisper computes word and segment timestamps separately, and they routinely disagree by a fraction of a second at the edges: in a real response, "Thanks" was timed 6.40-6.70s while its segment starts at 6.68s. Such a word matched no segment and silently disappeared from the editable transcript, captions, exports and AI features. It also shifted sentence boundaries in the segment, because sentences.ts maps text to words by count. Every word now goes to exactly one segment: the one it overlaps most, or the nearest one for a word in a gap. Assignments never go backwards, so order is preserved, and a segment's bounds grow to cover its words. On a 14-minute, 3-speaker conversation, missing words went from 22 (0.9%, 22 of 222 segments) to 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fix words silently dropped at Whisper segment boundaries
The packaged Mac app copies its pre-migrated app.db.template only on first launch, and the Prisma CLI isn't in the bundle, so an existing install never received new tables. Its database stopped at the first 6 migrations, and the updated server would crash at startup on the TranscriptionJob cleanup query. On startup the server now applies any migration not yet recorded in Prisma's own _prisma_migrations table, and records it in Prisma's format (sha256 checksum, millisecond timestamps). It is a no-op for databases that are already current, including dev databases managed with `prisma migrate dev`. Tested on a copy of a real older app database: the 5 missing migrations applied, a second run did nothing, existing projects were kept, the schema is identical to one built by `prisma migrate deploy`, and `prisma migrate status` reports it up to date. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Apps launched from Finder don't get Homebrew on their PATH, so a bare "ffprobe" isn't found. The app already passes an absolute FFMPEG_PATH for the same reason. Since the long-episode work, every transcription probes the audio's duration, so without this, transcription failed in the packaged app. The editing proxy and logo/caption sizing probe media too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mac app: migrate existing databases on startup, pass FFPROBE_PATH
A new page walks through the podcast features in the order a podcaster uses them: long uploads, transcript editing, speaker labels, the audio preset, publishing (chapters, show notes, transcript, MP3/WAV) and Shorts. It includes measured figures (an 18-minute episode transcribes in about a minute; speaker detection was 99.9% accurate on a synthetic 3-voice test) and the model behind each AI button. It doesn't quote costs that weren't measured; only Whisper's published per-minute price is given. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Home and Status: the podcast features, and an updated AI card. - AI-Assisted Editing: the new AI actions, a model table, and the rule that models return sentence picks rather than timestamps, with everything validated before use. - Architecture: corrected data layout (render/, not renders/; the proxy and audio track now actually exist), background jobs, and database upgrades on startup. - Getting Started: ffprobe requirement, MAX_UPLOAD_MB, and a tip about the bun --watch EBADF error in long dev sessions. - Download: updating keeps your projects, and troubleshooting for the 404 caused by a dev server holding port 3001. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The intro now describes the podcast-first product, the status list covers long episodes, speakers, audio, publishing and Shorts, and there is a note on how AI output is constrained. Setup mentions ffprobe, the port 3001 clash between the dev server and the Mac app, and the bun --watch EBADF error. It also notes that installing a newer app keeps existing projects. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Docs: podcast features, Podcast Workflow guide, README refresh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges everything on
devintomain: 43 commits from PRs #30–#37, each already reviewed and merged intodev.What's in it
gpt-4o-transcribe-diarize) kept consistent across chunks, rename everywhere, cut/restore everything one person saidgpt-5.6-luna), clips from a transcript selection, 9:16/1:1/16:9 reframing with an on-player crop preview, Shorts captionsFFPROBE_PATHEvery AI feature is an explicit button. Models return text and sentence picks, never timestamps or edits, and every response is validated before use.
Pre-release security check
Run over all 43 commits (full history, not just the final diff, so a key added and later removed would still be caught):
sk-keys, AWS/GitHub/Slack/Google keys, private keys, or hardcoded key/token/password values in any commit. No commit adds or removes ansk-string..env, database,settings.json,data/, media, build output, generated code or binaries. The only env-related change is a comment forMAX_UPLOAD_MBin.env.example..gitignorestill covers.env, databases, settings,server-bundle/, Tauritarget/and the generated Prisma client.~paths. The commit author email is the same one already on all 160 commits onmain.DATA_DIR.Known low-severity note (not a blocker): raising the upload limit to 10 GB also raised Bun's request-size limit for JSON endpoints, so an oversized JSON request could use a lot of memory. There's no auth and the app is local-only by design (claude.md §18), so anyone who can reach the server can already do more. A per-route body limit is a reasonable follow-up.
Verified on
devjust nowAfter merging
v*.*.*tag for CI to publish it. Existing installs upgrade their database automatically on first launch.main.🤖 Generated with Claude Code