Skip to content

Release: scenes, title cards, transitions, B-roll and padding - #46

Merged
amide-init merged 43 commits into
mainfrom
dev
Sep 27, 2026
Merged

amide-init merged 43 commits into
mainfrom
dev

Conversation

@amide-init

Copy link
Copy Markdown
Owner

Brings main up to date with dev for v0.3.0.

Included

Once this is merged, tagging v0.3.0 on main builds the macOS app and attaches it to the release.

🤖 Generated with Claude Code

amide-init and others added 30 commits September 26, 2026 00:44
Split ops existed in the type union but nothing used them. They become
scene boundaries: an optional title names the scene, source records
whether a person, the AI or shot detection placed it, and groupId lets a
batch of ops (e.g. applied suggestions) undo as one step. The route
rejects a split at or past the end of the video, since that would make an
empty scene; a split at 0 is allowed and only names the first scene.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Undo used to pop the last operation, so removing 40 filler words took 40
undos, and restoring a cut or renaming anything couldn't be undone at
all. History is now a stack of steps, each recording the ops it added and
removed: bulk actions (filler words, long pauses, cutting a speaker)
commit one step, restores are undoable, and replacing an op -- how a
scene gets renamed, since ops are never mutated -- swaps back on undo.
After a reload the steps are rebuilt from the persisted ops, grouping by
groupId.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scenes are computed, never stored: the ranges between split points, with
their position on the edited timeline so a fully cut scene collapses to
zero width but stays listed for restoring. New splits snap to the middle
of the nearest gap between words, so a boundary -- and later a card or
transition placed on it -- never lands inside a word.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Split at the playhead (S or the timeline button) or before any sentence
in the transcript. Scenes show as named blocks above the timeline and as
markers inline in the transcript; the Scenes panel renames, jumps to,
cuts, restores and merges them. Cmd/Ctrl+Z and Shift+Cmd/Ctrl+Z now
undo and redo from anywhere outside a text field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scenes: split the episode into named scenes
Cuts are joined with a chain of 30ms xfades. On ffmpeg 9 (9.0.2, the
current Homebrew ffmpeg-full the Mac app uses), xfade discards the second
input's frames before its offset instead of holding them, so every
segment after a cut started late: a 2s + 10s join rendered 10.0s long,
with the second segment's first ~2s missing, and the loss compounded
with each cut. Audio (acrossfade) was unaffected, so exports also drifted
out of sync.

Pinning each trimmed segment to the source's frame rate with fps= makes
the join land where the offset says: the same join renders 11.98s, and a
three-segment export's frames match the source at every segment start.
The render now always probes the source's rate (probeVideoStream), which
also provides the size the logo and caption layout needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each trimmed segment is a whole number of frames, so it can come out up
to a frame shorter than the length the xfade offsets assume. When the
stream joined so far runs out before its next join's transition ends,
ffmpeg 9 ends that join there and every segment after it is lost --
measured: a 5s + 10s + 2s chain rendered 14.9s with the last segment
missing.

Every segment but the last now gets 0.1s of cloned frames (tpad). xfade
drops the first input's frames after the transition, so the padding
never reaches the output: a four-segment export now renders 15.88s video
against 15.91s audio (within a frame), and the same three-segment chain
renders all 16.95s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A card is a full-screen title card inserted into the program before the
content at its source time: a scene start, 0 for an intro, the duration
for an outro. It carries a template (title, chapter, quote, outro), a
title and optional subtitle, 1-10s on screen and a solid background.

Card text is burned in through a generated ASS file, so the schema
refuses braces, backslashes and control characters -- they'd be read as
override tags or break the event line. The route rejects a card placed
past the end of the video.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cards add time that isn't in the source, so edited time (source minus
cuts) is no longer what plays. program.ts places each card on the edited
timeline and maps edited time to program time (edited plus cards), so
everything that already works in edited time -- captions, chapters --
shifts into program time without knowing about cards, and cuts and cards
compose. A card whose whole scene is cut goes with it, rather than
announcing the next scene; intro and outro cards always stay.

With no cards the program is exactly the playable ranges, so nothing
changes for existing projects. Duplicated across client and server like
cuts.ts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
layout.ts turns a card into positioned lines as percentages of the frame
(shared with the client preview, so both place the same card the same
way at any size), picks dark or light text from the background's
luminance, and fades the text in and out. ass.ts writes every card's
text into one ASS file timed on the program timeline, with the output
size as its script resolution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each card becomes an ffmpeg color source at the source's size and frame
rate, joined into the same xfade chain as the cuts, with silence for its
duration on the audio side; the cards' text is burned in over the whole
program in one subtitles pass, before captions and the logo. With cards
present, segments are normalized to one pixel and audio format so they
can join, and color grading moves onto each footage segment so a preset
doesn't recolor the cards. Projects without cards render the same graph
as before.

Captions and embedded chapters move onto the program timeline; a chapter
that starts where a card plays starts with the card. MP3/WAV exports get
the cards' silence so chapters stay in sync. Clips skip cards.

Verified with ffmpeg 9.0.2 on a synthetic clip with an intro, a cut, a
chapter card and an outro: every card and segment lands where the program
model says (video 23.857s, audio 23.850s), and a beep every second lands
at its predicted time before and after each card.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Publish panel's chapters, YouTube chapter list and the exported
transcript's timestamps now include title cards, so they match what the
export actually plays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The preview can't insert frames into a <video>, so useCardPlayback pauses
the video where a card plays, shows CardPreview on a clock of its own
(same layout percentages as the export, in cqh units), then resumes;
cards at the same moment play back to back, and outros play once the
last kept range ends. The scrubber and timecode run on program time.
Paused exactly where a card plays, the playhead sits at the card's start,
since the card is still ahead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The track is laid out on program time: footage with each card as a block
in its own color between it. Clicking a card opens it in the player;
scene blocks start at their card, and pause markers and scene boundaries
are placed around cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each scene gets an intro/title card button, and there's an Outro
section. The editor offers the four templates, title and subtitle, a 1-
10s duration and preset or custom colors; changes preview live over the
player and save as one undo step when you close it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Records the three clocks (source, edited, program), how cards render and
preview, and the frame-rate and padding rules the render graph needs on
ffmpeg 9, so nobody removes them as noise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fix: exports drop footage after every cut on ffmpeg 9
Title cards: intro, chapter and outro cards between scenes
A transition styles how the scene starting at its time joins what plays
before it: dip to black, dip to white, or crossfade with the scene's
title card. At 0 it fades the episode in; at the end, out. 0.2-2s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Transitions are placed like cards: one per moment (the latest wins),
following the scene's content if its start was cut, and dropped with the
scene if it's cut entirely. Footage is split at a transition inside a
kept range so it has a join to live on, and every join at the moment --
into a card and out of it -- gets the style. A crossfade needs a card on
one side; between two stretches of footage that follow each other in the
source it would show nothing, so it falls back to a cut.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Captions and chapters are placed on program time, so a transition that
ate time would knock everything after it out of sync. A dip instead fades
each side's own edge to the color (audio fades too, capped at 0.3s since
speech sits right by scene boundaries) and keeps the usual 30ms join; a
crossfade's overlap is added to the card it blends with, so the footage
around it keeps its timing. The video and audio graphs now share one
per-item layout (render length, join overlap, edge fades) so they can't
drift apart. Projects with no cards or transitions render the same graph.

Verified with ffmpeg 9.0.2: a synthetic clip with a fade in, a white dip,
a crossfaded card and a fade out renders 22.99s (20s + 3s card, as
without transitions) and a beep every second lands at its predicted time
throughout, including right after the crossfaded card.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
transition-look.ts works out, for a program time, what color and/or
title card to lay over the player and how opaque, with the render's
timing and clamping. TransitionOverlay applies it from its own animation-
frame loop through refs, so fades are smooth without re-rendering the
editor every frame. Card and transition layers now sit under the logo,
as in the export.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each scene gets a Transition in picker (crossfade offered once it has a
title card), the first scene a Fade in and the outro a Fade out, each
with a length. On the timeline a diamond marks each transition and a
wedge each episode fade.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Transitions: dips, card crossfades, and episode fade in/out
Projects can now hold many images and video clips alongside their main
video. Uploads are raw bodies like the main upload, stored under a unique
prefix (camera files often share names), and probed on the way in --
anything ffmpeg can't read is refused now rather than at export. Asset
gains nullable name/width/height/duration columns (an additive migration,
applied at startup). A file can't be deleted while B-roll uses it, so an
export never references a missing file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An overlay shows a media-library file over the footage from start to end
in source time, so it follows cuts, either full screen or picture-in-
picture in a corner, optionally from an offset into a clip. The route
checks the file is in this project's library and the range is inside the
video; asset ids must be plain ids.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
B-roll maps through cuts (a partly cut range shortens it, a fully cut
one drops it) and past title cards like a caption does, starting after a
card at its start and ending before one at its end.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each file is an extra ffmpeg input: an image looped for its slot, a clip
trimmed from its offset that holds its last frame if it runs short. It's
scaled to fill the frame or to a 32%-wide inset in its corner, shifted to
its program start, and overlaid only for its slot -- after the join and
grade, before any clip reframe, card text, captions and the logo, whose
input indices move past the B-roll. File paths only appear as -i
arguments, never inside the filtergraph. B-roll audio is never used.

Verified with ffmpeg 9.0.2: a 2s clip over a 4s slot and a picture-in-
picture image, around a cut and an intro card, land at their predicted
program times; the clip holds its last frame, and audio sync is
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Upload images and clips (with progress), then add one as B-roll: over
the words selected in the transcript, or from the playhead. The panel
lists the B-roll in use, where full screen / picture-in-picture and the
corner can be changed; each change is one undo step. Deleting a file
that's in use shows why it can't be removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
amide-init and others added 13 commits September 26, 2026 02:00
BrollLayer shows B-roll for its slot, laid out like the export, from its
own animation-frame loop; a clip's position follows the program (offset
plus time into the slot, holding its last frame) and is re-seeked only
when it drifts. A B-roll lane under the timeline shows each piece.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
B-roll: a media library and image/clip overlays
Scene suggestions land better on a visual cut (a multicam switch, a
screen share starting), so ffmpeg scores frames at 320px wide and
reports every frame over the scene threshold; near-duplicates within a
second merge and the very ends are dropped. It runs as a 'shots'
TranscriptionJob that stores its times in a new resultJson column (an
additive migration), reading the 720p proxy when there is one. On a
synthetic clip with cuts at 5s and 10s it finds exactly those, none on
continuous footage, and scans a 14-minute file in under two seconds.

The transcribe and speaker routes refused to start while any job was
running; they now only wait for jobs that rewrite the transcript, so
shot detection doesn't block them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gpt-4o-mini reads the edited transcript and returns the sentences where
topics change, with titles -- sentence indices only, like chapters,
schema-validated. scenesFromPicks turns them into boundaries: whole
sentences, the first scene at 0 (a first pick near the start names it),
at least 45s apart, placed in the silence before the sentence, or on a
detected shot change within 1.5s that isn't inside a word. The route
only returns suggestions; nothing is applied server-side.

The prompt spells out the sentence range and asks for the last scene
late in the episode: without that, gpt-4o-mini put every scene in the
first 90 seconds of a 14-minute transcript.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Scenes panel gets Suggest scenes, From chapters (built client-side
from the Publish panel's chapters, no AI call) and shot-change
detection. Suggestions show as a review list -- keep or drop each, edit
titles, optionally add a numbered chapter card to each -- and as dashed
lines on the timeline, with detected cuts as ticks. Apply adds the kept
splits and cards as one undo step; a suggestion within a second of an
existing split renames it rather than leaving a sliver of a scene.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A new Scenes, Cards & B-roll guide page, a scenes step in the podcast
workflow, status and README entries, the grouped undo, and the scene
endpoints in claude.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scene suggestions: AI topics, chapters and shot-change detection, plus docs
A `pad` operation shrinks one scene's footage and centres it on a solid
border, keyed by the scene's start like transitions. The output size
never changes, so every segment still matches for the xfade chain; the
padded scene's edges become ordinary cut joins, so the program's length
doesn't change either. Padded projects are graded per segment (as with
cards) so the border keeps its exact color, and only a validated
'#RRGGBB' ever reaches the pad filter.

program.ts changes are mirrored into the client copy, which the preview
will use next.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each scene row in the Scenes panel gets a Padding size and a border
color, stored as one `pad` operation per scene, so every change is one
undo step. Merging a scene drops its padding with its boundary.

The preview shrinks the <video> and shows the border color behind it,
switched on the video's live time in an animation-frame loop (like
TransitionOverlay), so it changes exactly at the scene boundary. A
focused clip previews without padding, as clips export without it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Splitting by playhead means scrubbing to the right spot first; in a
transcript-first editor the natural place to mark a scene boundary is
the text itself. Clicking a word now also sets a cursor before it, and
pressing / starts a new scene there -- in the gap before the word, via
the same splitAt as every other split, so the minimum-scene-length guard
and undo work unchanged. With no cursor, / splits at the playhead like S.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Per-scene frame padding and / to split in the transcript
@amide-init
amide-init merged commit 32c17d4 into main Sep 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant