Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
00db19e
feat(server): add TranscriptionJob model for background transcription
amide-init Sep 24, 2026
fee7583
feat(server): add ffmpeg builders for audio extraction, chunking and …
amide-init Sep 24, 2026
4cde01d
feat: transcribe long episodes in background chunks
amide-init Sep 24, 2026
2b60a2c
feat: play a 720p proxy and draw the waveform from the speech track
amide-init Sep 24, 2026
463deb6
docs: document the long-episode transcription pipeline and new endpoints
amide-init Sep 24, 2026
7b8961e
Merge pull request #30 from amide-init/feat/long-episodes
amide-init Sep 24, 2026
05ee58e
feat(server): persist per-project podcast audio settings
amide-init Sep 24, 2026
8115d6c
feat(server): apply podcast audio cleanup and loudness targets at export
amide-init Sep 24, 2026
a482e9a
feat(server): render a 15s before/after audio preview
amide-init Sep 24, 2026
76dfbaa
feat(client): add an Audio panel with podcast preset and preview
amide-init Sep 24, 2026
0bd8a96
docs: document podcast audio cleanup and the audio-preview endpoint
amide-init Sep 24, 2026
e244784
Merge pull request #31 from amide-init/feat/audio-quality
amide-init Sep 24, 2026
797adec
feat(server): add PublishingMeta and an export format to render jobs
amide-init Sep 24, 2026
ab6fc2d
feat(server): chapter and edited-transcript logic for publishing
amide-init Sep 24, 2026
bbfae03
feat(server): generate chapters and show notes, export transcripts
amide-init Sep 24, 2026
fcbb0dc
feat(server): export MP3/WAV and embed chapters in every export
amide-init Sep 24, 2026
1e5d5f2
feat(client): add a Publish panel and MP3/WAV export
amide-init Sep 24, 2026
1f158e3
docs: document podcast publishing and the new endpoints
amide-init Sep 24, 2026
2794398
Merge pull request #32 from amide-init/feat/publishing
amide-init Sep 24, 2026
55c33a4
refactor(server): share speech-audio extraction and chunk planning
amide-init Sep 24, 2026
8e1714e
feat(server): add a kind to TranscriptionJob
amide-init Sep 24, 2026
d2223e5
feat(server): diarization service and speaker alignment
amide-init Sep 24, 2026
e06ad77
feat(server): detect speakers in the background and rename them in bulk
amide-init Sep 24, 2026
f7c0fa6
feat(client): speakers bar with detect, rename, and cut-all
amide-init Sep 24, 2026
0233ecb
docs: document speaker detection and its endpoints
amide-init Sep 24, 2026
84317d3
Merge pull request #33 from amide-init/feat/speakers
amide-init Sep 24, 2026
ce059a7
feat(server): add the Clip model for Shorts
amide-init Sep 24, 2026
1eb794f
feat(server): lay out karaoke captions for any frame; Shorts cues
amide-init Sep 24, 2026
0ebb861
feat(server): reframe renders to a clip's aspect ratio
amide-init Sep 24, 2026
038f0c6
feat(server): clip rules -- clip-as-cuts and highlight validation
amide-init Sep 24, 2026
0f50792
feat(server): find highlights and render clips as Shorts
amide-init Sep 24, 2026
dc2d422
feat(client): Clips panel with crop preview on the player
amide-init Sep 24, 2026
d248d99
docs: document clips / Shorts and the new endpoints
amide-init Sep 24, 2026
2ea7aa2
Merge pull request #34 from amide-init/feat/clips
amide-init Sep 24, 2026
4e176d4
fix(server): stop dropping words at Whisper segment boundaries
amide-init Sep 24, 2026
d3121ab
Merge pull request #35 from amide-init/fix/whisper-dropped-words
amide-init Sep 24, 2026
c57cd61
fix(server): apply pending migrations at startup
amide-init Sep 24, 2026
57b7cec
fix(tauri): pass FFPROBE_PATH to the bundled server
amide-init Sep 24, 2026
dcd12b3
Merge pull request #36 from amide-init/fix/mac-app-migrations
amide-init Sep 24, 2026
621b13f
docs: add a Podcast Workflow guide
amide-init Sep 24, 2026
b575c10
docs: bring the docs site up to date with the podcast features
amide-init Sep 24, 2026
7b486c3
docs(readme): describe the podcast features and dev gotchas
amide-init Sep 24, 2026
6408ec3
Merge pull request #37 from amide-init/docs/podcaster-features
amide-init Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 57 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,17 @@

![transcriptcut editor: video preview, transcript with word-level cuts, and a frame-thumbnail timeline](docs/public/screenshots/hero.png)

Edit video by editing its transcript, with a few explicit AI-assisted
actions (filler-word and long-pause removal) layered on top — **not** a
free-text "ask the AI to edit this" chat bar; an early version of that was
built and worked, then deliberately removed (see Status below and
`claude.md` section 3 for why). **Local-first and open source** — runs
entirely on your machine, no cloud account required. See the
[docs](https://amide-init.github.io/transcriptcut/) for more, and the
Edit video by editing its transcript, built for **video podcasts**:
long multi-speaker episodes in, podcast-ready audio, chapters, show notes,
an MP3 for your feed and vertical Shorts out. AI helps through explicit,
scoped buttons (filler words, speakers, chapters, show notes, highlights),
**not** a free-text "ask the AI to edit this" chat bar; an early version
of that was built and worked, then deliberately removed (see `claude.md`
section 3 for why). **Local-first and open source** — runs entirely on
your machine, no cloud account required. See the
[docs](https://amide-init.github.io/transcriptcut/), especially the
[Podcast Workflow](https://amide-init.github.io/transcriptcut/guide/podcasting)
guide, and the
[issue tracker](https://github.com/amide-init/transcriptcut/issues) for
current build status.

Expand All @@ -37,8 +41,31 @@ Editing tools beyond manual transcript cuts:
- **Captions** — font, size, color, position, and background-box styling,
plus an optional word-by-word karaoke-style highlight.

Export renders the final MP4 with every preview effect above (filters,
properties, logo, captions) baked in — the CSS-preview ↔ ffmpeg-export
For podcasts:

- **Long episodes** — uploads up to 10 GB; transcription runs in the
background in ~10-minute chunks split at pauses, with progress; a 720p
editing proxy keeps big 4K originals smooth in the browser.
- **Speakers** — "Detect speakers" labels who said what; rename a speaker
everywhere at once, or cut (and restore) everything one person said.
- **Audio** — a one-click podcast preset: -16 / -14 LUFS loudness,
noise reduction, speaker leveling and a rumble filter, with a
15-second before/after preview.
- **Publishing** — AI chapters (built into MP4/MP3 exports and copyable as
YouTube timestamps), editable show notes and title ideas, TXT/Markdown
transcripts, and MP3/WAV export.
- **Clips / Shorts** — AI picks standalone 15-90s highlights, or clip a
transcript selection; render 9:16 / 1:1 / 16:9 with a positionable crop
and big word-by-word captions.

AI models only ever return text and sentence picks, never timestamps or
edits: times always come from the transcript's own word timestamps, and
every response is validated before use. See the
[AI guide](https://amide-init.github.io/transcriptcut/guide/ai-editing)
for which model does what.

Export renders the final MP4 (or MP3/WAV) with every preview effect above (filters,
properties, logo, captions, audio cleanup) baked in — the CSS-preview ↔ ffmpeg-export
math for each is unit-tested (see Testing below) to catch the preview and
the actual export drifting apart, which happened in practice for several
of these (see the issue tracker's closed bugs).
Expand Down Expand Up @@ -82,9 +109,10 @@ pnpm run dev

## Running the app

Requires [FFmpeg](https://ffmpeg.org/download.html) and
Requires [FFmpeg](https://ffmpeg.org/download.html) (with `ffprobe`) and
[Bun](https://bun.sh) (the backend's runtime) on your machine, plus an
OpenAI API key (used server-side only, for transcription).
OpenAI API key (used server-side only, for transcription and the AI
buttons).

### macOS quick start

Expand Down Expand Up @@ -117,6 +145,16 @@ Open [http://localhost:5173](http://localhost:5173). See
[`server/.env.example`](./server/.env.example) for what each variable
does, and [`CONTRIBUTING.md`](./CONTRIBUTING.md) for more.

Two things to know during development:

- **The dev server and the Mac app both use port 3001**, so quit one
before starting the other. If the app shows "404 Not Found", a dev
server is still running.
- **`bun --watch` leaks a few file handles on every reload.** After a long
session with many server edits, FFmpeg can fail to start with
`EBADF: bad file descriptor, posix_spawn`. Restart `pnpm run dev` to
clear it.

### Download

A native macOS app is also available (Apple Silicon only) — no need to
Expand All @@ -133,7 +171,9 @@ cd client
pnpm exec tauri build
```

Produces `client/src-tauri/target/release/bundle/{macos,dmg}/`. See
Produces `client/src-tauri/target/release/bundle/{macos,dmg}/`. Installing
a newer build over an older one keeps your projects: the server upgrades
the app's database on launch. See
`claude.md` section 35 for how it's wired together, and
[`.github/workflows/build-macos-app.yml`](./.github/workflows/build-macos-app.yml)
for how CI builds and releases it on a version tag push.
Expand All @@ -142,8 +182,11 @@ for how CI builds and releases it on a version tag push.

Deterministic unit tests (Vitest) cover the core editing logic — timeline
cut/trim math, transcript-to-caption mapping, the ffmpeg render-plan
builder, path-traversal safety, and the CSS-preview/ffmpeg-export formula
parity work described in the issue tracker:
builder (cuts, audio cleanup, reframing, chapter metadata), chunked
transcription and speaker alignment, chapter and highlight validation,
database migrations, path-traversal safety, and the
CSS-preview/ffmpeg-export formula parity work described in the issue
tracker:

```bash
pnpm run test # runs both client and server test suites
Expand Down
88 changes: 85 additions & 3 deletions claude.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,7 +217,12 @@ Do not use GPT-5.6 Luna for every request.
Transcription is a separate concern from editing and isn't routed through
the GPT-4o-mini / GPT-5.6 Luna split above. Use OpenAI Whisper
(`whisper-1`, `verbose_json`, word + segment timestamp granularity) —
already implemented in `lib/ai/transcribe.ts`. Keep it behind a clear
already implemented in `lib/ai/transcribe.ts`. Long episodes are handled
by `lib/ai/transcription-job.ts`: it extracts a mono 16kHz 32kbps speech
track, splits it into ~10-minute chunks at detected silences
(`lib/ai/chunking.ts`), transcribes the chunks in parallel, and merges them
back onto the source timeline, so source file size never hits Whisper's
25MB upload cap. Keep it behind a clear
service interface (per section 29) so the provider can be swapped later.

---
Expand Down Expand Up @@ -575,6 +580,62 @@ child process on the same machine as everything else (see section 18).

---

## Podcast audio cleanup

Per-project `AudioSettings` (`server/src/types/audio-settings.ts`, stored as
`Project.audioSettingsJson`): loudness target (off / -16 / -14 LUFS), noise
reduction (afftdn), speaker leveling (dynaudnorm) and an 80Hz high-pass.
Enums/booleans only, mapped to fixed filter strings in
`lib/ffmpeg/audio-filters.ts`. Loudness is set by measured gain + a peak
limiter at -2 dBTP, not loudnorm's linear mode (which silently falls back
to dynamic mode and undershoots on peaky speech); `render-job.ts#resolveLoudnessGain`
measures and corrects with secant steps. Defaults are all off, so exports
are unchanged unless a user opts in.

## Podcast publishing

`PublishingMeta` (one per project) holds AI-generated, user-editable
chapters and show notes (`server/src/types/publishing.ts`). The model
(gpt-4o-mini) only returns text and *sentence indices* -- chapter starts
come from those sentences' own timestamps, never from a time the model
wrote. Chapters are stored in **source** time and mapped onto the edited
timeline when shown or exported (`lib/publishing/chapters.ts`), so later
cuts don't strand them. Every export (MP4, and MP3 at 44.1kHz/192k) embeds
the episode title and chapters via a generated FFMETADATA file; WAV can't
hold chapters. AI output never changes the edit.

## Speaker detection

An explicit "Detect speakers" action, not part of transcription: the
diarization model (`gpt-4o-transcribe-diarize`, behind `lib/ai/diarize.ts`)
runs at about half real time. It returns speaker-labeled spans but no word
timings, so Whisper stays the source of truth for words; diarization only
decides who said each word (`lib/ai/speakers.ts#assignSpeakers`), splitting
segments where the speaker changes. Labels are only consistent within one
API call, so `lib/ai/diarization-job.ts` keeps them consistent across
~10-minute chunks two ways: reference clips of the first chunk's voices
(`known_speaker_names`), and a 45s overlap between chunks used to match any
label the references missed. Results land in the existing `segment.speaker`
field ("Speaker 1", ...), so captions, exports and show notes use them with
no changes. Cutting a speaker is ordinary cut operations. Known limit:
similar-sounding voices can merge into one speaker.

## Clips / Shorts

`Clip` rows are source-time ranges with an aspect (9:16, 1:1, 16:9), a
horizontal crop position and a captions toggle. "Find highlights" is the
one feature routed to GPT-5.6 Luna (`gpt-5.6-luna`, section 4: selecting
important sections); like chapters, the model returns sentence indices,
never times, and `lib/clips/clips.ts#highlightsFromPicks` enforces length
(15-90s), whole sentences and no overlaps. A clip renders through the
normal export as the project **plus two extra cuts** (everything before
and after it), so the project's own edits and audio cleanup apply inside
it; `plan.ts` then crops/scales to the clip frame, and captions are
re-cut into 3-4 word Shorts cues laid out for that frame (the ASS script
resolution must match the output aspect, or libass stretches glyphs).

---

# 13. Architecture (local-first)

No cloud account is required to run or develop this project.
Expand Down Expand Up @@ -714,6 +775,8 @@ Transcript — belongs to Project
EditOperation — belongs to Project

RenderJob — belongs to Project

TranscriptionJob — belongs to Project (background chunked transcription status/progress)
```

No `User` model, no ownership checks for v1 (see section 2) — every
Expand Down Expand Up @@ -810,15 +873,34 @@ GET /api/projects
GET /api/projects/:id

POST /api/projects/:id/upload
POST /api/projects/:id/transcribe
POST /api/projects/:id/transcribe (202 + jobId; runs in the background)
GET /api/projects/:id/transcribe/:jobId (status + chunk progress)

GET /api/projects/:id/video[?variant=proxy]
GET /api/projects/:id/audio (extracted speech track, for the waveform)

GET /api/projects/:id/transcript

POST /api/projects/:id/operations
DELETE /api/projects/:id/operations/:operationId

POST /api/projects/:id/render
POST /api/projects/:id/render (body: format mp4|mp3|wav, captions)
GET /api/projects/:id/render/:jobId
POST /api/projects/:id/render/audio-preview (15s before/after sample of the audio settings)

GET /api/projects/:id/publishing (chapters on the edited timeline + show notes)
POST /api/projects/:id/publishing/chapters (AI-generate) PUT (save user edits)
POST /api/projects/:id/publishing/show-notes (AI-generate) PUT (save user edits)
GET /api/projects/:id/transcript/export?format=txt|md

POST /api/projects/:id/speakers/detect (202 + jobId; background speaker detection)
GET /api/projects/:id/speakers/detect/:jobId
POST /api/projects/:id/speakers/rename ({from, to|null} on every segment)

GET /api/projects/:id/clips
POST /api/projects/:id/clips/highlights (GPT-5.6 Luna picks; replaces earlier AI picks)
POST /api/projects/:id/clips PATCH/DELETE /api/projects/:id/clips/:clipId
(render a clip: POST /api/projects/:id/render with {clipId})
```

(No `POST /api/projects/:id/ai/edit` — the free-text AI command bar this
Expand Down
7 changes: 7 additions & 0 deletions client/src-tauri/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,13 @@ pub fn run() {
"FFMPEG_PATH",
"/opt/homebrew/opt/ffmpeg-full/bin/ffmpeg",
)
// Needed since transcription (audio duration), the editing
// proxy and the logo/captions sizing all probe media -- a
// bare "ffprobe" isn't on a GUI app's PATH, same as ffmpeg.
.env(
"FFPROBE_PATH",
"/opt/homebrew/opt/ffmpeg-full/bin/ffprobe",
)
.env(
"CLIENT_DIST_DIR",
server_dir.join("client-dist").display().to_string(),
Expand Down
Loading
Loading