AI Video Studio with Claude Code + Remotion
🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill
🎬 Use it: Talking-head video · AI story video · Deadpan comedy video
🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline
The jump cut is the house style of talking-head video. Every breath, every "uh", every half-second of the speaker gathering their thoughts is cut out, and what remains is a dense stream of speech. Done by hand, it's the most tedious edit there is: scrub to the end of a sentence, find where the silence starts, cut, find where speech resumes, cut, delete, repeat forty times per minute of footage.
Detecting silence automatically is the easy half. FFmpeg has a filter for it, and you'll see the whole command in a minute. The hard half is everything downstream: the moment you remove 1.3 seconds from the middle of a clip, every word timestamp after that point is 1.3 seconds wrong, every sound effect anchored to a word lands on the wrong syllable, and the subtitles from part 2 drift further out of sync with every cut.
This post builds the whole jump-cut stage: detecting silence, deciding what to keep, mapping time from the original clip to the edited timeline, and the punch-in zoom that makes a single static camera look like a multi-angle edit. It ends with the variant for deadpan comedy, where the rule flips and some silences are the whole point.
Before cutting anything: preflight
Footage arrives as one file per scene: raw/s1.mov, raw/s2.mov, and so on. The first stage inspects every file with ffprobe and writes what it finds to media.json — duration, resolution, frame rate, and whether there's an audio stream at all. It stops the pipeline on the problems that can't be fixed in editing and warns on the ones that can:
| Check | Severity | Why |
|---|---|---|
| A scene has no file | ERROR | Lists exactly which sN is missing |
| A clip has no audio stream | ERROR | Almost always a microphone that wasn't connected. No edit can fix it; the scene must be re-shot |
| Duration unreadable | ERROR | Corrupt or still-copying file |
| Landscape clip | WARN | It will be cropped hard to 9:16 |
| Height under 1920 | WARN | Soft once it's scaled to the frame |
| Length differs >50% from the script's estimate | WARN | Informational: the timeline always follows the real footage |
Checking for audio first sounds like paranoia until the first time a render finishes ten minutes later with a silent scene in the middle. Failing in two seconds instead is the whole point of a preflight.
Detecting silence with FFmpeg
FFmpeg's silencedetect audio filter reports every stretch where the level stays below a noise floor for at least a minimum duration. It runs against the normalized 16 kHz WAV from the previous stage:
ffmpeg -i audio/s1.wav -af silencedetect=noise=-32dB:d=0.35 -f null -
[silencedetect @ 0x...] silence_start: 0
[silencedetect @ 0x...] silence_end: 0.9 | silence_duration: 0.9
[silencedetect @ 0x...] silence_start: 2.1
[silencedetect @ 0x...] silence_end: 2.7 | silence_duration: 0.6
[silencedetect @ 0x...] silence_start: 5.2
One gotcha costs people an afternoon: FFmpeg writes this log to stderr, not stdout. A helper that captures only stdout (Node's execFileSync, for instance) gets an empty string and silently finds no silences. Capture both streams:
/** ffmpeg logs (including silencedetect) go to STDERR. Capture both streams. */
export const shBoth = (cmd, args) => {
const r = spawnSync(cmd, args, { encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 });
return `${r.stdout ?? ''}${r.stderr ?? ''}`;
};
Parsing is a small state machine over the log lines. Note the last silence_start above has no matching silence_end: the clip ended while silent, so that open interval is closed at the clip's duration.
const findSilences = (wav) => {
const log = shBoth('ffmpeg', ['-i', wav, '-af', `silencedetect=noise=${db}dB:d=${minSilence}`, '-f', 'null', '-']);
const silences = [];
let open = null;
for (const line of log.split('\n')) {
const s = line.match(/silence_start:\s*(-?[\d.]+)/);
if (s) open = Math.max(0, Number(s[1]));
const e = line.match(/silence_end:\s*(-?[\d.]+)/);
if (e && open !== null) {
silences.push([open, Number(e[1])]);
open = null;
}
}
return { silences, trailing: open }; // trailing: silence that ran to the end
};
From silences to segments worth keeping
The edit is defined by what you keep, not what you cut, because that's what the renderer needs. Kept segments are the complement of the silences, widened slightly at both ends:
const keep = [];
let cursor = 0;
for (const [s, e] of silences) {
const segEnd = Math.min(dur, s + pad); // let the last syllable ring out
if (segEnd - cursor > 0.12) keep.push([cursor, segEnd]);
cursor = Math.max(0, Math.min(dur, e - pad)); // start a hair before speech resumes
}
if (dur - cursor > 0.12) keep.push([cursor, dur]);
Three details carry the quality:
- Padding (80 ms by default) on both sides of each cut. Cut exactly at the detected boundary and you clip the soft consonant at the start of a word and the decay at the end of the last one. It sounds chopped even when you can't say why.
- A minimum segment length of 120 ms. Anything shorter is a click or a breath that the detector caught between two silences; keeping it produces a one-frame flash of video.
- No cutting inside speech. The detector only ever proposes cuts inside silences, so a cut can never land in the middle of a word — even when the speaker says "uh". Cutting mid-word sounds far worse than leaving a filler in.
And one safety net: if no speech is found at all (an empty keep), the whole clip is kept with a warning. A mis-set threshold should never silently delete a scene.
A worked example
Take a 6.0-second clip with the three silences from the log above — [0, 0.9], [2.1, 2.7], and a trailing one from 5.2 to the end — and the default 80 ms pad:
| Silence | Segment proposed | Kept? | Cursor moves to |
|---|---|---|---|
[0, 0.9] | [0, 0.08] | No — under 120 ms | 0.82 |
[2.1, 2.7] | [0.82, 2.18] | Yes | 2.62 |
[5.2, 6.0] | [2.62, 5.28] | Yes | 5.92 |
| end of clip | [5.92, 6.0] | No — under 120 ms | — |
The result is two segments, [[0.82, 2.18], [2.62, 5.28]]: 4.02 seconds kept, 1.98 seconds removed. That's what lands in cuts.json:
{
"scenes": {
"s1": { "durationSec": 6.0, "keep": [[0.82, 2.18], [2.62, 5.28]], "removedSec": 1.98 }
}
}
Tuning, and the escape hatch
The defaults suit a lapel mic in a quiet room. When they don't suit your footage:
| Flag | Default | Adjust when |
|---|---|---|
--silence-db | -32 | Cuts are missed (room noise reads as speech) → lower to -40. Cuts bite into words → raise to -26 |
--min-silence | 0.35 | The edit feels choppy → raise to 0.5 |
--pad | 0.08 | Words sound clipped → raise to 0.15 |
--no-cut | — | Keep every clip intact |
If more than half of a clip disappears, the stage warns: that's almost always a quiet microphone making speech look like silence, and --silence-db=-40 usually fixes it.
The most useful property of this stage is that cuts.json is meant to be edited by hand. If the detector trims one breath too aggressively, change one number and re-run from the props stage. Nothing upstream — audio extraction, transcription — needs to run again. That's why the cut list is data in a file, not something computed on the fly during rendering.
Turning segments into Remotion clips
Each kept segment becomes one clip on the timeline. Seconds become frames, and the clip records where it sits on the final timeline (fromFrame) and which part of the source file it plays (trimBefore/trimAfter):
const sec = (s) => Math.round(s * fps);
for (const [a, b] of cut.keep) {
const trimBefore = sec(a);
const trimAfter = sec(b);
const durationInFrames = trimAfter - trimBefore;
if (durationInFrames <= 0) continue;
clips.push({
src: m.src,
fromFrame: cursor, // position on the edited timeline
durationInFrames,
trimBefore, // where in the source to start
trimAfter, // where in the source to stop
zoom: clipIndex % 2 === 0 ? 1 : ZOOM,
sceneId: m.sceneId,
});
segments.push({ srcStart: a, srcEnd: b, frames: durationInFrames });
cursor += durationInFrames;
clipIndex++;
}
In Remotion, each clip is a Sequence wrapping an OffthreadVideo with the trim applied. OffthreadVideo extracts exact frames with FFmpeg during rendering instead of relying on the browser's video element, which is what you want when cuts must be frame-accurate. The speech plays from the footage itself, so the audio is cut exactly where the picture is:
<OffthreadVideo
src={staticFile(clip.src)}
trimBefore={clip.trimBefore}
trimAfter={clip.trimAfter}
style={{
width: '100%',
height: '100%',
objectFit: 'cover',
transform: `scale(${clip.zoom})`,
}}
/>
The function that keeps subtitles in sync
Here is the part that's easy to get almost right. Whisper's timestamps describe the original clip. The renderer needs positions on the edited timeline, where the removed silences no longer exist. Every timestamp has to be pushed through the cuts:
// Map a time in the ORIGINAL clip to a frame on the EDITED timeline.
// The easiest place in the whole pipeline to be off by a few frames.
const toTimeline = (srcSec) => {
let acc = 0; // frames of kept footage before this point
for (const s of segments) {
if (srcSec < s.srcStart) return sceneStart + acc; // fell in a removed gap
if (srcSec <= s.srcEnd) {
return sceneStart + acc + Math.round((srcSec - s.srcStart) * fps);
}
acc += s.frames; // skip over this whole segment
}
return sceneStart + acc; // after the last segment
};
Walk through it with the example above at 30 fps. The first kept segment, [0.82, 2.18], spans source frames 25 to 65 — 40 frames. The second, [2.62, 5.28], spans frames 79 to 158 — 79 frames.
Now take a word Whisper heard at 3.0 seconds in the original clip:
- It's past the end of segment one, so the 40 frames of segment one are added to
acc. - It falls inside segment two: 3.0 − 2.62 = 0.38 s into it, which is 11 frames.
- Result: frame 51 from the start of the scene.
Without the mapping, that word would be drawn at frame 90 (3.0 × 30) — a full 1.3 seconds after it's actually spoken in the edit. With four or five cuts per scene, the error compounds into subtitles that describe a sentence the speaker finished long ago.
The first branch handles a subtle case: a timestamp that falls inside a removed silence (Whisper occasionally places a word boundary there). It snaps to the start of the next kept segment, which is exactly where the speech resumes on screen. Every mapped word is then clamped to its own scene's range, so a slightly-long final word can't leak into the next scene.
The same function serves every time-anchored element: word timestamps, caption pages, and sound effects anchored to words. Putting the mapping in one function — instead of re-deriving it in three places — is what keeps them consistent with each other.
Punch-in zoom: one camera, many angles
A jump cut on a locked-off camera has a known weakness: the speaker's head twitches slightly between two frames that are almost the same shot. Editors hide it with a punch-in — alternating between the full frame and a slightly tighter crop at each cut. The cut then reads as a change of camera angle rather than a glitch.
The pipeline does this for free: even-numbered clips play at scale 1.0, odd-numbered ones at 1.06.
const ZOOM = 1.06; // 1.03 if it feels jumpy, 1.10 if it feels flat
zoom: clipIndex % 2 === 0 ? 1 : ZOOM,
That's also why the shooting guide asks for 4K. Scaling 1080p footage to 106% throws away resolution; scaling 4K footage down to a 1080-wide frame, then 6% back in, still has plenty of pixels to spare.
The deadpan variant: when silence is the joke
Pipeline C produces deadpan comedy — a performer explaining something absurd with total seriousness. Run that footage through the standard cutter and the comedy dies, for two reasons. The slow, deliberate delivery becomes rushed when every pause is cut to zero. And the twist scene depends on a two-second silence after the punchline — the stare into the camera is the joke.
The comedy script format gives the editor two instructions to handle this.
edit.leaveGapSec: 0.3 applies to the whole video. Silences between sentences are still cut, but 0.3 s of each is left behind, split across both sides of the cut. Silences that are already shorter than what would be left are skipped entirely — cutting them would only create two overlapping segments:
const inner = s > 0.05 && e < dur - 0.05; // not at the very start or end
if (inner && e - s <= leaveGap + 2 * pad) continue; // too short to bother cutting
const half = inner ? leaveGap / 2 : 0;
const segEnd = Math.min(dur, s + pad + half);
// ...
cursor = Math.max(0, Math.min(dur, e - pad - half));
keepPauses: true applies to individual scenes, usually the twist. Silence inside the scene is never cut at all. Only the head and tail are trimmed, and even then up to 2.5 seconds of silence is kept before the first word — because "stare silently, then speak" is a beat, not dead air:
const MAX_LEAD_PAUSE = 2.5;
const MAX_TAIL_PAUSE = 1.0;
const lead = silences.find(([s]) => s <= 0.05);
const tail = silences.find(([, e]) => e >= dur - 0.05 && !(lead && e === lead[1]));
const speechStart = lead ? lead[1] : 0;
const speechEnd = tail ? tail[0] : dur;
keep = [[Math.max(0, speechStart - MAX_LEAD_PAUSE), Math.min(dur, speechEnd + MAX_TAIL_PAUSE)]];
The twist scene also gets a snap zoom: in four frames the frame lurches to 140%, centered on the performer's eye line, and holds there. The zoom origin sits at 38% from the top rather than at the center, because a seated performer's eyes are above the middle of the frame:
const SNAP_FRAMES = 4;
const SNAP_SCALE = 1.4;
const scale = interpolate(frame + fx.sceneOffset, [0, SNAP_FRAMES], [1, SNAP_SCALE], {
extrapolateLeft: 'clamp',
extrapolateRight: 'clamp',
easing: Easing.out(Easing.cubic),
});
// transformOrigin: '50% 38%'
The zoom is driven by the frame offset within the scene (sceneOffset), not within the clip, so if the scene still contains a cut, the zoom continues smoothly across it instead of restarting. A gentler push mode — a slow creep from 100% to 114% across the whole scene — suits the "serious lecture" parts.
Anchoring sound effects to the edit
Sound effects in the comedy format are anchored to moments, not timecodes: the start of a scene, its end, its longest pause, or a specific word in the dialogue. All of those resolve against the edited timeline, after the mapping above:
function sfxFrame(spec) {
if (spec.word) {
// the first word of the phrase, compared without accents
const key = norm(tokenize(spec.word)[0]);
const hit = words.find((w) => norm(w.text) === key);
if (hit) return hit.fromFrame;
return sceneStart; // and warn that the word wasn't found
}
if (spec.at === 'end') return Math.max(sceneStart, sceneEnd - sec(0.6));
if (spec.at === 'pause') {
// the longest gap between words, including the silence before the first word
let best = { at: sceneStart, len: words.length ? words[0].fromFrame - sceneStart : 0 };
for (let i = 1; i < words.length; i++) {
const len = words[i].fromFrame - words[i - 1].toFrame;
if (len > best.len) best = { at: words[i - 1].toFrame, len };
}
return best.at + sec(0.1);
}
return sceneStart;
}
at: "pause" is what puts crickets chirping into the silence after the punchline, wherever that silence ends up after editing. word: "180" puts a thud exactly on the number in "a 180-day contract." The script author describes the comedy; the editor works out the frame.
What the stage reports
The props builder finishes with a one-line summary — clips, caption pages, titles, total length — and two warnings worth reading:
- The finished video is much longer or shorter than the script estimated. That's normal. The timeline follows the real footage; the script's timings were a target for filming, not a constraint.
- Caption pages were dropped. The cuts were so aggressive that several words collapsed onto the same frames. Re-run detection with
--silence-db=-40or--no-cut.
Takeaways
- Let FFmpeg find silence, but define the edit as segments to keep, padded at both ends and never shorter than a few frames.
- Store the cut list as an editable file. Hand-fixing one cut should never require re-running analysis.
- Every timestamp in the system goes through one mapping function from source time to edited time. Subtitles, captions and sound effects stay in sync because they share it.
- Alternate a small punch-in zoom at each cut to turn a static camera into an "edited" look.
- Make silence handling a per-scene editorial decision. In comedy, some pauses are the content.
The music under all of this — and the crickets, thuds and record scratches — isn't downloaded from anywhere. Part 4 shows how the pipeline synthesizes all of it from scratch in plain JavaScript.

