AI Video Studio with Claude Code + Remotion
🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill
🎬 Use it: Talking-head video · AI story video · Deadpan comedy video
🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline
Part 2 produced a script and a shooting guide. You've filmed one clip per scene — raw/s1.mov, raw/s2.mov and so on. This part builds tiktok-video, the skill that turns those clips into a finished 1080×1920 MP4 with:
- jump cuts that remove the silence between sentences,
- word-by-word animated subtitles whose spelling comes from the script and whose timing comes from speech recognition,
- a block title at the start of each scene,
- background music and sound effects, synthesized by code,
- a punch-in zoom at each cut and a progress bar.
It's the biggest skill in the series, and unlike Part 2 it involves almost no creative writing. Claude's job here is to orchestrate: run six scripts in order, stop at the first error, and translate the warnings into plain advice. So most of this part is about the scripts and the Remotion composition. The full source is in the project repository under .claude/skills/tiktok-video/ and remotion/src/; this post walks through what each piece does and the code that matters.
The pipeline
script.json + raw/*.mov
├─[1] preflight.mjs → media.json clips exist, have audio, are vertical
├─[2] extract-audio.mjs → audio/sN.wav 16 kHz mono, loudness-normalized
├─[3] transcribe.mjs → timing.json word timestamps (whisper.cpp), text from the script
├─[4] detect-cuts.mjs → cuts.json segments to keep — hand-editable
├─[5] build-props.mjs → props.json everything re-timed to the real footage
└─[6] render.mjs → out/<dir>.mp4
Every stage is a standalone script that reads the previous stage's file and writes its own. That costs a little ceremony and buys three things: you can open any intermediate file to see what went wrong, re-run one stage without the others, and fix a bad cut by editing one line of cuts.json.
The skill folder:
.claude/skills/tiktok-video/
├── SKILL.md
├── references/
│ ├── pipeline.md ← every stage, every flag, every intermediate file
│ ├── design-system.md ← colors, fonts, safe areas, caption presets
│ └── troubleshooting.md ← symptom → cause → fix
└── scripts/
├── lib.mjs shared helpers
├── preflight.mjs extract-audio.mjs transcribe.mjs align.mjs
├── detect-cuts.mjs build-props.mjs render.mjs
├── audio-synth.mjs music + SFX synthesizer
└── make-test-footage.mjs fake clips for testing without a camera
Step 1: Add the speech-recognition package
Word timing comes from whisper.cpp, which Remotion wraps in a package that downloads, builds and runs it from Node. Add it at the same exact version as the rest of Remotion:
cd remotion
pnpm add @remotion/install-whisper-cpp@4.0.512
Whisper itself and its model (about 1.5 GB for medium) are downloaded on first use into .whisper/ at the project root — already in .gitignore from Part 1.
Step 2: Shared helpers — lib.mjs
Every stage needs the same few things, so they live in one dependency-free module. Two helpers deserve a closer look.
Slug resolution. Folder names are long (2026-08-18-09-23-co-cong-mai-sat), and typing them is miserable. Every script accepts any suffix and finds the folder, picking the newest if several match:
export const resolveSlug = (arg) => {
if (existsSync(contentDir(arg))) return arg;
const hits = readdirSync(join(ROOT, 'content'), { withFileTypes: true })
.filter((d) => d.isDirectory() && !d.name.startsWith('_'))
.map((d) => d.name)
.filter((n) => n.endsWith(`-${arg}`))
.sort(); // fixed-width timestamps → sorts by time
if (hits.length === 0) die(`no content/ folder matches "${arg}"`);
const picked = hits[hits.length - 1];
if (hits.length > 1) info(`"${arg}" matches ${hits.length} folders, using the newest: ${picked}`);
return picked;
};
Capturing FFmpeg's log. FFmpeg writes its diagnostics — including the silence detector's output — to stderr. Node's execFileSync returns only stdout, so a naive helper silently gets nothing. Use spawnSync and join both streams:
export const shBoth = (cmd, args) => {
const r = spawnSync(cmd, args, { encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 });
return `${r.stdout ?? ''}${r.stderr ?? ''}`;
};
The rest is small: ok/info/warn/die print with a fixed prefix (so Claude can scan for ERROR and WARN), flagValue('silence-db', '-32') reads --silence-db=-40-style flags, isFresh(out, src) compares modification times for caching, and norm/tokenize normalize Vietnamese text for comparison.
Step 3: The contract between scripts and Remotion — props.json
Write the TypeScript type first; it's the spec the props builder must satisfy. All frame numbers are absolute positions on the final timeline:
export type VideoProps = {
slug: string;
fps: number;
durationInFrames: number;
bgm: { src: string; volume: number } | null; // one track for the whole video
music?: MusicSegment[]; // or one per scene (Part 5)
clips: Clip[]; // footage segments that survived the jump cuts
titles: Title[]; // the big block titles
captions: CaptionPage[];// subtitle pages of 2–3 words
sfx: Sfx[];
};
export type Clip = {
src: string; // relative to content/, e.g. "<dir>/raw/s1.mov"
fromFrame: number; // where it sits on the timeline
durationInFrames: number;
trimBefore: number; // which part of the source file it plays
trimAfter: number;
zoom: number; // 1.0 or 1.06, alternating
sceneId: string;
};
The composition never reads script.json and never does arithmetic on seconds. If the video looks wrong, either props.json is wrong (a script bug) or the rendering of a correct props.json is wrong (a component bug). That split makes every bug easy to place.
Step 4: The composition
Register it in Root.tsx next to HelloVideo, with its length taken from the props:
<Composition
id="TikTokVideo"
component={TikTokVideo}
width={1080}
height={1920}
fps={30}
durationInFrames={placeholder.durationInFrames}
defaultProps={placeholder}
calculateMetadata={({ props }) => ({ durationInFrames: props.durationInFrames, fps: props.fps })}
/>
TikTokVideo.tsx is nothing but layers of Sequences:
export const TikTokVideo: React.FC<VideoProps> = (props) => (
<AbsoluteFill style={{ backgroundColor: COLORS.ink, fontFamily }}>
{props.clips.map((clip, i) => (
<Sequence key={`clip-${i}`} from={clip.fromFrame} durationInFrames={clip.durationInFrames}>
<SceneClip clip={clip} />
</Sequence>
))}
{props.captions.map((page, i) => (
<Sequence key={`cap-${i}`} from={page.fromFrame} durationInFrames={page.durationInFrames}>
<Subtitles page={page} />
</Sequence>
))}
{props.titles.map((title, i) => (
<Sequence key={`title-${i}`} from={title.fromFrame} durationInFrames={title.durationInFrames}>
<TitleCard title={title} />
</Sequence>
))}
<ProgressBar />
{props.bgm ? <Audio src={staticFile(props.bgm.src)} volume={props.bgm.volume} loop /> : null}
{props.sfx.map((s, i) => (
<Sequence key={`sfx-${i}`} from={s.atFrame}>
<Audio src={staticFile(s.src)} volume={s.volume} />
</Sequence>
))}
</AbsoluteFill>
);
The components, briefly:
SceneClipplays one footage segment withOffthreadVideo(frame-exact extraction during render),trimBefore/trimAfterfrom the clip,objectFit: 'cover', andtransform: scale(zoom). The speech audio comes from the footage itself.Subtitlesdraws a page of 2–3 words at about 74% of the frame height. The word being spoken gets a rounded yellow box behind it, a small spring "pop", and dark text; upcoming words are slightly dimmed. The details — and why the dimming stops at 72% opacity — are in the subtitles deep dive.TitleCardshowscaption.textat about 22% height for the first 2.6 seconds of each scene. Each role gets its own entrance — the hook slides in from the left, story titles rise, the punchline assembles word by word with a beat before the highlighted phrase, the CTA lifts up with a bouncing arrow — because viewers remember motion, and four identical entrances become predictable by the third scene.ProgressBaris eight pixels at the top:
export const ProgressBar: React.FC = () => {
const frame = useCurrentFrame();
const { durationInFrames } = useVideoConfig();
const pct = Math.min(1, frame / Math.max(1, durationInFrames - 1));
return (
<AbsoluteFill style={{ justifyContent: 'flex-start' }}>
<div style={{ height: 8, width: '100%', backgroundColor: 'rgba(255,255,255,0.18)' }}>
<div style={{ height: '100%', width: `${pct * 100}%`, backgroundColor: COLORS.yellow }} />
</div>
</AbsoluteFill>
);
};
Keep every color, size and position in one tokens.ts. Two values in it matter for every vertical video: the platform's own interface covers about 150 px at the top and 220 px at the bottom (SAFE = { top: 150, bottom: 220 }), and the middle of the frame is the speaker's face — so titles go in the upper third and subtitles in the lower third, never in the center. And every outlined text uses paint-order: stroke fill; without it, the outline is painted over the letters and Vietnamese diacritics (ế, ộ, ữ) clog into black blobs.
Before you change any of these components, have Claude read /remotion-markup (or /remotion-captions for subtitles). The skill says so explicitly.
Step 5: The six stages
Stage 1 — preflight.mjs: fail in two seconds, not ten minutes
For each scene, find raw/<id>.{mov,MOV,mp4,MP4,m4v}, probe it with ffprobe, and record duration, size and frame rate in media.json. Stop on what editing can't fix; warn on what it can:
const probe = ffprobe(path);
const v = probe.streams.find((s) => s.codec_type === 'video');
const a = probe.streams.find((s) => s.codec_type === 'audio');
if (!a) fail(`${scene.id}: NO AUDIO. The microphone almost certainly wasn't connected — re-shoot.`);
// iPhones store a landscape sensor frame plus a ±90° rotation matrix.
// The *displayed* frame is portrait, so swap width and height when rotated.
const rotation = Number(v.side_data_list?.find((d) => d.rotation != null)?.rotation ?? v.tags?.rotate ?? 0);
const quarterTurn = Math.abs(rotation) % 180 === 90;
const width = Number(quarterTurn ? v.height : v.width);
const height = Number(quarterTurn ? v.width : v.height);
if (width > height) warn(`${scene.id}: filmed in LANDSCAPE — it will be cropped hard to 9:16`);
That rotation check is the kind of thing you only learn by testing with real phone footage: without it, every iPhone clip reports as landscape.
Stage 2 — extract-audio.mjs: one clean WAV per scene
ffmpeg -y -i raw/s1.mov -vn -ac 1 -ar 16000 -c:a pcm_s16le \
-af loudnorm=I=-14:TP=-1.5:LRA=11 audio/s1.wav
16 kHz mono 16-bit is the format whisper.cpp requires. loudnorm brings each scene to −14 LUFS so clips recorded at slightly different mic distances sound continuous. The stage skips any WAV newer than its source clip; --force redoes all of them.
Stage 3 — transcribe.mjs: timing from Whisper, text from the script
Install whisper.cpp and the model once, then transcribe each scene with token-level timestamps, in Vietnamese, primed with the scene's own lines:
await installWhisperCpp({ to: whisperPath, version: '1.5.5' });
await downloadWhisperModel({ model: 'medium', folder: whisperPath });
const raw = await transcribe({
model: 'medium', whisperPath, whisperCppVersion: '1.5.5',
inputPath: wav, tokenLevelTimestamps: true, language: 'vi', printOutput: false,
additionalArgs: [['--prompt', scene.voiceover]],
});
const { captions } = toCaptions({ whisperCppOutput: raw });
const heard = mergeSubwordTokens(captions); // " nh"+"ẹ"+"o" → "nhẹo"
const aligned = alignWords(scene.voiceover, heard, durationMs);
The crucial rule: subtitles never show Whisper's text. Whisper gets Vietnamese diacritics wrong constantly, and a misspelled subtitle is the most visible error a video can have. align.mjs matches what Whisper heard against the script with the Needleman–Wunsch algorithm on accent-stripped text, then takes the spelling from the script and only the timestamps from Whisper. If too little matches, it falls back to spreading the words evenly — wrong rhythm, but never wrong text. The full algorithm, its scoring and its self-test are in the subtitles deep dive. Ship align.mjs with its test (node align.mjs --test) and make the skill re-run it whenever the file changes.
Two flags make the stage practical: --fallback skips Whisper entirely (useful before you've downloaded 1.5 GB), and the script refuses any *.en model up front, since English-only models can't transcribe Vietnamese.
Stage 4 — detect-cuts.mjs: the jump cuts
Run FFmpeg's silence detector on each WAV, parse the log, and keep the complement of the silences with a small pad on each side:
const log = shBoth('ffmpeg', ['-i', wav, '-af', `silencedetect=noise=${db}dB:d=${minSilence}`, '-f', 'null', '-']);
// parse silence_start / silence_end pairs → silences: [[s, e], ...]
const keep = [];
let cursor = 0;
for (const [s, e] of silences) {
const segEnd = Math.min(dur, s + pad);
if (segEnd - cursor > 0.12) keep.push([cursor, segEnd]);
cursor = Math.max(0, Math.min(dur, e - pad));
}
if (dur - cursor > 0.12) keep.push([cursor, dur]);
Defaults: --silence-db=-32, --min-silence=0.35, --pad=0.08. If no speech is found at all, keep the whole clip with a warning — a wrong threshold must never silently delete a scene. Write the result to cuts.json, and don't add it to .gitignore: it's the one intermediate file you'll edit by hand, and its history is worth keeping. Tuning and the comedy variant are covered in the jump-cuts deep dive.
Stage 5 — build-props.mjs: put it all on one timeline
This is the only stage with real arithmetic. For each scene, in order:
- Turn each kept segment into a
Clip(seconds → frames, alternating zoom 1.0 / 1.06), advancing a timeline cursor. - Map every word timestamp from source time to timeline frame through the kept segments — the one function the whole sync depends on:
const toTimeline = (srcSec) => {
let acc = 0;
for (const s of segments) {
if (srcSec < s.srcStart) return sceneStart + acc; // inside a removed silence
if (srcSec <= s.srcEnd) return sceneStart + acc + Math.round((srcSec - s.srcStart) * fps);
acc += s.frames;
}
return sceneStart + acc;
};
- Group words into subtitle pages: at most 3 words, a new page after a 0.35 s pause, no page longer than 1.6 s, each held up to 0.4 s after its last word.
- Add the block title (2.6 s from the start of the scene) and the scene's sound effects.
- Resolve music and SFX names to files, synthesizing any that are missing.
Step 5 deserves a note. Rather than ship audio files, the project synthesizes four background tracks and fourteen effects from code (FM electric piano, plucked strings, filtered-noise drums) the first time they're requested. Files you put in content/_assets/bgm/ or sfx/ yourself always take priority. The props builder just calls:
import { ensureMusic, ensureSfx } from './audio-synth.mjs';
if (ensureMusic(script.audio.bgm)) bgm = { src: `_assets/bgm/${script.audio.bgm}.mp3`, volume: script.audio.bgmVolume };
How the synthesizer works is its own deep dive. You can start with plain MP3s in _assets/ and add synthesis later.
Finish by printing a one-line summary (clips, pages, titles, seconds) and the warnings that matter: subtitle pages dropped because cuts were too aggressive, or a finished length very different from the script's estimate (normal — the footage is the truth).
Stage 6 — render.mjs
Same as Part 1, with TikTokVideo, the props file, and content/ as the public directory — plus a --studio flag that opens Remotion Studio on the same props so you can scrub the timeline:
const common = ['--public-dir', join(ROOT, 'content'), `--props=${propsPath}`];
if (hasFlag('studio')) {
spawnSync('pnpm', ['exec', 'remotion', 'studio', ...common], { cwd: remotionDir, stdio: 'inherit' });
process.exit(0);
}
spawnSync('pnpm', ['exec', 'remotion', 'render', 'TikTokVideo', outFile, ...common], { cwd: remotionDir, stdio: 'inherit' });
Step 6: Test without a camera — make-test-footage.mjs
You'll change this pipeline often, and filming test clips every time isn't realistic. This script generates one fake clip per scene: a colored background, a sine tone as "speech", and one second of silence at each end so the cutter has something to do:
const bg = busy
? `testsrc2=s=1080x1920:d=${dur}:r=30` // busy, high-contrast pattern
: `color=c=${COLORS[i % COLORS.length]}:s=1080x1920:d=${dur}:r=30`;
sh('ffmpeg', [
'-y', '-f', 'lavfi', '-i', bg,
'-f', 'lavfi', '-i', `sine=f=${180 + i * 40}:d=${speech}`,
'-af', `adelay=1000|1000,apad=whole_dur=${dur}`, // 1 s silence before and after
'-c:v', 'libx264', '-pix_fmt', 'yuv420p', '-preset', 'ultrafast', '-c:a', 'aac',
'-t', String(dur), join(rawDir, `${scene.id}.mov`),
]);
Use --busy when testing subtitle readability: text is readable on any flat color, so a flat background proves nothing. FFmpeg's testsrc2 pattern is about the worst background real footage will ever give you. And put a loud warning in the script's header and the skill: it overwrites raw/.
Step 7: The orchestrator — SKILL.md
With the scripts in place, the skill is short. Its job is order, stopping, and honest reporting:
---
name: tiktok-video
description: Edit a 9:16 TikTok video from content/<dir>/script.json and self-recorded
footage in content/<dir>/raw/. Runs a 6-stage pipeline — preflight, audio extraction,
whisper word timing, silence jump cuts, props build, Remotion render — into
out/<dir>.mp4 with animated subtitles. Use when the user says "edit the video",
"render the video", or has finished filming.
---
# TikTok Video Builder
This skill is an ORCHESTRATOR. It contains no Remotion knowledge — before editing
remotion/src/, read /remotion-markup (layout, animation) or /remotion-captions.
## Steps
0. First time only: `cd remotion && pnpm install`; check `which ffmpeg`.
1. Pick the slug from the request, or `ls content/`. Exactly one folder → use it, don't ask.
2. Run in order and STOP at the first ERROR:
S=.claude/skills/tiktok-video/scripts
node $S/preflight.mjs <slug>
node $S/extract-audio.mjs <slug>
node $S/transcribe.mjs <slug>
node $S/detect-cuts.mjs <slug>
node $S/build-props.mjs <slug>
node $S/render.mjs <slug>
3. Read the WARNs. Always report these to the user:
- match below 70% in transcribe → they drifted far from the script; subtitle timing may slip
- more than 50% cut in detect-cuts → microphone too quiet; suggest --silence-db=-40
- "fell back" → subtitles are correct but won't follow the speech rhythm
4. Report: the mp4 path and duration; that a length different from the script is normal;
every warning above.
## Hard rules
- Never edit script.json — it belongs to tiktok-script. A wrong script → re-run that skill.
- Subtitle text always comes from voiceover, never from Whisper.
- Missing music or SFX never blocks the render.
- make-test-footage.mjs OVERWRITES raw/. Never run it when real footage exists.
"Pipeline finished" and "video is good" are different claims, and step 3 is what keeps the skill from conflating them. A run that falls back to evenly spaced subtitles still produces an MP4; without that step, Claude would report success and you'd discover the drift on your phone.
Step 8: Run it end to end
First with fake footage, using the script folder from Part 2:
S=.claude/skills/tiktok-video/scripts
node $S/make-test-footage.mjs co-cong-mai-sat --busy
node $S/preflight.mjs co-cong-mai-sat
node $S/extract-audio.mjs co-cong-mai-sat
node $S/transcribe.mjs co-cong-mai-sat --fallback
node $S/detect-cuts.mjs co-cong-mai-sat
node $S/build-props.mjs co-cong-mai-sat
node $S/render.mjs co-cong-mai-sat
Each stage prints one OK line. detect-cuts should report about two seconds removed per scene (the fake silences), and the MP4 should show a busy test pattern, alternating zoom at each cut, a title at the start of each scene and evenly timed subtitles. Then copy your real clips into raw/ and, in Claude Code:
/tiktok-video co-cong-mai-sat
The first real run downloads whisper.cpp and the model. After that, a 40-second video takes a few minutes end to end, most of it rendering.
What to put in the references
The three reference files are for the moments when something isn't right, so write them as lookup tables:
pipeline.md— every stage, its flags with defaults and "change it when…", its output file, and whether that file is safe to edit by hand.design-system.md— the token table, the safe-area diagram, caption and title presets per role, and the reasons behind the non-obvious numbers (the 0.09 outline ratio, the 0.72 minimum opacity).troubleshooting.md— symptom → cause → fix, per stage: "file HAS NO AUDIO", "more than 50% cut", "accents render as boxes", "black screen, no footage" (emptyclipsbecause everything was cut).
The skill doesn't load them by default; it points to them. That keeps the procedure short while making sure Claude finds the right fix when you report a problem.
What you've built
You now have a complete selfie-video studio: Part 2 writes the script, you film, this part edits. The patterns that made it manageable:
- One stage, one script, one output file. Debuggable, cacheable, re-runnable.
- A typed props contract between the scripts and Remotion.
- The composition is pure layout; all timing arithmetic happens in one builder.
- A skill that orchestrates and reports honestly instead of re-explaining Remotion.
Part 4 removes the camera: a skill that writes a fable, narrates it with a local Vietnamese text-to-speech model, illustrates it with a local image model, and edits it — all on your Mac.

