AI Video Studio with Claude Code + Remotion
🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill
🎬 Use it: Talking-head video · AI story video · Deadpan comedy video
🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline
Editing short vertical videos is the part of content creation that refuses to scale. Writing a 40-second script takes ten minutes. Filming it takes another ten. Then comes an hour in an editor: cutting the dead air between sentences, typing subtitles word by word, nudging each caption until it lands on the right syllable, hunting for background music you are actually allowed to use, and placing a "pop" sound exactly where the joke lands. Do that three times a week and editing becomes the job.
This series walks through a working system that removes the editing hour almost entirely. You give Claude Code a sentence — a proverb, a pickup line, a short fable — and you get back a filmable script. You record yourself (or don't record at all), run one command, and get a finished 1080×1920 MP4 with jump cuts, animated word-by-word subtitles, background music and sound effects.
The interesting part isn't that "AI edits video." It doesn't. The interesting part is the architecture that makes this reliable: Claude writes the creative data, deterministic scripts do everything else, and Remotion turns that data into frames. This first post covers that architecture — the ideas you can reuse in any Claude Code automation, not just video. The next four posts go deep on the hardest individual problems.
What the finished system does
There are three production lines, each driven by its own Claude Code skill:
| Pipeline | You provide | You get | Human effort |
|---|---|---|---|
| A — Selfie talking head | A quote or opinion | A 30–45s script + a shooting guide, then a finished edit of your footage | Film 4–6 short clips |
| B — AI illustrated story | A short fable or life lesson | A narrated video with AI illustrations, fully generated on your machine | Approve the script |
| C — Deadpan comedy | A saying, a pickup line or a story | A 5-act comedy script as a print-ready PDF, then a finished edit of your performance | Perform 5–12 short clips |
All three end in the same place: a vertical MP4 rendered by Remotion. Pipelines A and C share one editing engine; pipeline B has its own, because it has no camera footage at all.
Remotion in one paragraph
Remotion lets you write a video as a React component. Your component receives the current frame number through useCurrentFrame() and returns what that frame should look like. A video is just a function from frame number to pixels, so anything you can lay out with HTML and CSS — text, images, video clips, SVG — can be animated with ordinary code, previewed in a browser-based Studio, and rendered to MP4 from the command line:
import { AbsoluteFill, interpolate, useCurrentFrame } from 'remotion';
export const Hello: React.FC = () => {
const frame = useCurrentFrame();
const opacity = interpolate(frame, [0, 30], [0, 1], { extrapolateRight: 'clamp' });
return (
<AbsoluteFill style={{ justifyContent: 'center', alignItems: 'center', fontSize: 120, opacity }}>
Xin chào
</AbsoluteFill>
);
};
That is the entire mental model. Everything in this series is built from Sequence (show something from frame X for N frames), OffthreadVideo and Img (media), Audio, and interpolate/spring (motion).
A licensing note before you build on it: Remotion is free for individuals and small companies, but larger companies need a paid company license. Check the current terms on the Remotion site before using it commercially.
The core idea: video as data
The single decision that makes this whole system work is this: the Remotion composition never reads the script, and Claude never touches the timeline. Between them sits a single JSON file, props.json, that describes every clip, caption, title and sound with exact frame numbers. A chain of small Node scripts produces that file from the script and the raw media.
flowchart LR
U[You: a quote or a story] --> C[Claude Code skill]
C -->|writes + validates| S[script.json]
S --> D[Generated docs: script.md, shoot.md, PDF]
S --> P[Deterministic pipeline: Node + ffmpeg + whisper]
F[Your footage raw/s1.mov ...] --> P
P -->|writes| J[props.json]
J --> R[Remotion composition]
R --> O[out/video.mp4]
That split gives each participant the job it is good at:
- Claude is good at the creative, fuzzy work: finding a hook, writing dialogue with a real rhythm, picking where the twist goes, describing an illustration. It is not reliable at counting frames, measuring audio, or remembering that scene 3 was re-shot.
- Scripts are good at the precise, boring work: measuring silence in a WAV file, mapping a word's timestamp through a set of cuts, converting seconds to frames. They never get creative.
- Remotion is good at turning a fully specified description into pixels, deterministically, every time.
When something looks wrong in the final video, this separation tells you exactly where to look. Wrong words? The script. Words appear at the wrong moment? The timing stage. Wrong font? The composition. Nothing is ever "the AI was in a weird mood."
Anatomy of a Claude Code skill
A Claude Code skill is a folder with a SKILL.md file. The front matter has a name and a description; Claude reads the description to decide when the skill applies, and loads the full body only when it does. The body is a procedure written for Claude, and the folder can carry reference documents and scripts that the procedure tells Claude to read or run.
Each skill in this project follows the same layout:
.claude/skills/tiktok-script/
├── SKILL.md ← the procedure: steps, hard rules, what to output
├── references/
│ ├── framework.md ← the "craft": hook library, timing budget, anti-patterns
│ └── output-format.md ← the JSON contract, field by field
├── scripts/
│ ├── new-dir.mjs ← creates content/<date>-<time>-<slug>/
│ ├── validate-script.mjs ← schema + pacing checks, zero dependencies
│ └── build-docs.mjs ← script.json → script.md + shoot.md
└── assets/example/ ← one complete, valid example
The description in the front matter matters more than it looks. It is the only part Claude sees before deciding to use the skill, so it should say what the skill produces and list the phrases a user would actually type:
---
name: tiktok-script
description: Create a 9:16 selfie-style TikTok script (30–45s) from a proverb,
quote or opinion. Writes content/<date>-<time>-<slug>/script.json (the data
contract for the Remotion editing pipeline) and script.md. Use when the user
says "write a script", "make a video about this quote", or pastes a quote and
wants a short video.
---
The body then reads like a runbook: Step 1 — extract the input. Step 2 — read references/framework.md before writing a single word. Step 3 — brainstorm three hooks internally, keep the sharpest. … Splitting the craft knowledge into references/ keeps SKILL.md short enough that Claude follows it precisely, while the long creative guidance is only loaded at the step where it's needed.
Six design rules that make skills reliable
These are the patterns that turned "Claude sometimes writes a good script" into "every run produces a usable file." None of them are specific to video.
1. One JSON file is the only source of truth
Every pipeline has exactly one file that Claude writes — script.json, story.json or kichban.json — and everything human-readable is generated from it:
| Pipeline | Claude writes | Generated from it |
|---|---|---|
| Selfie | script.json | script.md (script table), shoot.md (shooting guide with cue cards) |
| AI story | story.json | story.md (scene table for approval) |
| Deadpan | kichban.json | kichban.html → kichban.pdf (the printable performance script) |
The skill says it bluntly: never write the .md files by hand, never edit them by hand. If Claude wrote both the JSON and the Markdown, they would drift apart within two edits — the Markdown would say one thing, the JSON (which the renderer consumes) another. Generating the docs makes that impossible, and it means formatting changes live in one script instead of in Claude's memory.
2. Validators are the contract, and Claude must run them
Each JSON format has a zero-dependency validator, and the skill makes running it a non-negotiable step, chained so that docs are only built from a valid file:
node .claude/skills/tiktok-script/scripts/validate-script.mjs content/<dir>/script.json && \
node .claude/skills/tiktok-script/scripts/build-docs.mjs content/<dir>/script.json
The validators check more than the schema. They encode the craft rules that are easy to violate while being creative:
- Scenes must be contiguous: each
startSecequals the previousendSec, the lastendSecequalsdurationSec. - The hook is at most 4 seconds; the call-to-action is 4–10 seconds.
- Speaking pace is 4.0–6.5 words per second for selfie videos (a WARN above 7), and a slower 2.5–4.5 for deadpan delivery.
- Every highlighted keyword must appear verbatim in the caption text, accents and case included — otherwise the renderer has nothing to highlight.
- In the AI-story format, image descriptions must be in English (the image model doesn't understand Vietnamese), and the validator rejects any Vietnamese diacritics in that field.
ERRORs block; WARNs must be either fixed or explained to the user. The rule in the skill is "don't report done while an ERROR remains," which turns the validator into a feedback loop Claude iterates against rather than a check you run afterwards.
3. Let scripts own the facts Claude can't know
Every video gets a folder named content/<YYYY-MM-DD>-<HH>-<MM>-<slug>/, so ls content/ sorts in the order things were written. Claude knows today's date but not the local time on your machine, and it will happily invent one. So the skill forbids typing the timestamp:
node .claude/skills/tiktok-script/scripts/new-dir.mjs nhan-chi-so-tinh-ban-thien
# → 2026-08-18-09-23-nhan-chi-so-tinh-ban-thien (folder created)
The script reads the system clock, appends -v2 if the same slug already exists in the same minute, creates the folder, and prints the name. The id field inside the JSON must match that folder name exactly, and the validator checks it. The general rule: whenever a value comes from the environment (time, file durations, what's on disk), a script produces it and Claude copies it.
4. A staged pipeline with files in between
The editing engine is six small scripts, each reading the previous stage's output and writing its own:
script.json + raw/*.mov
├─[1] preflight → media.json every clip present, has audio, vertical?
├─[2] extract-audio → audio/sN.wav 16 kHz mono, loudness-normalized
├─[3] transcribe → timing.json word timestamps from whisper.cpp
├─[4] detect-cuts → cuts.json the segments to keep (hand-editable)
├─[5] build-props → props.json everything re-timed to the real footage
└─[6] render → out/<dir>.mp4
This costs a little ceremony and pays for itself immediately:
- Debuggable. Every stage's output is a JSON file you can open. When subtitles drift, you look at
timing.jsonand see whether the timestamps are wrong or the mapping is. - Re-runnable. Stages cache by file modification time. Re-shoot scene 3, and only scene 3's audio is re-extracted.
- Hand-correctable. If the automatic cutter trims one breath too aggressively, you edit one line in
cuts.jsonand re-run from stage 5. No re-transcription, no re-analysis. - Fail early. Stage 1 refuses to continue if a clip has no audio track — almost certainly a disconnected microphone — instead of letting you discover it after a ten-minute render.
Every script also accepts a partial slug. Typing node render.mjs co-cong-mai-sat finds content/2026-08-18-09-23-co-cong-mai-sat/ on its own, picking the newest folder if several match.
5. Put a human gate in front of anything expensive
Some steps are cheap to redo, others are not. Generating one illustration locally takes about three and a half minutes; a 30-scene story is close to two hours of machine time. So the AI-story skill has an explicit stop:
Step 7 — STOP and get approval. Print the two-line title, the scene table, the estimated duration and the number of minutes image generation will take. Only run step 8 after the user agrees.
The deadpan-comedy skill does the same before producing the PDF: it prints the script table, the thumbnail text and two alternative directions, then waits. Feedback in plain language — "make the twist in scene 3 harsher", "cut it to 50 seconds", "go with direction 1" — is applied to the JSON, re-validated, and shown again. The PDF only exists once you say "ok, export." Without the gate, the cheapest possible mistake (a weak script) becomes the most expensive one.
6. Orchestrate; don't duplicate domain knowledge
The editing skill says up front: this skill contains no Remotion knowledge. When a component needs changing, it points Claude to the official Remotion skills (/remotion-markup for animation and layout, /remotion-captions for captions, /remotion-docs for API lookups). Copying framework documentation into your own skill means it goes stale the day the framework releases; delegating keeps your skill about your pipeline only.
The Remotion side: two compositions, one input each
The Remotion project registers two compositions — one for footage-based videos, one for illustrated stories. Both take their entire configuration from props, and both compute their duration from those props rather than hardcoding it:
<Composition
id="TikTokVideo"
component={TikTokVideo}
width={1080}
height={1920}
fps={30}
durationInFrames={placeholder.durationInFrames}
defaultProps={placeholder}
calculateMetadata={({ props }) => ({
durationInFrames: props.durationInFrames,
fps: props.fps,
})}
/>
calculateMetadata is what lets a 38-second video and a 3-minute video share one composition: the length is whatever props.json says. The empty placeholder props render a friendly "no props yet" screen that also prints ắ ầ ẫ ễ ộ ữ ỹ ặ ọ Đ đ — a one-glance check that the font is actually loading its Vietnamese subset.
Rendering points Remotion's public directory at the content/ folder, so staticFile('<dir>/raw/s1.mov') and staticFile('_assets/bgm/lofi.mp3') both resolve without copying media around:
pnpm exec remotion render TikTokVideo out/<dir>.mp4 \
--public-dir ../content --props=../content/<dir>/props.json
The composition itself is just layers, each a list of Sequences positioned by frame:
<AbsoluteFill style={{ backgroundColor: COLORS.ink, fontFamily }}>
{props.clips.map((clip, i) => (
<Sequence key={i} from={clip.fromFrame} durationInFrames={clip.durationInFrames}>
<SceneClip clip={clip} /> {/* footage segment + zoom */}
</Sequence>
))}
{props.captions.map((page, i) => (
<Sequence key={i} from={page.fromFrame} durationInFrames={page.durationInFrames}>
<Subtitles page={page} /> {/* 2–3 words, current word highlighted */}
</Sequence>
))}
{props.titles.map(/* big block titles */)}
<ProgressBar />
{props.sfx.map((s, i) => (
<Sequence key={i} from={s.atFrame}>
<Audio src={staticFile(s.src)} volume={s.volume} />
</Sequence>
))}
</AbsoluteFill>
Because the composition only ever sees frame numbers, it has no idea that the source clip was 11 seconds long and lost 3 seconds of silence. That complexity lives in one place — the props builder — which is the subject of part 3.
Setting it up
You need Node, pnpm, ffmpeg, and (for the AI-story pipeline) uv for Python tooling. On macOS:
brew install ffmpeg uv
npx create-video@latest remotion # pick the blank template, TypeScript
cd remotion && pnpm install
pnpm add @remotion/captions @remotion/google-fonts \
@remotion/install-whisper-cpp @remotion/layout-utils
Skills live in .claude/skills/<name>/ at the project root, so Claude Code discovers them as soon as you open the project. If you also want Remotion's official skills, install them the same way and reference them by name from your own.
A CLAUDE.md-level rule worth adding: pick one package manager and say so. The skills here say "use pnpm, never npm/npx", because mixing lockfiles in a Remotion project is a reliable way to end up with two React copies.
A complete run, end to end
Here is what pipeline A looks like from the user's side.
1. Write the script.
/tiktok-script "Nhân chi sơ, tính bản thiện. Không cà khịa, rất ngứa mồm."
Claude reads the framework reference, brainstorms three hooks internally, picks one, creates the folder, writes script.json, validates it until clean, and generates two documents. It replies with the script table and points you at shoot.md — the file to open when you pick up the camera. That file has the setup checklist, the recommended filming order (hook last, once you're warmed up) and a cue card per scene: three to five short phrases taken from your own lines, meant to be read at a glance next to the lens, not a teleprompter.
2. Film. One file per scene, named by scene id: raw/s1.mov, raw/s2.mov, … This naming convention is the contract that makes everything downstream possible. It gives the pipeline the true length of each block, natural cut boundaries, and the ability to re-shoot one scene without redoing the rest.
3. Edit.
/tiktok-video nhan-chi-so-tinh-ban-thien
The six stages run in order and stop at the first ERROR. The skill then reads the warnings back to you in plain language — a low transcription match rate, a scene that lost more than half its length to silence trimming, a fallback to evenly spaced subtitles — because "the pipeline finished" and "the video is good" are different claims.
Reusing one engine for a second format
The deadpan-comedy pipeline writes a different script format — five acts, per-scene music changes, sound effects anchored to specific words, scenes that must keep their awkward silences. Rather than building a second editing engine, every stage reads its script through one function:
/** Read the script of a content/<dir>/, translating from kichban.json if needed. */
export const loadScript = (dir) => {
const kPath = join(dir, 'kichban.json');
const sPath = join(dir, 'script.json');
if (!existsSync(kPath)) return readJson(sPath);
// A kichban.json is present → script.json is always a derived translation.
if (existsSync(sPath) && statSync(sPath).mtimeMs >= statSync(kPath).mtimeMs) {
const s = readJson(sPath);
if (s.derivedFrom === 'kichban.json') return s;
}
const script = convertKichban(readJson(kPath));
writeJson(sPath, script);
return script;
};
The converter maps the five comedy acts onto the four roles the engine already understands (hook, story, punchline, cta), carries over the new per-scene fields (music, keepPauses, zoom, anchored sfx), and stamps derivedFrom so nobody edits the translation by mistake. The engine gained a handful of optional behaviors — gentler silence trimming, music that switches per scene — but no second code path. It's the adapter pattern, and it's what lets a hobby project grow a third format without growing a third pipeline.
Testing without a camera
Every pipeline can be exercised end to end without real media:
S=.claude/skills/tiktok-video/scripts
node $S/make-test-footage.mjs <slug> --busy # synthetic clips (OVERWRITES raw/)
node $S/transcribe.mjs <slug> --fallback # skip whisper, spread words evenly
node $S/detect-cuts.mjs <slug> && node $S/build-props.mjs <slug> && node $S/render.mjs <slug>
The --busy flag deserves a mention. A solid-color test background proves nothing about whether subtitles are readable, because everything is readable on flat gray. --busy uses ffmpeg's testsrc2 pattern — saturated, high-contrast and moving — which is about the worst background real footage will ever give you. One of the subtitle design rules in part 2 — how faint an upcoming word is allowed to get — was only discovered because of it.
For the illustrated-story pipeline, make-test-images.mjs produces one flat color per scene so you can check layout and pacing in seconds, and render.mjs <slug> --frame=120 renders a single frame to PNG instead of the whole video.
What to take away
Strip away the video specifics and the recipe applies to any Claude Code automation that produces something precise:
- Let Claude write one structured file of creative decisions; generate everything human-readable from it.
- Encode your quality bar in a validator and make running it part of the procedure.
- Anything that comes from the environment — time, durations, file contents — is measured by a script, never guessed.
- Break the work into stages with inspectable files between them.
- Put a human approval gate in front of the expensive step.
- Keep skills about your workflow; delegate framework knowledge to the framework's own docs or skills.
The rest of the series takes the hard parts one at a time. Part 2 tackles the problem that nearly sank the project: speech recognition that gets Vietnamese diacritics wrong, and how to get word-perfect subtitles from it anyway.

