AI Video Studio with Claude Code + Remotion
🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill
🎬 Use it: Talking-head video · AI story video · Deadpan comedy video
🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline
The first three parts built a studio for videos where you're on camera. This part builds one where nobody is. ai-story takes a short fable — an old man planting trees he'll never sit under, a ferryman who never asks where his passengers are going — and produces a narrated, illustrated vertical video in a well-known short-video format:
- a bold two-line title, frozen at the top for the whole video,
- an illustration in the middle that changes every few seconds with a slow zoom and smooth crossfades,
- a calm Vietnamese narrator,
- one subtitle line at a time underneath.
Everything runs on your Mac: narration from VieNeu-TTS, illustrations from Z-Image Turbo running in mflux on Apple's MLX framework, editing in Remotion. No API keys, no per-video cost. You need an Apple Silicon Mac with at least 16 GB of memory; the timings in this post were measured on an M3 Pro with 18 GB.
The skill's design is shaped by one number: an illustration takes about three and a half minutes. A 30-scene story is close to two hours of image generation, so the skill must get the script right before it draws anything. The full source is in the repository under .claude/skills/ai-story/ and remotion/src/story/.
.claude/skills/ai-story/
├── SKILL.md
├── references/
│ ├── storytelling.md ← structure, titles, scene cuts, anti-patterns
│ ├── output-format.md ← the story.json contract
│ ├── design-system.md ← layout numbers, measured from a reference video
│ └── setup.md ← install, voices, styles, speed, troubleshooting
├── scripts/
│ ├── setup.mjs new-dir.mjs validate-story.mjs build-docs.mjs
│ ├── tts.mjs tts_worker.py voice-sample.mjs
│ ├── images.mjs make-test-images.mjs
│ ├── build-props.mjs render.mjs lib.mjs
└── assets/
├── styles.json ← art-style presets
└── example/story.json
Step 1: The setup script
Two tools need installing, and the skill should do it rather than leave you a list of commands. setup.mjs checks for FFmpeg and uv, creates an isolated Python 3.11 environment for the TTS, installs VieNeu-TTS into it, and installs mflux as a uv tool:
if (!has('ffmpeg')) die('missing ffmpeg. Install with: brew install ffmpeg');
if (!has('uv')) die('missing uv. Install with: brew install uv');
run('uv', ['venv', join(toolDir(), 'venv'), '--python', '3.11'], 'create Python 3.11 venv');
run('uv', ['pip', 'install', '--python', py, '--upgrade', 'vieneu'], 'install VieNeu-TTS');
// torch is NOT needed for the built-in voices (they run on ONNX), but voice cloning
// requires it: the reference-voice encoder is a torch model. About 1 GB extra.
if (!hasFlag('no-clone')) run('uv', ['pip', 'install', '--python', py, 'torch', 'torchaudio'], 'install torch');
run('uv', ['tool', 'install', '--upgrade', 'mflux'], 'install mflux');
if (!has('mflux-generate-z-image-turbo')) {
warn('mflux is installed but not on PATH. Add to ~/.zshrc: export PATH="$HOME/.local/bin:$PATH"');
}
toolDir() is .ai-story/ at the project root (already in .gitignore). Add a --check mode that prints what's present and tries import vieneu in the venv; the skill tells Claude to run it whenever something fails.
Models download on first real use: about 300 MB for the TTS, and for the images, one choice saves you 27 GB. mflux's built-in name z-image-turbo points to full-precision weights — a 33 GB download that's quantized locally. The mflux author publishes a pre-quantized 4-bit build, filipstrand/Z-Image-Turbo-mflux-4bit, at 5.5 GB, which is plenty for flat illustration styles. Use that as the default (it needs mflux 0.13 or newer). Setting HF_TOKEN to a free Hugging Face read token speeds the download up considerably.
Step 2: The contract — story.json
{
"id": "2026-08-18-21-46-nguoi-trong-cay",
"title": { "line1": "NGƯỜI TRỒNG CÂY", "line2": "KHÔNG NGỒI DƯỚI BÓNG" },
"voice": "Thanh Bình",
"style": "doodle",
"seed": 20260818,
"cast": [
{
"id": "lao",
"look": "An old white stick-figure man with a smooth bald oval head, a short white beard, two tiny black dot eyes and a calm patient expression, wearing a simple brown vest over a white shirt and loose grey trousers rolled up to the knee, thin black stick arms and legs."
}
],
"scenes": [
{
"id": "s01",
"lines": ["Có một ông lão đào hố bên vệ đường đất, dưới cái nắng tháng sáu."],
"image": "He is digging a small hole with a worn shovel beside a dusty dirt road under a harsh summer sun, dry yellow fields stretching to the horizon, heat haze in the distance.",
"cast": ["lao"]
}
]
}
Three rules in this contract come straight from the models' limitations, and each becomes a validator check:
- Narration (
lines) in Vietnamese; image descriptions (image,cast[].look) in English. The image model doesn't understand Vietnamese. The validator rejects any Vietnamese diacritic in an image field. - Every recurring character is declared once in
castand referenced by id. The image model has no memory between scenes; the only way the old man looks the same in scene 7 is if scene 7's prompt describes him in exactly the same words. A goodlookcovers build, hair, eyes, top, bottom and one signature prop; the validator warns below 40 characters. The same person at two ages gets two entries. - One entry in
lines= one subtitle line = one audio file. That's what gives the video perfect subtitle timing (Step 5). Each line should be a full sentence under about 85 characters; over 110 is an error because it won't fit in two subtitle lines.
The title is two uppercase lines under 20 characters each (30 is an error) — it stays on screen for the whole video, so it has to be short enough to read at a glance.
Step 3: The craft reference — storytelling.md
As in Part 2, the reference is where the quality comes from. The core of it is a five-beat structure with a time budget, and one rule about endings:
| Beat | Share | Job |
|---|---|---|
| Open | 10% | Put the character in a specific place |
| Build | 30% | Show what they do, steadily |
| Clash | 25% | Someone mocks them, or something goes wrong |
| Turn | 25% | Time passes or the truth comes out |
| Settle | 10% | An action, not a lesson |
"He stood up. Then he went to find a shovel." beats "So let us all be grateful to those who came before." The reference also covers:
- title patterns (paradox, "who fears whom", "what it costs") and never spoiling the ending;
- when to cut to a new scene — when the picture must change, not because a sentence is long; one scene is 1–3 lines and 2.5–6 seconds;
- the genre's visual vocabulary for abstract ideas: a thought bubble with a picture inside, a fork in the road for a choice, the same landscape across seasons for passing time;
- a ban on attributing invented quotes to real people.
Step 4: Validate and STOP — the approval gate
validate-story.mjs enforces the contract and, just as importantly, estimates the cost so the approval step has real numbers. VieNeu-TTS reads at about 18.5 characters per second (measured, within 5%), and an image takes about 200 seconds:
const CHARS_PER_SEC = 18.5;
const estSec = totalChars / CHARS_PER_SEC + scenes.length * 0.4; // + pauses between scenes
sceneSecs.forEach((sec, i) => {
if (sec > 8) warn(`scenes[${i}] holds one picture for ~${sec.toFixed(1)}s — split it`);
if (sec < 1.2) warn(`scenes[${i}] is only ~${sec.toFixed(1)}s — the picture changes before anyone sees it`);
});
console.log(` ${scenes.length} scenes · ~${Math.round(estSec)}s`);
console.log(` image generation will take about ${Math.ceil((scenes.length * 200) / 60)} minutes`);
build-docs.mjs turns the JSON into story.md, a table of start time, duration, narration and scene description. Then the skill does something none of the earlier skills did — it stops:
### Step 7 — STOP and get approval
This is the most expensive step to skip. Images take ~3.5 minutes each; a 30-scene story
is nearly 2 hours. Editing the script after drawing throws those hours away.
Print: the two-line title, the scene table from story.md, the estimated duration and
the image-generation minutes. Ask the user to approve or request changes.
Only run step 8 after the user agrees.
That gate is the most valuable line in the skill. Approving a script takes a minute; regenerating thirty illustrations takes two hours.
Step 5: Narration with exact timing — tts.mjs + tts_worker.py
The selfie pipeline needed speech recognition to find out when each word was said. This one doesn't, because it creates the audio itself — one line at a time. Each line's duration is then a measured fact (samples ÷ sample rate), and the subtitle for that line can switch exactly when the voice does.
The TTS is a Python library and the pipeline is Node, so the bridge is a small worker that takes a JSON job on stdin and returns JSON on stdout (logs go to stderr so they stream without corrupting the result):
job = json.load(sys.stdin)
tts = Vieneu() # ONNX path; faster than MPS on Apple Silicon
voice = tts.add_voice("user", job["refAudio"], denoise=True) if job.get("refAudio") else job["voice"]
results = []
for item in job["lines"]:
wav = tts.infer(item["text"], voice=voice)
path = out_dir / f"{item['id']}.wav"
tts.save(wav, path)
results.append({"id": item["id"], "file": path.name, "sec": round(len(wav) / tts.sample_rate, 4)})
json.dump({"sampleRate": tts.sample_rate, "lines": results}, sys.stdout, ensure_ascii=False)
sys.stdout.flush(); sys.stderr.flush()
os._exit(0) # skip destructors: torch + ONNX occasionally crash on teardown AFTER success
The Node side then does three things:
- Trims silence from each file using a threshold relative to that file's own peak (6% of peak RMS, in 10 ms windows, keeping 40 ms of pad). VieNeu adds up to 500 ms of uneven silence per file, and a fixed dB threshold either kept audible breaths or clipped soft consonants.
- Lays out the timeline: 0.35 s lead-in, 0.16 s between lines in the same scene, 0.38 s between scenes, 0.9 s tail.
- Mixes everything in one FFmpeg pass, placing each line at its exact millisecond with
adelayand normalizing to −16 LUFS:
const delays = items.map((it, i) => `[${i}:a]adelay=${Math.round(it.startSec * 1000)}:all=1[a${i}]`).join(';');
const mix = items.map((_, i) => `[a${i}]`).join('');
const filter = `${delays};${mix}amix=inputs=${items.length}:normalize=0:dropout_transition=0[m];` +
`[m]apad=whole_dur=${totalSec},loudnorm=I=-16:TP=-1.5:LRA=11[out]`;
It writes audio/voice.wav and timing.json — the start and end of every line — which becomes the timing contract for the whole video. Measured speed: about 6 s to load the model, then 6.4× real time, so three minutes of narration takes about half a minute.
There are 14 built-in voices across Northern, Central and Southern accents; for fables, use the four storytelling voices (the news voices sound like a bulletin). --ref=my-voice.wav clones a voice from 5–10 seconds of clean speech. The repository also has voice-sample.mjs, which finds and cleans the best sample window in an existing video.
Step 6: Illustrations — images.mjs + styles.json
Art styles are data, not code. Each preset in styles.json has a long positive prompt, a negative prompt and a step count:
"doodle": {
"prompt": "flat 2D vector cartoon illustration, simple stick-figure explainer animation style. Drawing rules: each head is a large smooth white oval with two tiny black dot eyes ... Characters, props and background are all drawn with the same clean bold black outlines and flat colours ...",
"negative": "photograph, photorealistic, 3d render, ... text, letters, words, caption, watermark, ... extra person, blank featureless white figure, ...",
"steps": 8
}
Ship three: doodle (stick figures — by far the easiest to keep consistent), tranh (painted storybook) and muc (ink wash). Ban text in the negative prompt: the model misspells Vietnamese, and words belong in the subtitle layer.
images.mjs builds each prompt in a deliberate order — style, then the scene, then an explicit head count in words, then each present character's look — and calls mflux:
const count = cast.length
? `Exactly ${NUM[cast.length]} character${cast.length > 1 ? 's' : ''} in the frame, no one else.`
: '';
const prompt = [style.prompt, `Scene: ${sc.image}`, count, who].filter(Boolean).join(' ');
spawnSync('mflux-generate-z-image-turbo', [
'--model', 'filipstrand/Z-Image-Turbo-mflux-4bit',
'--prompt', prompt, '--negative-prompt', style.negative,
'--width', '1216', '--height', '784', // the panel's 1.552 ratio, with room to zoom
'--steps', String(style.steps),
'--seed', String(seed), // same seed for every scene → steadier characters
'--output', out,
]);
Without the head count, the model often added a blank white figure: it read the style's drawing rules as a description of another character. And there's no --guidance flag on purpose — Z-Image Turbo is a distilled model that ignores it (the same seed at 1.8 and 3.0 produced pixel-identical images).
Above all, make the stage resumable. It skips any scene whose image already exists, so an interrupted run loses nothing, and --only=s07 redraws one scene after you edit its description. --dry prints every prompt without drawing.
Step 7: The composition — StoryVideo
Register a second composition next to TikTokVideo, again with length from props. It has four layers:
export const StoryVideo: React.FC<StoryProps> = (props) => (
<AbsoluteFill style={{ backgroundColor: STORY.bg }}>
<ImagePanel shots={props.shots} />
<TitleBar title={props.title} />
<StoryCaptions captions={props.captions} />
{props.voice ? <Audio src={staticFile(props.voice.src)} /> : null}
{props.bgm ? <Audio src={staticFile(props.bgm.src)} volume={props.bgm.volume} loop /> : null}
</AbsoluteFill>
);
The layout was measured from a reference video in the genre and scaled to 1080×1920: a flat blue-grey background, title at 14.7–29.2% of the height, image panel at 31.8–68%, subtitles from about 70%, and the bottom quarter left empty for the platform's buttons.
TitleBaruses the Anton font (one of very few heavy condensed fonts with a Vietnamese subset), lime line one and red line two, and one shared size for both lines: each line is measured withfitTextfrom@remotion/layout-utilsand both use the smaller result, capped at 120 px. Line height is 1.4 rather than 1.3, or stacked diacritics on line two (Ồ, Ữ) hit line one. Add the package:pnpm add @remotion/layout-utils@4.0.512.ImagePanelcrossfades over 12 frames by fading only the incoming image on top of a fully opaque outgoing one — fading both reveals the background mid-transition — and gives each image a slow zoom that alternates direction (1.02→1.11, then 1.11→1.02) and origin, so it never looks mechanical.StoryCaptionsshows one line at a time in Nunito italic withtext-wrap: balance, timed straight fromtiming.json, lingering 0.28 s after the voice but never past the next line.
build-props.mjs is simple because there's nothing to guess: lines become captions, each scene's image starts when its first line starts and runs until the next scene's, and the background track (if you have one in _assets/bgm/) plays at 10% volume.
render.mjs adds one flag worth copying: --frame=240 renders a single frame to PNG with remotion still, so you can check layout in seconds instead of rendering the whole video.
Step 8: Test without waiting two hours
make-test-images.mjs writes one flat-colored placeholder per scene, so you can check the whole edit — layout, pacing, subtitle timing — before generating anything:
S=.claude/skills/ai-story/scripts
node $S/tts.mjs <slug>
node $S/make-test-images.mjs <slug> # placeholders — OVERWRITES images/
node $S/build-props.mjs <slug>
node $S/render.mjs <slug> --frame=240
Step 9: The procedure — SKILL.md
---
name: ai-story
description: Turn a life-lesson story into a 9:16 video made entirely by AI — narration
read by a local TTS, illustrations drawn by a local image model, no filming, no
recording. Writes content/<dir>/story.json and renders out/<dir>.mp4. Use when the user
pastes a story or a lesson, or says "AI story video", "philosophy video".
---
## Steps
0. First time: node .claude/skills/ai-story/scripts/setup.mjs (and --check any time).
1. Extract input: story or theme (required); voice (default "Thanh Bình"); voiceRef;
style (default "doodle"). A bare theme → invent a concrete story and SAY SO.
2. Read references/storytelling.md before writing anything.
3. Brainstorm internally: 3 two-line titles → keep the sharpest; the twist; the lesson.
4. node .claude/skills/ai-story/scripts/new-dir.mjs <slug>
5. Write story.json. Narration in Vietnamese, image and look in ENGLISH. Every recurring
character in cast.
6. validate-story.mjs && build-docs.mjs. Fix every ERROR.
7. STOP. Show the title, the scene table, the estimated duration and image minutes.
Wait for approval.
8. tts.mjs → images.mjs (long; can run in the background) → build-props.mjs → render.mjs
9. Report: mp4 path, real duration, scene count, and every WARN — a scene over 8 s,
a failed image (re-run with --only=), a line over 7 s.
The order in step 8 is deliberate: narration first, because it's fast and its timing file is needed by everything else; images second, because they're slow and can run in the background while you do something else.
What you've built
A second production line that needs no camera and no microphone, built on the same foundations as the first — one JSON contract, validators, generated docs, resumable stages, a typed props file — plus two new ideas worth reusing elsewhere:
- Estimate the cost in the validator and put a human approval gate in front of the expensive step.
- Let model limitations shape the schema: no memory → declare characters once; English-only → validate the language.
For a deeper look at the TTS trimming, the prompt order and the measured design numbers, see the local AI pipeline deep dive. Part 5 builds the last skill: a deadpan comedy writer that exports a printable PDF and plugs into the editing pipeline you built in Part 3.

