Subtitles That Never Misspell: Whisper for Timing, Your Script for Text

Cover Image for Subtitles That Never Misspell: Whisper for Timing, Your Script for Text
Video AI5 min read

AI Video Studio with Claude Code + Remotion

🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill

🎬 Use it: Talking-head video · AI story video · Deadpan comedy video

🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline

Animated word-by-word subtitles are the signature of short-form video: two or three words at a time, the word being spoken lit up in a colored box, the whole line bouncing slightly as the speaker moves through it. They keep people watching with the sound off, and they keep people reading with the sound on.

The standard recipe is to run speech recognition, take its words and timestamps, and draw them. For English that works well enough. For Vietnamese it produces a steady stream of visible mistakes, because Vietnamese is a tonal language that writes its tones as diacritics, and recognition models get them wrong all the time. ngứa mồm comes back as ngưa mom. chưa kịp comes back as chức hiệp. Each one is a misspelling in giant bold letters in the middle of the screen — the single most visible error a video can have.

This post shows how the pipeline gets word-perfect subtitles from imperfect recognition, with every word appearing at the moment it's spoken. The trick is one sentence long; the details that make it hold up on real footage take the rest of the post.

The insight: you already know the words

In this pipeline, the speaker is reading lines that Claude wrote into script.json. The correct text of every scene — accents, capitalization and punctuation included — is sitting right there in scene.voiceover. What we don't know is when each word is spoken.

So the two sources split the job:

Source of truth
Which words, spelled howscene.voiceover from the script
When each word starts and endsWord timestamps from Whisper

Whisper never contributes a single character to the screen. It contributes only timestamps. The job becomes an alignment problem: match the words Whisper heard to the words in the script, then copy the timestamps across.

There's a complication that makes this harder than a simple zip. The shooting guide deliberately tells the speaker not to memorize the script — "say it in your own words" — because memorized lines sound recited. So real speech drifts: an extra "à" here, a skipped word there, a phrase repeated after a stumble. The aligner has to tolerate wrong accents, inserted words and missing words, all at once.

Step 1: Clean audio in the format Whisper expects

Each scene's clip becomes a WAV file in the exact format whisper.cpp requires — 16 kHz, mono, 16-bit PCM — with loudness normalization applied on the way:

ffmpeg -y -i raw/s1.mov -vn -ac 1 -ar 16000 -c:a pcm_s16le \
       -af loudnorm=I=-14:TP=-1.5:LRA=11 audio/s1.wav

loudnorm does more than it appears to. Five scenes filmed separately will have the microphone at five slightly different distances from the speaker's mouth, and the volume jump at each cut is obvious. Normalizing every scene to −14 LUFS — the common target for social platforms — makes the joins sound continuous, and gives the silence detector in part 3 a consistent level to work with.

Step 2: Word timestamps from whisper.cpp

Remotion ships a helper package, @remotion/install-whisper-cpp, that downloads and builds whisper.cpp and its models into a folder of your choice and wraps transcription in a Node API. The transcription stage uses it like this:

import { installWhisperCpp, downloadWhisperModel, transcribe, toCaptions }
  from '@remotion/install-whisper-cpp';

const whisperPath = join(ROOT, '.whisper');
await installWhisperCpp({ to: whisperPath, version: '1.5.5' });
await downloadWhisperModel({ model: 'medium', folder: whisperPath });

const raw = await transcribe({
  model: 'medium',
  whisperPath,
  whisperCppVersion: '1.5.5',
  inputPath: 'audio/s1.wav',
  tokenLevelTimestamps: true,
  language: 'vi',
  printOutput: false,
  // Prime the decoder with the script itself.
  additionalArgs: [['--prompt', scene.voiceover]],
});

const { captions } = toCaptions({ whisperCppOutput: raw });

Three choices here matter:

Model size. medium (about 1.5 GB, downloaded once and reused for every video) is the practical minimum for Vietnamese; large-v3-turbo is better if your machine can take it. tiny, base and small make so many errors that the aligner has too few anchors to work with. And any model ending in .en is English-only — the script refuses those up front rather than letting you download 1.5 GB to find out.

Prompting with the script. whisper.cpp's --prompt biases the decoder toward the vocabulary you give it. Passing the scene's own lines noticeably increases how often Whisper lands on the script's actual words — "chín giờ" comes back as "9 giờ" rather than the meaningless "chỉnh dơ", so at least "giờ" now matches. We don't need Whisper to be right; we need it to be right often enough to anchor the timeline, and every extra matched word is one more anchor.

Token-level timestamps. With tokenLevelTimestamps: true you get a timestamp per token instead of per segment, which is what word-level subtitles need. It also introduces the next problem.

Step 3: Merge subword tokens back into words

Whisper's tokenizer is byte-pair encoding, and Vietnamese syllables are often split into pieces: " nh" + "ẹ" + "o", or "H" + "ôm". If you align those fragments directly against script words, every fragment looks like an extra word the speaker inserted, and the match rate collapses.

The rule for merging is simple: a token that doesn't start with a space continues the previous word.

export const mergeSubwordTokens = (tokens) => {
  const words = [];
  for (const t of tokens) {
    const prev = words.at(-1);
    if (prev && !/^\s/.test(t.text)) {
      prev.text += t.text;                       // glue onto the previous word
      prev.endMs = Math.max(prev.endMs, t.endMs); // and extend its end time
    } else {
      words.push({ text: t.text.trim(), startMs: t.startMs, endMs: t.endMs });
    }
  }
  return words.filter((w) => norm(w.text).length > 0); // drop pure punctuation
};

A merged word keeps the start time of its first fragment and the end time of its last, which is exactly the span the word occupies in the audio.

Step 4: Compare words without their accents

Since the whole problem is wrong diacritics, comparison happens on a normalized form that has none: decompose to NFD, strip the combining marks, map đ to d, lowercase, and drop everything that isn't a letter or digit.

/** Strip Vietnamese accents + lowercase + drop non-alphanumerics.
 *  For COMPARING only — never for display. */
export const norm = (s) =>
  String(s ?? '')
    .normalize('NFD')
    .replace(/[̀-ͯ]/g, '')
    .replace(/đ/g, 'd')
    .replace(/Đ/g, 'D')
    .toLowerCase()
    .replace(/[^a-z0-9]/g, '');

After normalization, ngứa, ngưa and ngua are all ngua. The original strings are never thrown away — the script word is what goes on screen; norm only decides whether two words are "the same."

Normalization alone isn't enough, because Whisper also mishears whole syllables. So the scoring has three levels: an exact normalized match, a near match (Levenshtein similarity of at least 0.6 between the normalized forms), and a mismatch.

const MATCH = 3;      // same word once accents are stripped
const NEAR = 1;       // similar enough to be a mishearing
const MISMATCH = -2;
const GAP = -1;       // a word present on one side only

const similarity = (a, b) => 1 - levenshtein(a, b) / Math.max(a.length, b.length);
const score = (a, b) => (a === b ? MATCH : similarity(a, b) >= 0.6 ? NEAR : MISMATCH);

Step 5: Align with Needleman–Wunsch

Matching two sequences that may each contain extra or missing elements is the classic global sequence alignment problem, and the classic answer is the Needleman–Wunsch algorithm — originally designed for DNA, perfectly suited to "the script" versus "what was actually said." It fills a dynamic-programming table where each cell holds the best possible score for aligning the first i script words with the first j heard words:

const A = script.map(norm);           // script words, normalized
const B = heard.map((h) => norm(h.text)); // Whisper words, normalized

const dp = Array.from({ length: A.length + 1 }, () => new Float64Array(B.length + 1));
for (let i = 1; i <= A.length; i++) dp[i][0] = i * GAP;
for (let j = 1; j <= B.length; j++) dp[0][j] = j * GAP;

for (let i = 1; i <= A.length; i++) {
  for (let j = 1; j <= B.length; j++) {
    dp[i][j] = Math.max(
      dp[i - 1][j - 1] + score(A[i - 1], B[j - 1]), // pair script word i with heard word j
      dp[i - 1][j] + GAP,                           // script word i was not heard
      dp[i][j - 1] + GAP,                           // heard word j is not in the script
    );
  }
}

The three moves map directly onto the three things a speaker can do: say a scripted word (possibly mangled), skip one, or add one. The scoring makes pairing two similar words always better than leaving them unpaired, and makes an outright mismatch costly enough that the algorithm prefers a gap.

Walking back from the bottom-right cell recovers the alignment, and this is where the rule "text from the script, time from Whisper" is enforced:

while (i > 0 || j > 0) {
  if (i > 0 && j > 0 && dp[i][j] === dp[i - 1][j - 1] + score(A[i - 1], B[j - 1])) {
    const good = score(A[i - 1], B[j - 1]) > MISMATCH;
    if (good) matches++;
    out.push({
      text: script[i - 1],                         // TEXT from the script
      startMs: good ? heard[j - 1].startMs : null, // TIME from Whisper
      endMs: good ? heard[j - 1].endMs : null,
    });
    i--; j--;
  } else if (i > 0 && dp[i][j] === dp[i - 1][j] + GAP) {
    // Whisper didn't hear this word — still show it, fill its time in later.
    out.push({ text: script[i - 1], startMs: null, endMs: null });
    i--;
  } else {
    // The speaker said something off-script. Do NOT show it: its only
    // spelling comes from Whisper, and Whisper's spelling is what we don't trust.
    j--;
  }
}
out.reverse();

The third branch is a deliberate trade-off worth stating plainly: ad-libbed words don't get subtitles. Showing them would mean showing Whisper's spelling, which is the exact failure this design exists to prevent. A subtitle that briefly lags an improvised aside is far less noticeable than a misspelled one.

Step 6: Fill the gaps and keep time moving forward

After alignment, some script words have timestamps and some have null — the words Whisper missed or mangled beyond recognition. Those are filled by interpolation:

  • Words before the first known timestamp are spread evenly between 0 and that timestamp.
  • Words after the last known timestamp are spread between it and the end of the scene (with a floor of 120 ms per word, so a late cluster isn't squashed to nothing).
  • A run of unknown words in the middle is spread evenly between the end of the known word before it and the start of the known word after it.

A final pass guarantees what the renderer depends on: every word starts no earlier than the previous one ended, and every word lasts at least 60 ms.

for (let q = 1; q < out.length; q++) {
  if (out[q].startMs < out[q - 1].endMs) out[q].startMs = out[q - 1].endMs;
  if (out[q].endMs <= out[q].startMs) out[q].endMs = out[q].startMs + 60;
}

Step 7: Know when to give up — on timing, never on spelling

Alignment reports a matchRatio: the fraction of script words that found a good partner in what Whisper heard. That number drives an explicit set of modes, recorded in timing.json per scene:

ModeMeaningWhat you see
alignedNormal. Below 70% match, the editing skill warns you the speaker drifted far from the scriptCorrect words, correct timing
fallback-low-matchUnder 40% matched — wrong file, wrong language, or very noisyCorrect words, evenly spaced
fallback-no-audioWhisper heard nothingCorrect words, evenly spaced
fallback-even--fallback flag: Whisper skipped entirelyCorrect words, evenly spaced

Every fallback divides the scene duration evenly across the script's words. The subtitles won't hit each syllable, but they will never be misspelled. That's the trade-off the whole design makes on purpose: when something has to give, it's the rhythm, not the text.

The --fallback mode has a second use: you can build a complete draft of a video before downloading the 1.5 GB model at all.

Testing the aligner

An algorithm like this fails silently — a slightly wrong traceback produces subtitles that are mostly right — so it carries its own self-test, runnable with node align.mjs --test. The suite covers seven scenarios in 22 assertions, and the most important one is the reason the file exists:

console.log('2. Whisper drops the accents — TEXT must come from the script');
const r = alignWords('Không cà khịa rất ngứa mồm', [
  w('Khong', 0, 300), w('ca', 300, 500), w('khia', 500, 800),
  w('rat', 800, 1000), w('ngua', 1000, 1300), w('mom', 1300, 1600),
], 1600);
check('shown with correct accents', r.words.map((x) => x.text).join(' ') === 'Không cà khịa rất ngứa mồm');
check('still counted as aligned', r.mode === 'aligned');
check('matchRatio = 1', r.matchRatio === 1);

The others cover an inserted filler word (it must not appear, and the next word must take its timestamp from after the filler), dropped words (they must be interpolated between their neighbors, never lost), garbage input (must fall back), silence, and BPE fragments (must merge before aligning, or the match rate collapses). The rule in the skill is simple: change align.mjs, re-run the test.

From timed words to TikTok-style pages

Timed words aren't subtitles yet. They're grouped into pages of two or three words, the unit that appears on screen together. The grouping starts a new page whenever any of these is true:

const MAX_WORDS_PER_PAGE = 3;   // never more than three words on screen
const MAX_GAP_SEC = 0.35;       // a pause longer than this starts a new page
const MAX_PAGE_SEC = 1.6;       // a page can't stay up longer than this
const PAGE_HOLD_SEC = 0.4;      // linger after the last word ends

Remotion has a ready-made helper for this, createTikTokStyleCaptions in @remotion/captions, and it's a good default. It groups by a time threshold only, though, and this design needed a hard cap on words per page, so the grouping is twenty lines of custom code in the props builder instead.

Each page stays on screen until the next one replaces it, plus at most 0.4 s. There's deliberately no minimum duration: when aggressive jump cuts squash several words onto the same frame, forcing a page up to a minimum length makes it overlap the next one, and two caption lines stack on top of each other mid-screen. A page that would last under two frames is dropped and counted, and the builder warns if it drops any.

Rendering a word that "lights up"

The subtitle component draws each word in a page as an inline-block with an absolutely positioned rounded box behind it. The box's visibility is computed from the current frame relative to the word's own start and end:

const CLAMP = { extrapolateLeft: 'clamp', extrapolateRight: 'clamp' } as const;

const rise = interpolate(frame, [token.fromFrame - 2, token.fromFrame + 2], [0, 1], CLAMP);
const fall = interpolate(frame, [token.toFrame - 1, token.toFrame + 4], [1, 0], CLAMP);
const box = Math.min(rise, fall); // 0 = idle, 1 = currently spoken

Two separate ramps combined with Math.min look roundabout, but there's a reason. The obvious version — a single four-point interpolate from fade-in to fade-out — crashes on very short words, because the builder only guarantees a word lasts two frames, and interpolate requires a strictly increasing input range. Two independent ramps can't violate that.

Everything else is driven by that one box value:

// A spring that overshoots then settles reads as a "pop", not a twitch.
const pop = spring({ frame: frame - token.fromFrame, fps, config: { damping: 10, mass: 0.5, stiffness: 210 } });
const scale = 1 + 0.13 * box * pop;

// Upcoming words are dimmed so the eye can read ahead — but never below 0.72.
const opacity = interpolate(frame, [token.fromFrame - 5, token.fromFrame], [0.72, 1], CLAMP);

// Text goes from white (or the highlight color) to near-black as the box appears…
const color = interpolateColors(box, [0, 1], [idle, COLORS.ink]);
// …and the outline melts into the box color, or dark text with a dark outline
// on a yellow box turns into a blob.
const strokeColor = interpolateColors(box, [0, 1], [COLORS.ink, preset.active]);

The 0.72 floor is a lesson from testing on that deliberately ugly --busy background. The first version dimmed upcoming words to 0.45. On flat test backgrounds it looked elegant. On real, busy footage the dimmed words vanished, because CSS opacity fades the black outline along with the fill — and the outline is the only thing separating white text from a bright background. The word animation also originally toggled scale(1.12) on and off with no easing; on screen that read as a twitch rather than a bounce, which is why every transition now goes through spring or interpolate.

A page as a whole enters with a slight rise and scale — and deliberately without a fade. Pages butt up against each other with no gap, and a fade would leave a few frames of half-transparent text at every join, which reads as flicker.

Making Vietnamese text survive a busy background

Legible captions over real footage take three layers, and all three are needed:

export const stroke = (fontSize: number, color = COLORS.ink) => ({
  WebkitTextStroke: `${Math.round(fontSize * 0.09)}px ${color}`,
  paintOrder: 'stroke fill' as const,
});
  1. An outline at 9% of the font size. Thicker, and at weight 800 the stacked diacritics start fusing into the letter; thinner, and it doesn't separate text from real footage.
  2. paint-order: stroke fill — the line that matters most for Vietnamese. By default -webkit-text-stroke paints the outline on top of the glyph, eating half its thickness. Latin text survives that. Vietnamese doesn't: the stacked marks in ế, ộ, ữ, ỹ are thin, and an outline drawn over them clogs them into black lumps. Painting the stroke first and the fill on top keeps every mark crisp.
  3. A layered drop shadow, whose opacity is driven to zero while a word's highlight box is visible — a dark halo on a yellow box looks dirty.

The font is Be Vietnam Pro, designed specifically for Vietnamese, loaded with its Vietnamese subset explicitly:

import { loadFont } from '@remotion/google-fonts/BeVietnamPro';

export const { fontFamily } = loadFont('normal', {
  weights: ['600', '800'],
  subsets: ['vietnamese', 'latin'], // without 'vietnamese', accents break or fall back
});

Leave out the vietnamese subset and the accented characters either render as boxes or fall back to a system font mid-word. Several popular display fonts lack some Vietnamese glyphs at some weights; test any substitute with a string like ắ ầ ẫ ễ ộ ữ ỹ ặ ọ Đ đ before committing to it.

Where text goes on a vertical frame

A 1080×1920 frame isn't all yours. The platform's own interface covers roughly the top 150 px and the bottom 220 px, and the speaker's face occupies the middle. So there are two text layers, kept strictly apart:

Block titleAnimated subtitles
Text fromscene.caption.text (written for the screen)scene.voiceover (what's said)
What it isA headline for the whole block, readable with sound offThe speech itself, 2–3 words at a time
Position~22–24% of the height~74% of the height
On screenThe first 2.6 s of each sceneThe whole scene, following speech

Subtitle size also scales with the scene's role: the quote in the punchline is the biggest (76 px), the storytelling scenes the smallest (58 px) to leave room for the face, and the call-to-action switches its highlight box from yellow to green. The caption isn't a constant overlay; it has a rhythm that follows the script's structure.

The pattern, generalized

The technique here applies to any time you have a reliable text and an unreliable recognizer: karaoke lyrics against a vocal track, lecture slides against a recording, a known transcript against a noisy audio file.

  • Let each source contribute only what it's reliable at.
  • Compare on a normalized form; display the original.
  • Use a global alignment that tolerates insertions, deletions and substitutions together.
  • Fill the gaps by interpolation and enforce the invariants the consumer needs.
  • When quality is too low, degrade in the dimension users notice least.

Word timing, of course, is measured on the original clip. Once silence is cut out of the middle of that clip, every timestamp after the cut is wrong. Part 3 covers jump cuts — and the one function that keeps these subtitles in sync after them.