Royalty-Free Music and Sound Effects, Synthesized in Plain JavaScript

Cover Image for Royalty-Free Music and Sound Effects, Synthesized in Plain JavaScript
Video AI5 min read

AI Video Studio with Claude Code + Remotion

🛠 Build it: 1. Setup & your first skill · 2. The script-writing skill · 3. The editing skill · 4. The local AI story skill · 5. The comedy script skill

🎬 Use it: Talking-head video · AI story video · Deadpan comedy video

🔬 Under the hood: Architecture · Subtitles · Jump cuts · Music & SFX · Local AI pipeline

Background music is where short-video automation usually stops being automatic. Every video needs a bed of music under the voice and a handful of sound effects on the beats — a record scratch when the mood turns, crickets in an awkward silence, a "pop" on a keyword. Downloaded tracks bring licensing questions, content-ID claims that can mute or demonetize a video, and a folder of files that has to travel with the project. Trending platform sounds can't be used in an automated render at all.

This pipeline takes a different route: every piece of music and every sound effect is synthesized by code. About 650 lines of dependency-free JavaScript produce four looping background tracks and fourteen sound effects, written straight to MP3. The output is deterministic, it's yours outright, and a script can ask for funky or record-scratch by name and know the file will exist.

Be clear about expectations up front: this is not going to fool anyone into thinking a band recorded it. It sounds like a competent, slightly retro synthesizer — "good enough" to sit at 10–15% volume under speech. For deadpan comedy, the slightly cheesy quality actually helps. If you want a real track for a particular video, drop an MP3 with the same name into the assets folder; the pipeline always prefers your file and never overwrites it.

This post walks through how it works, from writing a WAV file by hand to a lo-fi hip-hop loop with swing and vinyl crackle.

Audio is just an array of numbers

Digital audio is a sequence of samples: numbers between −1 and 1 describing where the speaker cone should be, 44,100 times per second. A sine wave at 440 Hz — the A above middle C — is one line of math:

const SR = 44100;
const TAU = Math.PI * 2;
const buf = (sec) => new Float32Array(Math.ceil(sec * SR));

const tone = buf(1);
for (let i = 0; i < tone.length; i++) {
  tone[i] = Math.sin(TAU * 440 * (i / SR)) * 0.5;
}

Every instrument below is a variation of that loop: a formula for the waveform, an envelope that shapes loudness over time (fast attack, exponential decay), and sometimes a filter. A song is many short buffers added into one long one at the right times:

/** Add `src` into `dst` starting at `atSec`, scaled by `gain`. */
const mix = (dst, src, atSec, gain = 1) => {
  const off = Math.round(atSec * SR);
  for (let i = 0; i < src.length; i++) {
    const j = off + i;
    if (j >= 0 && j < dst.length) dst[j] += src[i] * gain;
  }
};

Musical pitch uses MIDI note numbers — 60 is middle C, each step is a semitone, 69 is A440:

const midi = (n) => 440 * 2 ** ((n - 69) / 12);

Writing the file without a library

A WAV file is a 44-byte header followed by raw 16-bit samples. Writing one by hand takes a dozen lines, and it means the synthesizer needs nothing but Node:

const writeWav = (path, samples) => {
  const data = Buffer.alloc(samples.length * 2);
  for (let i = 0; i < samples.length; i++) {
    const s = Math.max(-1, Math.min(1, samples[i]));     // clip to the valid range
    data.writeInt16LE(Math.round(s * 32767), i * 2);
  }
  const h = Buffer.alloc(44);
  h.write('RIFF', 0); h.writeUInt32LE(36 + data.length, 4); h.write('WAVE', 8);
  h.write('fmt ', 12); h.writeUInt32LE(16, 16);
  h.writeUInt16LE(1, 20);                 // PCM
  h.writeUInt16LE(1, 22);                 // mono
  h.writeUInt32LE(SR, 24); h.writeUInt32LE(SR * 2, 28);
  h.writeUInt16LE(2, 32); h.writeUInt16LE(16, 34);
  h.write('data', 36); h.writeUInt32LE(data.length, 40);
  writeFileSync(path, Buffer.concat([h, data]));
};

FFmpeg — already a dependency of the pipeline — then converts it to a 192 kbps stereo MP3, the format Remotion's Audio component plays.

Randomness you can reproduce

Drums, cymbals and many effects are built from noise. If the noise came from Math.random(), every regeneration would sound slightly different, and "the snare sounded better yesterday" would be unreproducible. A tiny seeded generator (xorshift) makes the same name always produce the same file:

const rng = (seed) => {
  let s = seed >>> 0 || 1;
  return () => {
    s ^= s << 13;
    s ^= s >>> 17;
    s ^= s << 5;
    return ((s >>> 0) / 4294967296) * 2 - 1; // uniform in [-1, 1)
  };
};

Each instrument call takes its own seed, so two hi-hats on different beats get different noise while staying identical from run to run.

A filter for everything: the biquad

Most of the character of these sounds comes from one component: a second-order IIR filter, the biquad, using the well-known coefficient formulas from Robert Bristow-Johnson's Audio EQ Cookbook. A low-pass darkens a sound, a high-pass thins it, a band-pass isolates a region — turn white noise into a snare, a hi-hat or a shaker by choosing which band to keep.

/** RBJ biquad. Returns a per-sample function; `.set(freq)` retunes it live. */
const biquad = (type, freq, q = 0.707) => {
  let b0, b1, b2, a1, a2;
  const set = (f) => {
    const w = (TAU * Math.min(f, SR * 0.45)) / SR;
    const cs = Math.cos(w);
    const al = Math.sin(w) / (2 * q);
    const a0 = 1 + al;
    if (type === 'lowpass')       { b0 = (1 - cs) / 2 / a0; b1 = (1 - cs) / a0; b2 = b0; }
    else if (type === 'highpass') { b0 = (1 + cs) / 2 / a0; b1 = -(1 + cs) / a0; b2 = b0; }
    else /* bandpass */           { b0 = al / a0; b1 = 0; b2 = -al / a0; }
    a1 = (-2 * cs) / a0;
    a2 = (1 - al) / a0;
  };
  set(freq);
  let x1 = 0, x2 = 0, y1 = 0, y2 = 0;
  const run = (x) => {
    const y = b0 * x + b1 * x1 + b2 * x2 - a1 * y1 - a2 * y2;
    x2 = x1; x1 = x; y2 = y1; y1 = y;
    return y;
  };
  run.set = set;
  return run;
};

The .set() method matters: changing the cutoff on every sample is how you get a filter sweep, the "wow" in a funk bass or the scrubbing sound of a record scratch.

The instruments

Each instrument is a function that returns a buffer for one note. Here are the techniques behind them — each a classic of sound synthesis, each only a few lines.

Electric piano: FM synthesis

The warm, bell-like Rhodes sound that defines lo-fi music comes from frequency modulation: one sine wave wobbles the phase of another at the same frequency. The strength of that wobble (the modulation index) starts high and decays quickly, so each note begins bright and mellows — exactly how a struck tine behaves:

const rhodes = (freq, dur, vel = 0.5, detune = 0) => {
  const f = freq * (1 + detune);
  const b = buf(dur + 0.6);
  for (let i = 0; i < b.length; i++) {
    const t = i / SR;
    const env = Math.min(1, t / 0.004) * Math.exp(-t * 1.6)
              * (t > dur ? Math.exp(-(t - dur) * 9) : 1);     // key release
    const idx = 0.25 + 1.4 * Math.exp(-t * 6);                // brightness decays
    const tine = 0.12 * Math.sin(TAU * f * 7 * t) * Math.exp(-t * 18); // metallic attack
    const trem = 1 + 0.08 * Math.sin(TAU * 4.5 * t);          // gentle tremolo
    b[i] = (Math.sin(TAU * f * t + idx * Math.sin(TAU * f * t)) + tine) * env * vel * trem;
  }
  return b;
};

Guitar and pizzicato: Karplus–Strong

The Karplus–Strong algorithm is one of the most satisfying tricks in audio. Fill a short ring buffer with random noise — its length sets the pitch — then loop through it, replacing each sample with the average of itself and its neighbor. The averaging is a low-pass filter applied once per cycle, so high frequencies die faster than low ones, and the noise turns into a convincingly plucked string within milliseconds:

const pluck = (freq, dur, vel = 0.5, bright = 0.5, seed = 1) => {
  const r = rng(seed);
  const n = Math.max(2, Math.round(SR / freq));  // period in samples = pitch
  const ring = new Float32Array(n);
  for (let i = 0; i < n; i++) ring[i] = r();     // the "pluck": a burst of noise
  const b = buf(dur + 0.3);
  const damp = 0.5 - 0.004 * (1 - bright);
  let p = 0;
  for (let i = 0; i < b.length; i++) {
    const nxt = (p + 1) % n;
    const v = ring[p];
    ring[p] = (v + ring[nxt]) * damp * 0.998;    // average neighbors → string decay
    p = nxt;
    b[i] = v * vel;
  }
  return filterBuf(b, 'lowpass', 1200 + 5000 * bright);
};

Bells: FM with a non-integer ratio

Use FM again, but modulate at 3.5× the note's frequency instead of 1×, and the overtones no longer line up with a harmonic series. The ear hears that inharmonicity as metal. That's the waltz's melody and the "ding" sound effect.

Funk bass: a resonant filter with its own envelope

A sawtooth blended with a square wave, run through a low-pass filter with high resonance (Q = 4), whose cutoff starts at about 2,400 Hz and falls to 180 Hz within a few tens of milliseconds. That falling, resonant cutoff is the "bwow" of a synth bass. A tanh at the end soft-clips it for warmth:

lp.set(180 + 2200 * Math.exp(-t * 22));          // cutoff envelope
b[i] = Math.tanh(lp(0.6 * saw + 0.4 * sq) * 1.5) * env * vel;

Drums from sine waves and noise

  • Kick: a sine wave whose pitch sweeps quickly from 120 Hz down to 45 Hz, with a 3 ms click on top for the beater. The sweep is what makes it a "thump" instead of a "boop".
  • Snare: band-passed noise around 1,900 Hz (the rattle of the wires) plus a short 185 Hz tone (the drum head).
  • Hi-hat: white noise through a 7 kHz high-pass with a very fast decay.
  • Vinyl crackle: near-silent noise with an occasional random spike, low-passed at 3 kHz so the pops sound dusty rather than digital — roughly one and a half pops per second.

Composing a loop: lo-fi at 76 BPM

With instruments in hand, a track is a set of nested loops over bars and beats. The lo-fi track plays a ii–V–I–vi progression in C (Dm9, G13, Cmaj9, Am9) at 76 BPM, with details that make it feel played rather than programmed:

lofi: () => {
  const bpm = 76;
  const beat = 60 / bpm;
  const bars = 16;
  const loop = bars * 4 * beat;
  const out = buf(loop + 2);           // extra room for ringing tails
  const prog = [
    [50, 53, 57, 60, 64],              // Dm9
    [43, 53, 59, 64],                  // G13
    [48, 52, 55, 59, 62],              // Cmaj9
    [45, 55, 60, 64, 71],              // Am9
  ];
  const r = rng(11);
  for (let bar = 0; bar < bars; bar++) {
    const t0 = bar * 4 * beat;
    const ch = prog[bar % 4];
    // Roll the chord like a human hand (18 ms apart), each note detuned a hair
    // like a worn tape machine.
    ch.slice(1).forEach((n, k) =>
      mix(out, rhodes(midi(n), beat * 2.6, 0.26, r() * 0.003), t0 + k * 0.018));
    mix(out, sineBass(midi(ch[0] - 12 < 28 ? ch[0] : ch[0] - 12), beat * 3.2, 0.42), t0);
    // Boom-bap: kick on 1 and the "and" of 3, snare on 2 and 4.
    mix(out, kick(0.6), t0);
    mix(out, kick(0.45), t0 + beat * 2.5);
    mix(out, snare(0.28), t0 + beat);
    mix(out, snare(0.28), t0 + beat * 3);
    // Swung eighth-note hats: every off-beat lands 60% of the way through the beat.
    for (let k = 0; k < 8; k++) {
      const sw = k % 2 ? beat * 0.6 : 0;
      mix(out, hat(k % 2 ? 0.07 : 0.11, 55, 20 + k), t0 + Math.floor(k / 2) * beat + sw);
    }
  }
  let mixd = foldLoop(out, loop);
  mixd = filterBuf(mixd, 'lowpass', 3400, 0.6);   // the muffled "lo-fi" top end
  mix(mixd, vinyl(loop), 0, 1);
  return normalizeRms(mixd, -19);
},

The three other tracks follow the same template with different ingredients:

TrackTempoIngredientsUsed for
lofi76 BPMRhodes chords, sine bass, swung boom-bap, vinylThoughtful openings
funky100 BPMResonant synth bass on sixteenths, clav stabs on off-beatsThe "serious explanation" of something absurd
waltz156 BPM, 3/4Plucked bass, organ "oom-pah-pah", bell melodyMock-solemn, documentary-style
romantic70 BPMArpeggiated plucked guitar, thin organ padLove-themed topics

Seamless loops: folding the tail

A video can be longer than a 16-bar loop, so the track has to repeat without a seam. The problem: the notes in the last bar ring past the loop boundary (that's what the extra two seconds in buf(loop + 2) are for). Cut them off and you hear the decay vanish at every repeat. Drop them and the first bar sounds dry compared to the rest.

The fix is to fold the overhang back onto the start, exactly as it would sound if the loop really had been playing before:

/** Wrap anything past `loopSec` (ringing notes, reverb) back onto the beginning. */
const foldLoop = (b, loopSec) => {
  const n = Math.round(loopSec * SR);
  const out = new Float32Array(n);
  for (let i = 0; i < b.length; i++) out[i % n] += b[i];
  return out;
};

Remotion's <Audio loop> then repeats the file indefinitely with no audible join.

Loudness: RMS, then a soft limiter

Adding dozens of notes together produces arbitrary levels. Each track is normalized to a target RMS (about −19 dBFS for the gentle tracks) so they sit at comparable loudness, and passed through tanh so the occasional peak bends instead of clipping:

const normalizeRms = (b, db = -18) => {
  let s = 0;
  for (const v of b) s += v * v;
  const rms = Math.sqrt(s / b.length) || 1;
  const g = 10 ** (db / 20) / rms;
  for (let i = 0; i < b.length; i++) b[i] = Math.tanh(b[i] * g * 1.2) / 1.2;
  return b;
};

Sound effects use peak normalization instead, since they're short events where the transient is the point.

Sound effects as tiny programs

Each effect is a short recipe. A few favorites:

Record scratch. Noise and a sawtooth through a band-pass filter whose center frequency swings back and forth — three pushes and pulls, each faster than the last — then stops dead, with a 20 ms fade so the stop doesn't click:

const sweep = Math.abs(Math.sin(TAU * (t < 0.3 ? 3.3 : 6) * t));
const f = 150 + 1900 * sweep;
bp.set(f);

Crickets. A 4.4 kHz tone chopped into pulses by a squared 32 Hz sine, repeated every 0.62 s, plus a second, quieter cricket at 4.95 kHz every 0.81 s. Because the two periods don't divide evenly, the pattern never quite repeats — which is what makes it sound like an actual empty field rather than a loop.

Thud. A falling low sine (the body of a glass set down on a table) plus a very short burst of band-passed noise around 900 Hz (the wood).

Boing. A sine whose pitch rises while a 13 Hz wobble on top decays — the sound of a spring settling.

The full set, available by name: record-scratch, crickets, thud, tick-tock, heartbeat, boing, ding, notification, pop, swoosh, riser, boom, sigh, typing.

Wiring it into the pipeline

The synthesizer is a module with two exported functions the props builder calls for every name a script mentions:

/** Returns the mp3 path, generating it if missing.
 *  null if the name can't be synthesized (and the user hasn't supplied a file). */
export const ensureMusic = (name, force = false) => {
  const p = join(assetsDir(), 'bgm', `${name}.mp3`);
  if (existsSync(p) && !force) return p;   // the user's own file always wins
  return render('bgm', name, force);
};

Three behaviors follow from that small function:

  • Generated on demand. A fresh checkout has no audio files at all. The first video that asks for lofi creates content/_assets/bgm/lofi.mp3; later videos reuse it.
  • Your file wins. Copy a real track to content/_assets/bgm/lofi.mp3 and it's used as-is. The synthesizer never overwrites an existing file unless you pass --force.
  • Unknown names don't block the render. If a script asks for a sound that can't be synthesized and isn't on disk, the video still renders without it, and the build prints which file to add.

It also works as a command-line tool:

node audio-synth.mjs --list             # every available name
node audio-synth.mjs --all              # generate everything that's missing
node audio-synth.mjs lofi --force       # regenerate one track

Music that changes with the scene

The comedy format specifies music per scene: lo-fi for the setup, dead silence for the twist, funk for the absurd explanation. The props builder walks the scenes, merges consecutive scenes that use the same track into one span, and drops the silent spans:

for (const span of sceneSpans) {
  const name = span.music ?? 'silent';
  const last = music[music.length - 1];
  if (last && last.name === name && last.to === span.from) last.to = span.to; // extend
  else music.push({ name, from: span.from, to: span.to });
}

Merging matters: without it, a track would restart from bar one at every scene boundary. In Remotion, each span becomes its own Sequence with a looping Audio, and the volume is a function of the frame so each span fades in and out over a quarter second — enough that a change of track doesn't "clunk":

const MusicTrack: React.FC<{ seg: MusicSegment }> = ({ seg }) => {
  const fade = Math.max(1, Math.min(seg.fadeFrames, Math.floor(seg.durationInFrames / 3)));
  return (
    <Audio
      src={staticFile(seg.src)}
      loop
      volume={(f) =>
        seg.volume *
        interpolate(f, [0, fade, seg.durationInFrames - fade, seg.durationInFrames], [0, 1, 1, 0], {
          extrapolateLeft: 'clamp',
          extrapolateRight: 'clamp',
        })
      }
    />
  );
};

The fade is capped at a third of the span's length, so a very short scene still has a moment at full volume. One thing the fade deliberately doesn't do: soften the cut into the twist. The script's "music stops dead" is a 0.25 s fade-out followed by true silence — close enough to an abrupt stop to land the beat, smooth enough not to click.

Under speech, music plays quietly: 0.10–0.15 of full volume for selfie videos and 0.13 for comedy, where the slow delivery needs more room. The voice track comes from the footage itself, normalized to −14 LUFS during audio extraction, so the balance is consistent from video to video.

Takeaways

  • Synthesis is a licensing strategy. Code-generated audio has no rights holder but you.
  • A handful of classic techniques — FM, Karplus–Strong, a biquad filter, filtered noise — covers an entire small band and a sound-effects library.
  • Seed your randomness so audio assets are reproducible.
  • Fold loop tails back to the start for seamless repetition.
  • Generate assets on demand, let user-supplied files take priority, and never block a render on a missing sound.

So far every video has needed a person on camera. Part 5 removes the camera entirely: a narrated, illustrated story video where the voice and every image are generated on your own Mac, with no API keys and no per-video cost.