Blog/Techniques

Add sound to a silent clip without losing control of the timing

Try automatic audio and separate Foley prompts on one clip, download the stems, then fix one mistimed sound through a pinned review in Voyager.

John Holliman

John Holliman, Cofounder & CTO, Moda

Sep 29, 2026·4 min read

An illustrated pitcher pours coffee into a ceramic cup on a saucer.

A cup lands, coffee pours, a spoon taps the rim. Those three moments give you a simple sound-design brief for a silent animation, product reel or game cinematic.

You can ask a model to score the whole clip, or generate separate effects and place them yourself. We tried both on the same eight-second animation. The useful distinction is control: can you move one impact without rebuilding everything else?

Try the automatic version

MMAudio generates audio conditioned on video and text. Give it the silent clip and describe the scene:

Prompt
A small ceramic cup lands gently on a ceramic saucer, coffee pours into the cup, then a metal teaspoon lightly touches the ceramic rim. Quiet indoor room ambience. Natural synchronized sound effects.

We used fal-ai/mmaudio-v2, eight seconds, seed 420, with music, speech, voices, singing and loud bangs in the negative prompt. This is an illustrated test scene, not a claim about how the model handles every kind of footage.

Two complete eight-second takes. Playback is matched to −24 LUFS; the workflows use different inputs and amounts of direction.
Description / transcript

First, MMAudio generates the complete soundtrack from the silent clip and scene prompt. Then the same clip plays with separate cup, pouring, spoon and room tracks. Compare the sounds at the visible contact moments.

The player compares the complete automatic output with our final layered mix, at matched playback loudness. Listen for contact timing, extra sounds and whether the materials fit. One example cannot establish a model winner.

Give each sound its own track

For the layered version, we generated four files with ElevenLabs Sound Effects: three actions and a quiet room bed. You can use these prompts with a sound model and place the files in any editor.

Cup contact, two-second generation:

Prompt
One small ceramic espresso cup placed gently on a ceramic saucer. A single clean porcelain contact with short delicate resonance. Close microphone in a quiet kitchen. One contact only. No voices, music, pouring or additional impacts.

Pour, three-second generation:

Prompt
A gentle continuous pour of coffee from a small pitcher into a ceramic espresso cup, close microphone, thin liquid stream and soft splashing. Starts immediately and lasts two seconds, then stops cleanly. No voices, music or cup impacts.

Spoon, two-second generation:

Prompt
A single light metal teaspoon tap on the rim of a small ceramic espresso cup. One delicate high ceramic clink with a short natural ringing tail. No voices, music or repeated taps.

Room tone, eight-second generation:

Prompt
Very quiet small kitchen room tone, soft steady indoor air, subtle warm room ambience. No distinct events, appliances, voices, footsteps, birds or music.

Keep the room quieter than the actions. Trim unwanted repeats and leading silence before mixing. Waveform analysis found a 0.37-second lead-in in our spoon file, which we trimmed. Audition the onset too: align the clink, not merely the file’s start.

CueVisible timingWhat to align
Cup2.00 secondsFirst impact
Pour3.00–5.00 secondsStart and stop of the stream
Spoon6.20 secondsFirst clink
RoomWhole clipLow continuous bed, with fades

Download the editable project for the silent clip, original generations, prepared stems, cue sheet and mix recipes. Its HTML comparison opens locally after you unzip it. The animation source is included too.

A silent close-up of the cup settling onto its saucer.
Download GIF
Silent timing reference: the cup reaches the saucer at 2.00 seconds in the full clip. A GIF cannot demonstrate the sound.

In Voyager, point to the contact frame

Separate tracks help in any editor. Voyager adds a direct path from a comment on the video to the source recipe that controls the mix.

For this revision exercise, we deliberately placed the cup cue at 2.40 seconds, 12 frames late at 30 fps. In the actual desktop app, we paused at the cup’s contact frame, pinned a comment and sent it to the agent:

Prompt
Align the cup impact with this contact frame and lower its volume from 0.65 to 0.50. Save mix-v2.json and export layered-v2.mp4. Preserve the pouring, spoon and room entries, all generated source files, mix-v1.json and layered-v1.mp4. Do not generate new audio.
24-second edit of the real desktop run and exported output. Waits are removed; saved reviews are revisited and still holds are labeled.
Description / transcript

The initial cup effect is deliberately 12 frames late. A comment pinned at 2.00 seconds asks the agent to align the impact and lower its gain from 0.65 to 0.50. The agent updates the recipe, preserves the other tracks and exports a new video.

The agent moved the cup from 2.40 to 2.00 seconds and reduced its gain from 0.65 to 0.50. The other three recipe entries stayed identical; hashes confirmed that the stems and previous export were unchanged.

Under the hood, the review carries the video reference, timestamp and position. The agent edits the local FFmpeg recipe and renders a new export. The useful shortcut is that the visual note already identifies the moment; you do not need to prepare a separate screenshot and timecode handoff.

Tested September 29, 2026, in Voyager development build fffef4cec with Codex and GPT-6-Astra. Six generation requests, including one source revision, total $0.046 at the documented public fal rates; this is an estimate, not the Moda account charge. Run notes include exact settings and limitations. Review measured timing and levels and inspected visual frames; it did not include perceptual listening, so we make no sound-quality ranking.

Reusable workflow and prerequisites

Add sound to a silent clip and revise one cue while preserving the other tracks.

Tested 2026-09-29 · Voyager Development build fffef4cec

Before you start

  • A silent clip and access to the chosen audio models.
  • An editor that can place separate audio tracks, or Python 3 and FFmpeg for the download.
  • Voyager with pinned video review for the demonstrated revision.

Six model requests total $0.046 at the public fal rates retrieved September 29, 2026. This is an estimate, not a receipt for the Moda account charge. Coding-agent access is separate. The revision used local rendering only.

Starting prompt

A single light metal teaspoon tap on the rim of a small ceramic espresso cup. One delicate high ceramic clink with a short natural ringing tail. No voices, music or repeated taps.

  1. Generate the sounds: Use the inline prompts for separate cup, pouring, spoon and room files, or try the complete scene prompt with MMAudio.
  2. Align the actual attacks: Trim leading silence and align the impacts to the visible contacts. Keep each cue on its own track and the room bed quiet.
  3. Revise one cue: In Voyager, pause at the contact, pin the timing request, attach the review and send it. Preserve the previous export and unaffected stems.

Expected outputs

  • An eight-second video with sound.
  • Separate original effects, prepared stems, cue sheet and editable mix recipes.
  • A revised export with the cup aligned at 2.00 seconds.
Test method and limitations

Tested 2026-09-29. Generated both workflows on an original illustrated clip. Recorded an actual desktop pin and agent revision. Verified unchanged source hashes and other cue entries, decoded exports and measured loudness-matched playback.

One stylized scene and different input workflows do not establish a model benchmark. Creative review inspected frame sequences, not continuous motion or perceptual audio. Exact Moda account charges were not returned.

Make your next project in Voyager

Create an account, download Voyager, and start making.

John Holliman

John Holliman

Cofounder & CTO, Moda

John is the CTO of Moda and a Y Combinator-backed startup founder. Previously CTO of Dover, John builds creative tools that let people direct agents, inspect their work, and refine the result while staying in control.