# Auto captions in DaVinci Resolve Free with AI

Make editable captions for DaVinci Resolve Free with an AI assistant. Follow the transcription, timing, SRT import, and export steps with downloadable example files.

Author: Anvisha Pai

Canonical: https://voyageragent.ai/blog/davinci-resolve-free-auto-captions-codex-fal

Published: 2026-09-17
Updated: 2026-09-17

DaVinci Resolve Free can import and edit subtitles, but its built-in **Create Subtitles from Audio** command was disabled in the Free 21.1 installation used for this guide. You can supply that missing transcription step with an AI assistant, then finish the captions in Resolve.

The workflow is straightforward: **give the assistant your video, have it create a timed subtitle file, and import that file into your edit**. The result is a subtitle track you can correct and style, plus captions you can burn into the exported video or save separately.

We'll use two tools alongside Resolve:

* **Codex** is OpenAI's coding assistant. Here, it handles the practical work: uploading the recording, calling the transcription service, processing the response, and writing the subtitle file. You describe the job in a prompt rather than writing the processing code yourself.
* **fal** is a service that hosts AI models, including speech-to-text models. Codex connects to it through **MCP**, a standard that lets an assistant call another service's tools. Think of MCP as the connection; fal does the speech recognition.

The file passed back to Resolve is an **SRT**: a plain-text file containing numbered captions, each with a start time and end time. It keeps the words and their timing editable.

This guide follows a real run on a 40-second synthetic English recording. The interesting part was not getting the words back. It was turning two long, overlapping transcript segments into **14 readable captions** without guessing when the words were spoken.

## What you'll need

* DaVinci Resolve Free and a video ready to caption.
* A local Codex session that can work with files on your computer.
* A fal account, an API key, and billing available for model calls.

“Resolve Free” describes the editor edition. Codex access and fal processing have their own costs. The instructions below ask the assistant to check pricing before submitting a job. This route uploads the recording to fal; use local transcription instead if the footage must remain on your machine.

To follow the exact example, download the [source video](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/caption-test-d615f07867.mp4) and [reviewed SRT](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/captions-6aa5b49fe8.srt). The [sample bundle](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/caption-tutorial-files-f112ec5424.zip) also includes the transcript, word timings, first-pass captions, a Python grouping script, and the editable Resolve project. Relink the bundled source video if the imported project shows missing media. You can inspect those files without paying to transcribe the clip again.

If you're deciding whether to upgrade for other editing features, see [DaVinci Resolve Free vs Studio](/blog/davinci-resolve-free-vs-studio). This tutorial only needs the Free edition's subtitle-import workflow.

## 1. Connect your AI assistant to the transcription service

First, open a terminal in the folder where you'll work on the video. This setup uses the Codex command-line app. If you haven't installed it, follow [OpenAI's Codex setup guide](https://developers.openai.com/codex/quickstart) first.

Create an API key in your [fal dashboard](https://fal.ai/dashboard/keys). An API key lets the assistant use your fal account; it is not something to paste into the conversation. Make it available as `FAL_KEY`, a setting the terminal passes to Codex. In macOS's default zsh shell, these commands ask for the key without displaying it or putting its value in your shell history:

```zsh
read -rs "FAL_KEY?Paste your fal API key: "
export FAL_KEY
```

Press Enter after pasting the key. Next, register the connection:

```bash
codex mcp add fal-ai \
  --url https://mcp.fal.ai/mcp \
  --bearer-token-env-var FAL_KEY
```

Start a new Codex session from the same environment. Check `codex mcp list` and the session's `/mcp` view to confirm the connection. The registration command refers to the environment variable rather than embedding its value in the command. Keep the key out of prompts and project files. A desktop app launched separately may not inherit the terminal's environment.

The command options are documented in [OpenAI's MCP guide](https://developers.openai.com/codex/mcp); fal documents the [hosted MCP endpoint](https://fal.ai/docs/documentation/setting-up/mcp).

fal's tools handle model discovery, schema inspection, pricing, uploads, and execution. They do **not** add native transcription to Resolve Free or install a Resolve bridge. If you want the agent to operate the editor too, that is a separate connection with its own edition and control limits; see [DaVinci Resolve MCP](/blog/davinci-resolve-mcp).

## 2. Ask the assistant to create editable captions

Put the finished cut in a dedicated folder. Transcribing the final cut avoids having to retime every caption after removing a scene. If you have an existing Resolve timeline, export a reference of that cut first. A speech-only audio export can reduce the upload size; use the same edit and starting point as the picture.

![DaVinci Resolve Free with the 40-second test video and its audio on a timeline starting at zero.](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/source-timeline-2be34fd4fb.webp)

Start with the cut you actually want to caption. This source timeline begins at 00:00:00:00 and has no subtitle track yet.

Here is a brief to adapt. Replace the file name, language, and description of the footage, and give the agent a spending limit appropriate to your clip:

```text
Make editable SRT captions for caption-test.mp4.

I use DaVinci Resolve Free. Use the connected fal MCP server
for speech recognition. Inspect the current model schema and
pricing first. Tell me the estimate before submitting paid jobs.

I authorize uploading this clip to fal for this task. It contains
synthetic English narration. Transcribe; do not translate or rewrite.

Save the original model response as transcript.json. I need real
start/end timestamps, not timings estimated from the word count.
If the result has only long segments, explain the problem and propose
an alignment step before making short captions.

Create captions.srt in UTF-8. Aim for one or two readable lines,
with no more than 42 characters per line where practical. Keep
names, dates and numbers together. Flag missing or overlapping
timestamps and unusually fast captions for review.

Leave my video unchanged. Give me the SRT, the model IDs used,
and a short list of corrections to inspect in Resolve.
```

The upload matters: this fal route sends the recording to a hosted service. For footage that must remain on your machine, use a local transcription route instead. That changes the transcription step, not the SRT handoff.

For a local file, the hosted MCP server cannot read your disk path. Its `upload_file` tool can prepare a signed upload URL; the agent then sends the actual bytes using its local HTTP tools. Only a successful upload makes the returned file URL usable as model input. An “upload prepared” response is not a completed upload.

## 3. Check the transcript and its timestamps

Ask the assistant which model it selected and whether the response includes timestamps. Model names and supported options can change, so this is worth checking before processing a long recording.

Our successful transcription used **`fal-ai/wizper`**, fal's Whisper v3 endpoint. An earlier lookup for `fal-ai/whisper` reported that older endpoint unavailable to the connected account. The assistant discovered the alternative and checked its input requirements before running it.

The submitted settings were:

```json
{
  "task": "transcribe",
  "language": "en",
  "chunk_level": "segment",
  "merge_chunks": false
}
```

The request also included the uploaded clip's `audio_url`. At the time of this run, the [Wizper schema](https://fal.ai/models/fal-ai/wizper/api) fixed `chunk_level` to `segment`. Asking for word timestamps in a prompt could not change that API constraint.

The returned text included the names and numbers we needed: “DaVinci Resolve,” “Voyager,” “24 frames per second,” and “9:30 on September 17th.” But its timing was unsuitable for short on-screen captions:

| Returned segment                 | Start    | End      | Problem                                 |
| -------------------------------- | -------- | -------- | --------------------------------------- |
| First paragraph of transcription | 0.000 s  | 25.940 s | Too much text for one readable cue      |
| Remaining transcription          | 25.740 s | 37.709 s | Starts before the previous segment ends |

Those values are from this one request. They are not a claim that every Wizper response behaves the same way. They do show why saving the raw response and checking its structure is part of the job.

Do not have the agent divide a 26-second paragraph into equal time slices. People pause, accelerate, and hold words for different lengths. A plausible-looking SRT can still be out of sync.

## 4. Get word timings and choose readable caption breaks

If your transcription already has usable word timings, you can move straight to grouping the captions. Our response did not, so it needed **forced alignment**: matching the supplied transcript back to the audio to find when each word begins and ends.

For this sample, the next call used **`fal-ai/elevenlabs/forced-alignment`** through the same fal MCP connection. It received the same uploaded recording and the returned transcript, and produced word-level timings. Its [API documentation](https://fal.ai/models/fal-ai/elevenlabs/forced-alignment/api) describes the `audio_url` and `text` inputs and the aligned output.

Forced alignment answers “where do these supplied words occur?” It does not prove that the supplied transcript is correct. Correct recognition mistakes before alignment where possible, and review the aligned result against the recording afterward.

Ask Codex to turn those timings into captions:

```text
Use the aligned words to create caption cues. Preserve the word
start/end times. Ignore whitespace-only alignment entries.

First check that every timestamp is finite, non-negative and ordered.
Do not silently fill in missing times or allow cues to overlap.

Group at sentence or phrase boundaries. Start with a four-second
maximum cue duration and two lines of about 42 characters each,
but report exceptions instead of breaking a person's name or number.

Compare the result to the recording. Keep “DaVinci Resolve” in one
cue and “September 17th” on one line. Save both the first pass and
the reviewed captions.srt so I can inspect the changes.
```

Our first grouping pass produced 12 cues. It met the basic duration and line-count checks but ended one cue with “DaVinci” and began the next with “Resolve.” The reviewed file has **14 cues**, with that phrase kept together. This was a caption-grouping error, not a speech-recognition error.

The reviewed date cue looks like this:

```srt
9
00:00:18,680 --> 00:00:21,860
The meeting starts at 9:30
on September 17th.
```

The line break is a readability choice. The cue's start and end come from the alignment. Do not confuse word-level timing with word-by-word animation: this file contains conventional subtitle cues, not animated karaoke text.

## 5. Bring the SRT into Resolve Free

Open the project containing the matching picture edit. On the Edit page:

1. Choose **File → Import → Subtitle** and open `captions.srt`.
2. Find the subtitle asset in the Media Pool.
3. Drag it onto a subtitle track above the video, aligned with the start of the recording you transcribed.
4. Select a subtitle cue to review its text and timing in the Inspector. Use the track's style controls for a consistent appearance.

Importing the file into the Media Pool is only the first step. It does not itself place the captions on your timeline. Plan to do the drag yourself: the fal connection produces files, not control of Resolve. In our session, the agent imported the file through the UI and a person placed it on the subtitle track.

![Resolve Free showing 14 subtitle clips above the video, with the date caption selected and its text and timecodes open in the Inspector.](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/caption-inspector-ad1ade71be.webp)

The reviewed date caption in Resolve Free. Each block on ST1 is editable; the Inspector exposes its text, start, end, and reading speed.

Our test timeline starts at `00:00:00:00`. If yours starts at `01:00:00:00`, place the SRT relative to the beginning of the corresponding media, and inspect the first spoken sentence. Do not add an hour to the SRT automatically. The correct placement depends on the reference you transcribed and how you insert it.

If the subtitles start correctly but drift later, check whether the transcription came from a different cut or a retimed version. If everything is shifted by the same amount, first check the placement of the subtitle asset.

## 6. Review and export the right deliverable

Before rendering, check the beginning, middle, and end against the audio. Look at short cues, numbers, names, pauses, and places where you changed the cut. Then watch a section muted: the captions should remain understandable without racing through unfinished phrases.

Resolve's Deliver page separates subtitle output from video output. Under the video render settings, expand **Subtitle Settings** and enable subtitle export. Choose the output that matches the destination:

| Output                                    | Use it when                                                                  |
| ----------------------------------------- | ---------------------------------------------------------------------------- |
| **Burn into video**                       | The text must always be visible, including in a social upload                |
| **As a separate file**, SRT               | You need an editable subtitle file or an uploadable caption track            |
| **As embedded captions**, where supported | Your delivery format and player support the required embedded caption format |

![DaVinci Resolve Deliver page with Export Subtitle enabled and its Format set to Burn into video.](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/burn-in-settings-11d7a448fd.webp)

For captions that stay visible in the picture, enable Export Subtitle and choose Burn into video. The screenshot shows the QuickTime H.264 settings used for our native export.

These are different deliverables. An SRT carries text and timing; it does not reliably preserve Resolve's visual styling. Burned captions preserve the rendered appearance but cannot be switched off or edited as text in the exported picture. Blackmagic's [Resolve reference manual](https://documents.blackmagicdesign.com/UserManuals/DaVinciResolveReferenceManual.pdf) documents subtitle tracks and delivery options.

Resolve exported all 14 reviewed cues back to SRT with their words and line breaks intact. The exported timestamps differed from the alignment-based file by at most 20 milliseconds, consistent with rounding to this 24 fps timeline. Resolve also added bold tags; an SRT is not always byte-for-byte identical after an editor round trip.

Keep the SRT and project even when the requested delivery is a burned-in MP4. Reopen the actual exported file and check that the captions are present. Seeing them in the Resolve viewer is not proof that the export included them.

Here is the captioned result from Resolve. The web copy below keeps Resolve's H.264 picture unchanged and converts its audio to AAC for MP4 playback.

![Video poster](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/resolve-render-check-02edf96c9d.webp)

[Play video](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/resolve-captioned-a13b2a8db0.mp4)

[Download the captioned MP4](https://cdn.moda.app/voyager/blog/davinci-resolve-free-auto-captions-codex-fal/resolve-captioned-a13b2a8db0.mp4).

## Keep the editable files

Codex coordinates the files and checks. fal supplies transcription and alignment. Resolve remains the place to review the edit, choose the subtitle appearance, and deliver the video. A fal MCP connection alone does not unlock Studio features or make Resolve's scripting API available in the Free edition.

This example is deliberately small: one synthetic voice, one 40-second source, and ordinary subtitles. It demonstrates the file handoff and the need to review model timing; it cannot establish accuracy on a podcast, speaker labeling, multilingual performance, or a cost advantage over Studio.

If you need animated text or a repeated captioning workflow across many projects, compare that requirement separately in [AI plugins for DaVinci Resolve](/blog/ai-plugins-for-davinci-resolve). For a single Free-edition edit, an inspected SRT is a practical starting point: the words remain editable, the timing is visible, and you can finish the job inside your existing project.

Test setup and limitations

The editor was DaVinci Resolve Free 21.1.0, build 17, on macOS. The source was a 1920×1080, 24 fps video with one synthetic English voice. It was reused from an earlier local-transcription check, which explains the “local transcription” label in the picture; the fal calls used that same recording.

Codex executed requests through fal's hosted MCP endpoint using a temporary HTTP client and an environment-provided credential. The reader's registration command was checked against the installed Codex CLI's help; it was not added to this machine's persistent configuration. A person placed the initial SRT on the track and selected the burn-in output mode. The agent applied the reviewed fal caption text and timings through the Inspector, exported the subtitle file, and started the native render. The SRT round trip was checked for every cue; a frame from the exported movie confirmed that captions were burned into the picture. Transcription and alignment responses are preserved in the downloadable bundle. This is a worked example, not a comparison with Studio's native transcription or a benchmark of model accuracy.


## Make your next project in Voyager

Create an account, download Voyager, and start making.

[Sign up](/sign-up)