Which AI video model can render text on screen
On-screen text is where AI video still breaks. We asked 19 models for the same animated title card, three words and a date over festival footage, and kept every result including the two refusals. Here is who can spell. The grid includes the September 18 additions; each model page dates its own run.
Anvisha Pai, Co-founder & CEO, Voyager
What the side-by-side shows
Five models set the card cleanly: Veo 3.1 and Veo 3.1 Fast typeset both lines and held them, Seedance 2.0 animated the words in and kept them intact, MiniMax H3 and Hailuo 2.3 were crisp from the cheapest end of the roster, and Gemini Omni Flash resolved to clean type after a bokeh opening. Kling is the clear failure in both tiers, producing 'MDDA SUMMER SESSIIONS' and 'MODA SUMER SOSSSIONS'; Sora 2 set thin type that wobbled and lost letters by the last frame; Veo 3.1 Lite overlapped its two lines and cropped one; Grok opened with a stray 'MME' before settling. Both Wan endpoints refused the prompt outright on fal's content checker, the only models to do so. Luma Ray 3.2, added later, shows the title faintly for about a second mid-clip and then loses it.
- 1.Veo 3.1 typeset both lines on a card and held them still over the crowd.
- 2.Seedance 2.0 animated the words in and kept every letter, with the date on its own line.
- 3.Veo 3.1 Fast the same clean type as the full model at a quarter of the price.
- 4.MiniMax H3 crisp lettering from a $0.30 clip.
- 5.Hailuo 2.3 correct text from the cheapest silent model on the roster.
Every video model on this prompt, 17 clips
One try each, nothing re-rolled or retouched, generated by us. Ranked models come first; the rest follow in family order. Open a model for the rest of its set.
›The prompt every model got
A title card reading 'MODA SUMMER SESSIONS' animates in over slow-motion footage of a festival crowd at dusk, then a smaller line 'Friday 12 June' fades in beneath it. Bold condensed type, tangerine on navy.
Set text that stays spelled correctly while it moves.
typeset both lines on a card and held them still over the crowd.
animated the words in and kept every letter, with the date on its own line.
the same clean type as the full model at a quarter of the price.
crisp lettering from a $0.30 clip.
correct text from the cheapest silent model on the roster.
The other prompts
- Which AI video model holds an art style
- Which AI video model handles a busy scene best
- Which AI video model handles a camera move best
- Which AI video model handles hands and objects best
- Which AI image-to-video model keeps the product intact
- Which AI video model gets physics right
- Which AI video model makes the best product turntable
- Which AI video model makes the best talking head
Frequently asked questions
Which AI video generator is best for text?
On our title card, Veo 3.1 and Veo 3.1 Fast, Seedance 2.0, MiniMax H3 and Hailuo 2.3 all rendered the words correctly. Kling 3.0, in both tiers, could not spell the title.
Can Kling render text?
Not reliably. Both Kling 3.0 Standard and Pro garbled the three-word title on this prompt. Add text afterwards in an editor instead.
Why did Wan refuse the prompt?
fal's content checker returned a policy error on the festival title card for both Wan 2.7 and Wan 3.0. No other vendor objected to the same prompt.
Voyager
Run any of these inside Voyager
Voyager is an agent for creative work that runs every model on this page. Brief it, and it picks the model for the job, makes the thing and hands back something you can edit. Create an account to get started.
