Gemini 3.1 Flash TTS
Gemini 3.1 Flash TTS is Google's text-to-speech model in the Gemini line, released in preview on 15 April 2026 as the successor to the 2.5 Flash TTS preview. It takes a script plus style instructions written in plain language, inline tags such as [laughing], and a speakers list for two-voice dialogue, across 30 named voices and more than 80 languages. On fal it costs $0.05 per thousand characters; on Google's own API it is billed by the token, $1 per million text tokens in and $20 per million audio tokens out, at 25 tokens per second of audio.
Anvisha Pai, Co-founder & CEO, Voyager
Google DeepMind · released
›What we asked for
Hi, I'm Gemini 3.1 Flash TTS, Google's voice model. I take stage directions written in plain language, but today I'm reading the script as written. Here is the same script every voice on this site reads, so you can compare us.
fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 16.4 seconds of mp3
At a glance
- Price
- $0.05 per thousand characters on fal, half of Eleven v3; our four scripts cost about five cents.
- Voices
- Thirty named voices, Kore by default, which we used; no cloning.
- Direction
- A style_instructions field in plain language plus inline tags; our scripts used neither.
- Dialogue
- A speakers list assigns a voice per speaker for two-person scripts.
- Languages
- More than 80, auto-detected.
Anvisha's take
We have not scored the four readings by ear; press play and judge the voice yourself. What we measured: Gemini 3.1 Flash TTS read the four scripts in 16 to 26 seconds of audio each and was the slowest to return, 10 to 19 seconds per reading, at $0.05 per thousand characters, half of Eleven v3 and MiniMax. It ran with the default Kore voice, no style instructions and no inline tags, so what you hear is the model's default delivery; its ranking on the text-to-speech leaderboards was earned with direction, which these scripts deliberately withhold.
Gemini 3.1 Flash TTS real output: the same 4 scripts every model gets
Every voice model on this site gets the same 4 scripts, so you can compare like with like. These are the readings Gemini 3.1 Flash TTS returned, one try each, nothing re-rolled or retouched. Generated on . The 4 readings took 13 to 19 seconds each and cost $0.0487 in total.
Pacing, warmth and whether the pauses land where a voice actor would put them.
›What we asked for
Meet the Ember mug. It keeps your coffee at exactly the temperature you choose, for up to ninety minutes, and it tells your phone when it's ready. Warm from the first sip to the last. Ember. Coffee, on your terms.
fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 16.1 seconds of mp3
Steady narration over a long paragraph, clean list rhythm and a natural close.
›What we asked for
A pour-over works in four steps. First, rinse the paper filter with hot water so it doesn't taste of paper. Second, add the grounds and pour just enough water to wet them, then wait thirty seconds while they bloom. Third, pour the rest of the water in slow circles. Fourth, wait about three minutes for it to drip through. That's it: a cup that tastes like the coffee, not the machine.
fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 25.8 seconds of mp3
Range: excitement, then a hushed aside, from the words alone with no tags.
›What we asked for
We did it! We actually did it! I can't believe it's finally over. Okay. Okay, deep breath. Let's not tell anyone until Monday, alright? Just... let me have this one quiet night.
fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 17.2 seconds of mp3
Reading numbers, times, currency, an email address and two hard names correctly.
›What we asked for
Your order ships on the 24th of September 2026 and arrives by 9:15 a.m. The total is $1,249.99, including VAT. Questions? Ask for Siobhan Nguyen or Ravi Parikh on 0800 555 0199, or email help@moda.app.
fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 24.0 seconds of mp3
How the cost under each picture was worked out: fal's Gemini 3.1 Flash TTS rate of $0.05 per 1,000 characters (fal model page, read 2026-09-10) for the 213 characters sent; fal does not return the charge.
Gemini 3.1 Flash TTS pricing
| Where | Unit | Price | Source |
|---|---|---|---|
| fal | 1,000 characters | $0.05 | fal model page, read |
| Gemini API | Million audio output tokens Plus $1 per million text tokens in; audio is 25 tokens a second, so about $0.50 per 1,000 seconds of speech. | $20.00 | Gemini API pricing, read |
Is Gemini 3.1 Flash TTS free?
Google's free API tier does not include the TTS model; the Gemini app can read responses aloud with it on paid plans. fal charges per character.
- •Gemini API: paid tier only for TTS.
- •fal: $0.05 per thousand characters.
Gemini API pricing, read
Where to use Gemini 3.1 Flash TTS
- Gemini API · Token-billed, with multi-speaker support.
- fal · $0.05 per thousand characters.
- Voyager · Runs inside Voyager, the agent for creative work. Private preview by waitlist.
Voyager
Run Gemini 3.1 Flash TTS inside Voyager
Voyager is an agent for creative work that runs every model in this reference, Gemini 3.1 Flash TTS included. Brief it, and it picks the model, makes the thing and hands back something you can edit. Create an account to get started.
What we noticed running Gemini 3.1 Flash TTS
- •Returned in 10 to 19 seconds per reading, the slowest of the four voice models.
- •16.1 seconds for the ad read, the longest of the four models on that script; the others took 13 to 16. See the ad read reading.
- •No style_instructions and no inline tags were sent; the model's headline feature is the direction these scripts do not give it.
Gemini 3.1 Flash TTS limits and API parameters
| Limit | Value |
|---|---|
| Cloning | None; 30 stock voices. fal API reference, read |
| Latency | Our readings took 10 to 19 seconds to return, the slowest of the four voice models. fal API reference, read |
›Every setting, for developers
fal (fal-ai/gemini-3.1-flash-tts)
fal API reference, read
| Parameter | Type | Values | Default |
|---|---|---|---|
| prompt | string | The script, with optional tags like [sigh] or [whispering] Required. | |
| style_instructions | string | Plain-language direction, for example 'speak warmly and slowly' | |
| voice | enum | 30 voices including Kore, Puck, Charon, Zephyr, Aoede | Kore |
| language_code | enum | 70-plus languages | auto |
| speakers | list | speaker_id and voice pairs for dialogue | |
| temperature | number | Randomness | 1 |
| output_format | enum | wav, mp3, ogg_opus | mp3 |
What Gemini 3.1 Flash TTS will not say
Google's Generative AI Prohibited Use Policy applies: no impersonation, fraud, harassment, sexual content involving minors or deceptive audio of real people. There is no cloning, which removes the main misuse; generated audio carries Google's SynthID watermark.
Alternatives to Gemini 3.1 Flash TTS
Frequently asked questions
What is Gemini 3.1 Flash TTS?
Google's text-to-speech model in the Gemini line, with 30 voices, plain-language style directions, two-speaker dialogue and more than 80 languages. It is on the Gemini API and on fal.
How much does Gemini TTS cost?
$0.05 per thousand characters on fal, or $20 per million audio output tokens plus $1 per million text tokens on Google's API, read on 10 September 2026.
Can Gemini TTS clone a voice?
No. It offers 30 stock voices only.
How this page is made
Prices, limits and settings are read from the linked vendor and host pages on the dates shown. Every picture was generated by us through fal with the request shown under it, and the original files are kept unedited. A generation that failed is shown as a failure. Model and vendor names are used to identify the products; no vendor imagery appears on this page.
