Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS is Google's text-to-speech model in the Gemini line, released in preview on 15 April 2026 as the successor to the 2.5 Flash TTS preview. It takes a script plus style instructions written in plain language, inline tags such as [laughing], and a speakers list for two-voice dialogue, across 30 named voices and more than 80 languages. On fal it costs $0.05 per thousand characters; on Google's own API it is billed by the token, $1 per million text tokens in and $20 per million audio tokens out, at 25 tokens per second of audio.

Anvisha Pai

Anvisha Pai, Co-founder & CEO, Voyager

Tested

Google DeepMind · released

Gemini 3.1 Flash TTS reading its own self-introduction, generated by us.
›What we asked for

Hi, I'm Gemini 3.1 Flash TTS, Google's voice model. I take stage directions written in plain language, but today I'm reading the script as written. Here is the same script every voice on this site reads, so you can compare us.

fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 16.4 seconds of mp3

At a glance

Price
$0.05 per thousand characters on fal, half of Eleven v3; our four scripts cost about five cents.
Voices
Thirty named voices, Kore by default, which we used; no cloning.
Direction
A style_instructions field in plain language plus inline tags; our scripts used neither.
Dialogue
A speakers list assigns a voice per speaker for two-person scripts.
Languages
More than 80, auto-detected.

Anvisha's take

We have not scored the four readings by ear; press play and judge the voice yourself. What we measured: Gemini 3.1 Flash TTS read the four scripts in 16 to 26 seconds of audio each and was the slowest to return, 10 to 19 seconds per reading, at $0.05 per thousand characters, half of Eleven v3 and MiniMax. It ran with the default Kore voice, no style instructions and no inline tags, so what you hear is the model's default delivery; its ranking on the text-to-speech leaderboards was earned with direction, which these scripts deliberately withhold.

Gemini 3.1 Flash TTS real output: the same 4 scripts every model gets

Every voice model on this site gets the same 4 scripts, so you can compare like with like. These are the readings Gemini 3.1 Flash TTS returned, one try each, nothing re-rolled or retouched. Generated on . The 4 readings took 13 to 19 seconds each and cost $0.0487 in total.

Ad read16s of audio · 13s wait · $0.0106

Pacing, warmth and whether the pauses land where a voice actor would put them.

›What we asked for

Meet the Ember mug. It keeps your coffee at exactly the temperature you choose, for up to ninety minutes, and it tells your phone when it's ready. Warm from the first sip to the last. Ember. Coffee, on your terms.

fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 16.1 seconds of mp3

Explainer paragraph26s of audio · 17s wait · $0.0192

Steady narration over a long paragraph, clean list rhythm and a natural close.

›What we asked for

A pour-over works in four steps. First, rinse the paper filter with hot water so it doesn't taste of paper. Second, add the grounds and pour just enough water to wet them, then wait thirty seconds while they bloom. Third, pour the rest of the water in slow circles. Fourth, wait about three minutes for it to drip through. That's it: a cup that tastes like the coffee, not the machine.

fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 25.8 seconds of mp3

Emotional line17s of audio · 13s wait · $0.0089

Range: excitement, then a hushed aside, from the words alone with no tags.

›What we asked for

We did it! We actually did it! I can't believe it's finally over. Okay. Okay, deep breath. Let's not tell anyone until Monday, alright? Just... let me have this one quiet night.

fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 17.2 seconds of mp3

Names, dates and prices24s of audio · 19s wait · $0.01

Reading numbers, times, currency, an email address and two hard names correctly.

›What we asked for

Your order ships on the 24th of September 2026 and arrives by 9:15 a.m. The total is $1,249.99, including VAT. Questions? Ask for Siobhan Nguyen or Ravi Parikh on 0800 555 0199, or email help@moda.app.

fal-ai/gemini-3.1-flash-tts
output_format: mp3, voice: Kore
returned 24.0 seconds of mp3

How the cost under each picture was worked out: fal's Gemini 3.1 Flash TTS rate of $0.05 per 1,000 characters (fal model page, read 2026-09-10) for the 213 characters sent; fal does not return the charge.

Gemini 3.1 Flash TTS pricing

WhereUnitPriceSource
fal1,000 characters$0.05fal model page, read
Gemini APIMillion audio output tokens
Plus $1 per million text tokens in; audio is 25 tokens a second, so about $0.50 per 1,000 seconds of speech.
$20.00Gemini API pricing, read

Is Gemini 3.1 Flash TTS free?

Google's free API tier does not include the TTS model; the Gemini app can read responses aloud with it on paid plans. fal charges per character.

  • •Gemini API: paid tier only for TTS.
  • •fal: $0.05 per thousand characters.

Gemini API pricing, read

Where to use Gemini 3.1 Flash TTS

  • Gemini API · Token-billed, with multi-speaker support.
  • fal · $0.05 per thousand characters.
  • Voyager · Runs inside Voyager, the agent for creative work. Private preview by waitlist.

Voyager

Run Gemini 3.1 Flash TTS inside Voyager

Voyager is an agent for creative work that runs every model in this reference, Gemini 3.1 Flash TTS included. Brief it, and it picks the model, makes the thing and hands back something you can edit. Create an account to get started.

What we noticed running Gemini 3.1 Flash TTS

  • •Returned in 10 to 19 seconds per reading, the slowest of the four voice models.
  • •16.1 seconds for the ad read, the longest of the four models on that script; the others took 13 to 16. See the ad read reading.
  • •No style_instructions and no inline tags were sent; the model's headline feature is the direction these scripts do not give it.

Gemini 3.1 Flash TTS limits and API parameters

LimitValue
CloningNone; 30 stock voices.
fal API reference, read
LatencyOur readings took 10 to 19 seconds to return, the slowest of the four voice models.
fal API reference, read
›Every setting, for developers

fal (fal-ai/gemini-3.1-flash-tts)

fal API reference, read

ParameterTypeValuesDefault
promptstringThe script, with optional tags like [sigh] or [whispering]
Required.
style_instructionsstringPlain-language direction, for example 'speak warmly and slowly'
voiceenum30 voices including Kore, Puck, Charon, Zephyr, AoedeKore
language_codeenum70-plus languagesauto
speakerslistspeaker_id and voice pairs for dialogue
temperaturenumberRandomness1
output_formatenumwav, mp3, ogg_opusmp3

What Gemini 3.1 Flash TTS will not say

Google's Generative AI Prohibited Use Policy applies: no impersonation, fraud, harassment, sexual content involving minors or deceptive audio of real people. There is no cloning, which removes the main misuse; generated audio carries Google's SynthID watermark.

Google Generative AI Prohibited Use Policy, read

Alternatives to Gemini 3.1 Flash TTS

Frequently asked questions

What is Gemini 3.1 Flash TTS?

Google's text-to-speech model in the Gemini line, with 30 voices, plain-language style directions, two-speaker dialogue and more than 80 languages. It is on the Gemini API and on fal.

How much does Gemini TTS cost?

$0.05 per thousand characters on fal, or $20 per million audio output tokens plus $1 per million text tokens on Google's API, read on 10 September 2026.

Can Gemini TTS clone a voice?

No. It offers 30 stock voices only.

How this page is made

Prices, limits and settings are read from the linked vendor and host pages on the dates shown. Every picture was generated by us through fal with the request shown under it, and the original files are kept unedited. A generation that failed is shown as a failure. Model and vendor names are used to identify the products; no vendor imagery appears on this page.

Anvisha Pai

Anvisha Pai

Co-founder & CEO, Voyager

Anvisha is the CEO of Voyager and a repeat, Y Combinator-backed startup founder. She was previously a PM at Dropbox. She believes nobody should need a design degree to make something that looks great.