Skip to main content
Coming soon

Gemini 3.8 Flash TTS API

Google's highest-quality speech model, released September 23, 2026: 3rd on the Artificial Analysis Text to Speech leaderboard, inline vocal bursts, two-speaker dialogue, and 130+ languages. We'll open it on Unifically once the integration is live.

What is Gemini 3.8 Flash TTS?

Gemini 3.8 Flash TTS is Google's highest-quality text-to-speech model, released on September 23, 2026. It turns text into spoken audio in 130+ languages, with vocal sounds, stage-style delivery, and two-speaker dialogue in a single request. It is coming to Unifically as google/gemini-3.8-flash-tts.

It is the biggest jump the Gemini voice line has made. On the Artificial Analysis Text to Speech leaderboard it holds 3rd place with an Elo of 1267, ahead of Qwen-Audio-3.0-TTS-Plus (1258) and far ahead of Gemini 3.1 Flash TTS (1203).

What's new in Gemini 3.8 Flash TTS

#3 on the Artificial Analysis TTS leaderboard

Gemini 3.8 Flash TTS scores an Elo of 1267 in blind listening tests, 3rd behind Eleven v4 (1319) and Sonic 3.6 (1276). Gemini 3.1 Flash TTS sits at 1203, so the new model gains 64 points on its predecessor.

Vocal bursts placed inline

Write `<laugh>`, `<sigh>`, `<gasp>`, `<breath>`, or `<long pause>` where the sound belongs and the model performs it at that exact spot. Short listener sounds like `|mhm|` and `|yeah|` make a two-person scene sound like a real talk.

2,000+ voices. The 30 curated Gemini voices stay, with the same names, and Google adds an extended library it puts at more than 2,000 voices across accents and character types.

Two-speaker scenes. One request renders a conversation between two prebuilt voices, with the turn-taking handled by the model.

130+ languages. Language is detected from the text, with regional accents such as Mexican Spanish and Quebec French in the voice library.

Style direction. A separate style note, like "whispered urgently", shapes the read without being spoken.

Best for

Games and animation

Character lines with laughs, sighs, and pauses placed exactly where the scene needs them.

Podcasts and explainers

Two-host scripts rendered in one request instead of two stitched files.

Audiobooks

Long narration where acting and pacing carry the listener.

Localized content

One script voiced in 130+ languages, with regional accents from the voice library.

Branded assistants

A fixed voice and style note that keeps the same persona across every reply.

Use cases

Turn a written outline into a two-host podcast episode, one voice per host, with the listener sounds and laughs written right into the script. Voice every character in an interactive story from a single model, with the emotion cues coming from the scene state. Render trailers and promos where one well-placed <gasp> or <long pause> carries the moment. Localize a product tour by translating the text and letting the model pick up each target language on its own.

Limitations

Gemini 3.8 Flash TTS is not live on Unifically yet, so this page is based on Google's release and independent leaderboards, not on our own runs. Dialogue tops out at two speakers per request. Each request takes up to 8,192 tokens of text, so a full audiobook needs to be split into chapters. Voice cloning and custom voice design live in a separate Google voice-management flow, and we have not confirmed that they will be part of the Unifically launch.

Gemini 3.8 Flash TTS vs Gemini 3.1 Flash TTS

3.8 wins on every public number we checked: 1267 against 1203 Elo on Artificial Analysis, and 130+ languages against 70+. The tag syntax changes too, from square-bracket directions like [whispering] in 3.1 to angle-bracket sounds like <laugh> in 3.8, so scripts written for 3.1 need a pass before they move over.

FAQs

People also ask

Not yet. The page is live ahead of the API, and the model will run as google/gemini-3.8-flash-tts on the same POST /v1/tasks endpoint as the other Gemini voices once it opens. ElevenLabs text to speech and dialogue are callable today.

Google released Gemini 3.8 Flash TTS on September 23, 2026, together with Gemini 3.8 Flash-Lite TTS. Both launched on the Gemini API and in Google AI Studio.

It sits 3rd on the Artificial Analysis Text to Speech leaderboard with an Elo of 1267, behind Eleven v4 at 1319 and Sonic 3.6 at 1276. That is 64 points above Gemini 3.1 Flash TTS at 1203.

You write short non-speech sounds in angle brackets right where they should happen, like <laugh>, <sigh>, <breath>, or <short pause>. The model performs the sound at that point instead of reading the tag. The square-bracket tags used by Gemini 3.1 Flash TTS are not part of the 3.8 syntax.

It keeps the 30 curated Gemini voices, such as Kore, Puck, and Charon, and adds an extended library that Google puts at more than 2,000 voices. It speaks more than 130 languages with automatic language detection.

Yes. One request can render a conversation between up to two speakers using prebuilt voices, with the turn-taking handled by the model.

Each request takes up to 8,192 input tokens of text, and Google serves up to 16,384 audio output tokens. Audio runs at about 25 tokens per second, so one request covers roughly ten minutes of speech.

Flash is the higher-quality model, ranked 3rd on Artificial Analysis against 6th for Flash-Lite, and it covers 130+ languages against 100+. Flash-Lite is tuned for throughput and low latency, for high-volume narration, dubbing, and voice agents.