
Gemini 3.8 Flash TTS: What Prompt-Built AI Voices Can Do
Gemini 3.8 Flash TTS and Flash-Lite TTS let developers design new AI voices from a text prompt in 100+ languages. Here's what both models can do.
Google has turned voice design into a writing exercise. On September 23, 2026, the company released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models that can invent a brand-new voice from a plain-language description, then perform a script line by line with laughs, sighs and pauses exactly where the writer wants them. Both are live today in the Gemini API and Google AI Studio.
- Two models: Flash TTS for creative voice and character work; Flash-Lite TTS for high-volume, cost-efficient jobs such as dubbing and voice agents
- Coverage: more than 100 languages and dialects, with over 2,000 production-ready voices
- Voice design: describe a voice in words, or replicate one from a 30-second sample after a built-in consent check
- Availability: both models in the Gemini API and AI Studio from September 23, with Gemini Enterprise support coming soon
What Is Gemini 3.8 Flash TTS?
Flash TTS is the creative member of the pair. Google pitches it at game studios, audiobook producers and podcasters who need characters rather than generic narrators. Instead of scrolling through a catalog, a developer writes a description and the model generates a matching voice. That voice can be saved and reused, so a character sounds the same in chapter one and chapter forty.
Flash-Lite TTS trades some of that creative range for throughput. Google positions it for dubbing, large content pipelines and real-time voice agents, where the cost per minute of audio matters more than bespoke character work. Inside Google's own products, Flash TTS is arriving in Gemini Notebook and Flash-Lite TTS in Google Vids.
Directing a Performance, Not Just Reading Text
The most interesting part of this AI voice generation release is the directing layer. Scripts can carry natural cues that tell the model how to deliver each line, and Google lists inline tags for vocal bursts and backchanneling, the little "mhm" and "yeah" sounds that make a conversation feel real. The models can also stage native two-speaker scenes and generate hours of continuous audio, which is what long-form podcast and audiobook work actually needs.
Where the Benchmarks Stand
Google says Flash TTS ranks first on Hume AI's Voice Design Benchmark with a score of 71.4, and that Flash and Flash-Lite hold the top two spots on Hume AI's Overall Quality Index. The company also reports leading Voice Arena positions for Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. These are Google's own reported rankings, so treat them as a strong opening claim until independent listeners weigh in.
How Does Gemini Voice Cloning Stay Safe?
Voice replication from a 30-second clip is the feature that needs the most care, and Google has built guardrails in from day one. Replication requires a consent verification step, every clip produced by either model carries an inaudible SynthID watermark, and replicated voices also carry C2PA content credentials. Google is not offering voice replication in Illinois, Texas, the European Economic Area, the UK, Switzerland or India for now.
That combination makes it easier for platforms and listeners to identify synthetic audio, and it is a good template for the rest of the text-to-speech industry to follow.
What It Means for Builders
For developers who have followed Google's voice work, this is a big step up from April's release, which we covered in our Gemini 3.1 Flash TTS launch report. It also pairs naturally with the real-time reasoning in Gemini 3.8 Live for voice agents: Live handles the conversation, while the new text-to-speech models handle the performance.
Google has not published pricing for either model yet, so teams should run small tests in AI Studio before committing a production pipeline. Indie game developers who could never afford a full voice cast, small publishers turning backlists into audiobooks, and educators producing lessons in dozens of languages all have a new, accessible tool today. For more model launches like this one, browse our AI coverage.
Sources: Google Blog — September 23, 2026; The Next Web — September 23, 2026.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
