# Google Unveils Gemini 3.8 Flash TTS and Flash-Lite TTS: A New Era in Expressive Text-to-Speech
Google has introduced two next-generation text-to-speech models under its Gemini Audio umbrella: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The company describes these as its most expressive audio generation models to date, engineered to serve two very different audiences — creative professionals who demand nuanced voice direction and production teams that need scalable, affordable speech synthesis.
Both models empower developers to control how each line of dialogue is delivered using plain, natural language instructions. This represents a significant leap beyond traditional TTS systems, which typically offer limited control over tone, pacing, and emotional delivery.
## Two Tiers, One Vision
The Gemini 3.8 release divides its TTS offering into two distinct tiers, each sharing the same underlying direction and control framework but optimized for different use cases.
### Gemini 3.8 Flash TTS
This model is purpose-built for deep creative work. It excels at character voice design, acting direction, and long-form narration. Target applications include video game dialogue, immersive audiobook production, scripted podcasts, and interactive entertainment experiences.
Key capabilities include granular control over acting cues, natural pacing variations, dialect switching within a single session, and backchannel responses that simulate attentive listening. The model also supports native two-speaker staging, meaning a single script can drive a full multi-turn conversation with distinctly separated voices without any additional configuration.
### Gemini 3.8 Flash-Lite TTS
Flash-Lite TTS is built for high-throughput, cost-conscious production environments. It is ideally suited for dubbing projects, large-scale audio content generation, and expressive voice agent deployments where volume and efficiency matter most.
While more streamlined than its sibling model, Flash-Lite TTS still delivers fine-grained control over tone, pacing, and expressive nuance. It ranks second on the Hume AI Overall Quality Index, confirming that cost efficiency does not come at the expense of quality.
## Getting Started
Both models are available now through the Gemini API and Google AI Studio. Access is API-only — there are no open weights available for self-hosting. Enterprise-grade API access through Gemini Enterprise is expected to launch in the near future. In the AI Studio playground, the models can be identified by the codes `gemini-3.8-flash-tts` and `gemini-3.8-flash-lite-tts`.
## Voice Design From Text Prompts
One of the most notable upgrades in the 3.8 release is the voice design system. Where earlier Gemini TTS models shipped with approximately 30 fixed voices, the new architecture opens up a dramatically larger voice ecosystem.
### Generative Voice Creation
Flash TTS can synthesize entirely new voices from text prompts that describe a character’s role, regional accent, and vocal characteristics. This feature works across more than 100 languages and dialects. Demonstrations from Google have included a high-energy Melbourne radio DJ, a synthetic monotone robot, and a dramatic Japanese dragon — each generated purely from descriptive text prompts.
### Extensive Voice Library
For developers who prefer ready-to-use options, the 3.8 release offers access to over 2,000 production-ready voices. This library includes regional varieties such as Mexican Spanish, Quebec French, and Scots English, providing authentic coverage for global content needs.
### Saving and Scaling Custom Voices
Custom voices created through generative design can be saved and reused across projects with minimal drift in quality or consistency. This makes it practical to maintain a unified vocal identity across an entire media property.
### Voice Remixing (Upcoming)
Google has announced that voice remixing will arrive soon, allowing users to adjust a library voice’s timbre, pitch, speaking pace, and accent through simple text prompts — without needing to generate a new voice from scratch.
## Directing the Performance
A standout feature of both models is the ability to receive and interpret stage directions embedded directly in a script. This allows developers to choreograph vocal performances with the same precision a director would use on a physical stage.
– **Long-Form Stability**: Voice quality, pacing, and tonal consistency hold across hours of continuous audio generation, making the models suitable for full-length audiobooks and serialized content.
– **Native Multi-Speaker Staging**: A single script can produce a natural-sounding conversation between two distinct speakers, with clean turn-taking and separation.
– **Vocal Bursts**: Non-verbal cues can be inserted using syntax like `
– **Backchanneling**: Active-listening interjections such as `|mhm|` and `|yeah|` can be scripted to create natural reaction beats and comedic timing, making dialogue feel genuinely interactive.
## Voice Replication and Content Safety
Voice replication allows developers to build a consistent vocal profile from a 30-second audio sample. The process requires the sample to be either the developer’s own voice or one for which they hold usage rights. Additionally, a verbal consent recording from the voice owner must be provided and matched against the reference speaker.
Every audio clip generated by Gemini Audio models carries a SynthID watermark — an imperceptible mark embedded directly into the audio signal. Replicated voices also come with C2PA content credentials, providing a verifiable chain of provenance. Google has outlined its broader safety approach in the Gemini 3.8 Audio model card.
It is worth noting that voice replication is currently not available in AI Studio for users located in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India.
## Benchmark Performance
Google has shared the following results from independent evaluations:
| Benchmark | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS |
|—|—|—|
| Hume AI Voice Design Benchmark | #1 overall (71.4) | — |
| Accent Modeling | Leading score (60.8) | — |
| Hume AI Overall Quality Index | #1 | #2 |
| Voice Arena (Blind Preference) | Top positions | Top positions |
Flash TTS has claimed the top spot on Voice Arena blind preference tests across multiple languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi. These results suggest that the generative voice design approach delivers perceptually superior quality compared to traditional fixed-voice TTS systems.
## Frequently Asked Questions
**What exactly is Gemini 3.8 Flash TTS?**
It is Google’s latest text-to-speech model designed for creative voice design and granular, line-by-line performance direction. It is accessible today through the Gemini API and Google AI Studio.
**How does Flash-Lite TTS differ from Flash TTS?**
Flash-Lite TTS is optimized for high-volume, cost-efficient production scenarios such as dubbing and voice agents. It trades some of the advanced creative direction features of Flash TTS for greater scalability and lower per-unit cost, while still ranking #2 on the Hume AI Overall Quality Index.
**Can I create a voice clone of my own voice?**
Yes. Voice replication requires a 30-second audio sample and a matching verbal consent recording from the voice owner. The feature is currently unavailable in certain regions, including the UK, EEA, and India.
**Are these models available for self-hosting?**
No. Both models are API-only at this time, with no open weights provided for local deployment.
**How many languages do these models support?**
The generative voice design feature works across more than 100 languages and dialects. The production voice library covers a wide range of regional varieties as well.
**What is SynthID?**
SynthID is Google’s imperceptible watermarking technology embedded directly into the audio output of Gemini Audio models. It provides a reliable method for identifying AI-generated content without affecting audio quality.
## Conclusion
The arrival of Gemini 3.8 Flash TTS and Flash-Lite TTS marks a meaningful shift in what text-to-speech technology can achieve. By splitting the offering between a creative-first model and a scale-first model, Google has addressed the needs of both independent creators and large-scale production teams. The combination of prompt-based voice design, script-level performance direction, built-in safety watermarks, and multilingual support positions these models as leading tools in the rapidly evolving landscape of AI-powered audio generation.
Thank you for reading



