Gemini 3.8 text-to-speech says hello
2026-09-27 · Google DeepMind
Gemini 3.8 Text-to-Speech Says Hello
On September 23, 2026, Google introduced two new text-to-speech (TTS) models to the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models transform voice generation from static presets into a dynamic creative studio. They are available across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Two New TTS Models
- Gemini 3.8 Flash TTS: Built for deep creative direction and character design. It allows users to create entirely new voices from scratch using natural language prompts and direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling.
- Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. It is optimized for high-volume dubbing, audio content creation, and expressive voice agents, offering fine-grained control over tone, pacing, and expressive nuance.
Create and Customize Your Own Voices
Users can scale up from 30 original voices to an infinite library:
- Generative voice design: Create bespoke voices from scratch by customizing role, accent, and voice characteristics across more than 100 languages and dialects using natural language prompting.
- Expansive voice library: Access over 2,000 production-ready voices with broad language coverage, including regional varieties like Mexican Spanish, Quebec French, and Scots English.
- Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample. This is backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect developers and vocal talent.
- Save and scale: Save and manage custom voices to ensure consistent performance and minimal drift across ongoing projects.
- Voice remixing (coming soon): Pick a voice from the library and fine-tune timbre, pitch, pace, and accent using prompts.
Direct the Performance, Line by Line
Once voices are selected, both TTS models provide precise control over delivery:
- Direct performance line by line: Write stage directions or let Gemini steer delivery with natural script cues, ranging from a calm customer service agent to a whispered suspense scene.
- Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift, making it ideal for podcasts and audiobooks.
- Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script while keeping both voices distinctly separated with natural conversational turn-taking.
- Scripted vocal bursts & backchanneling: Add realistic conversational texture using non-verbal cues (like `<laughs>`, `<sigh>`, `<gasp>`) and active-listening interjections (like `|mhm|` or `|yeah|`) for precise comedic timing and reaction beats.
Built-in Safety Tools
To ensure generated audio is used responsibly, Google has included built-in safety tools like SynthID watermarking to keep the generated audio secure and protect both developers and voice talent.