Google Introduces Gemini 3.8 TTS: Transforming Text-to-Speech into a Creative Voice Studio
Google 發表 Gemini 3.8 TTS:從語音合成邁向客製化「聲音工作室」的革命性升級

Google DeepMind has introduced two new models: Gemini 3.8 Flash TTS for deep creative control and Gemini 3.8 Flash-Lite TTS for cost-effective scale. Users can design custom voices using natural language prompts or clone a voice with a 30-second sample. Featuring line-by-line script directing, multi-speaker staging, and realistic backchanneling (like laughs and sighs), both models topped Hume AI's quality and design benchmarks.
Key points
Generative Voice Design
Create bespoke voices from scratch using natural language prompts to customize role, accent, and qualities across 100+ languages.
Line-by-Line Direction
Direct tone, pacing, and conversational cues line-by-line. Supports native two-speaker interactions and natural backchanneling.
Safeguarded Voice Cloning
Replicate consistent vocal profiles from a 30-second sample, secured with voice consent verification, SynthID, and C2PA credentials.
Benchmarked Excellence
Secured the #1 and #2 spots on Hume AI's Overall Quality Index and topped multi-language human evaluations on Voice Arena.
How it works
| Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | |
|---|---|---|
| Core Focus | 深度創意控制、角色塑造與情境執導 | 高吞吐量、低成本的大規模應用 |
| Key Use Cases | 遊戲配音、有聲書、多人口白 Podcast | 影片自動翻譯配音、對話式客服 Agent |
| Hume AI Quality Rank | 第 1 名 (兼具極致語音設計表現) | 第 2 名 (高效率下仍維持優秀音質) |
| Native Google Products | Gemini Notebook | Google Vids |
Why it matters
This update elevates TTS from static presets to a fully-realized digital vocal studio. Creators can dynamically tailor emotionally nuanced, localized speech for gaming, audiobooks, and branding with unprecedented control. Crucially, the integration of SynthID and strict consent verification models a responsible path forward for intellectual property protection in synthetic media.
Who it affects
- AI Developer
- Content Creator
- Enterprise Leader
- Product Manager
How to use it
- 1Dramatic Voice Acting for Games & Audiobooks
- 2Multi-Speaker Staging for Podcasts & Screenplays
- 3High-Volume, Cost-Effective Dubbing and Voice Agents
Limitations & caveats
- The prompt-based Voice Remixing feature for fine-tuning timbre and pitch is marked as 'coming soon' and not yet available.
- Rigorous verbal consent matching is safety-critical but may introduce friction to high-volume automated replication pipelines.
Related
Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking
聽見真假:VeriSpeak 揭示語音語言模型在事實查核中的「模態差距」
This study introduces VeriSpeak, a benchmark for speech-based fact-checking, revealing a critical text-speech modality gap in Large Audio Language Models (LALMs) and demonstrating how explicit reasoning with retrieval achieves 86.1% accuracy.