Aivora
Google DeepMindAudio AIIntermediate

Google Introduces Gemini 3.8 TTS: Transforming Text-to-Speech into a Creative Voice Studio

Google 發表 Gemini 3.8 TTS:從語音合成邁向客製化「聲音工作室」的革命性升級

2 min read
Google Introduces Gemini 3.8 TTS: Transforming Text-to-Speech into a Creative Voice Studio
The 30-second version

Google DeepMind has introduced two new models: Gemini 3.8 Flash TTS for deep creative control and Gemini 3.8 Flash-Lite TTS for cost-effective scale. Users can design custom voices using natural language prompts or clone a voice with a 30-second sample. Featuring line-by-line script directing, multi-speaker staging, and realistic backchanneling (like laughs and sighs), both models topped Hume AI's quality and design benchmarks.

Key points

01

Generative Voice Design

Create bespoke voices from scratch using natural language prompts to customize role, accent, and qualities across 100+ languages.

02

Line-by-Line Direction

Direct tone, pacing, and conversational cues line-by-line. Supports native two-speaker interactions and natural backchanneling.

03

Safeguarded Voice Cloning

Replicate consistent vocal profiles from a 30-second sample, secured with voice consent verification, SynthID, and C2PA credentials.

04

Benchmarked Excellence

Secured the #1 and #2 spots on Hume AI's Overall Quality Index and topped multi-language human evaluations on Voice Arena.

How it works

Comparison of Gemini 3.8 TTS Models
Gemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS
Core Focus深度創意控制、角色塑造與情境執導高吞吐量、低成本的大規模應用
Key Use Cases遊戲配音、有聲書、多人口白 Podcast影片自動翻譯配音、對話式客服 Agent
Hume AI Quality Rank第 1 名 (兼具極致語音設計表現)第 2 名 (高效率下仍維持優秀音質)
Native Google ProductsGemini NotebookGoogle Vids

Why it matters

This update elevates TTS from static presets to a fully-realized digital vocal studio. Creators can dynamically tailor emotionally nuanced, localized speech for gaming, audiobooks, and branding with unprecedented control. Crucially, the integration of SynthID and strict consent verification models a responsible path forward for intellectual property protection in synthetic media.

Who it affects

  • AI Developer
  • Content Creator
  • Enterprise Leader
  • Product Manager

How to use it

  1. 1Dramatic Voice Acting for Games & Audiobooks
  2. 2Multi-Speaker Staging for Podcasts & Screenplays
  3. 3High-Volume, Cost-Effective Dubbing and Voice Agents

Limitations & caveats

  • The prompt-based Voice Remixing feature for fine-tuning timbre and pitch is marked as 'coming soon' and not yet available.
  • Rigorous verbal consent matching is safety-critical but may introduce friction to high-volume automated replication pipelines.

Related

Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking
arXivAudio AI

Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking

聽見真假:VeriSpeak 揭示語音語言模型在事實查核中的「模態差距」

This study introduces VeriSpeak, a benchmark for speech-based fact-checking, revealing a critical text-speech modality gap in Large Audio Language Models (LALMs) and demonstrating how explicit reasoning with retrieval achieves 86.1% accuracy.

2 min read