Google DeepMind Introduces Gemini 3.8 Live with Live Avatar: Real-Time Visual AI Agents for Enterprise
Google DeepMind 發表 Gemini 3.8 Live with Live Avatar:具備即時視覺化身與背景任務處理能力的企業級 AI 代理

Building on Gemini 3.8 Live, Google DeepMind introduced Live Avatar to bring real-time visual personas to its dialogue models. It processes visual and audio inputs simultaneously to deliver natural expressions and precise lip-syncing. Crucially, it supports asynchronous tool calling, enabling the avatar to execute complex backend tasks (like checking in guests) in the background while maintaining an uninterrupted conversation. It supports seamless switching across 97 languages and integrates SynthID watermarks for safety.
Key points
Real-time Multimodal Presence
Pairs real-time video generation with speech, delivering precise lip-syncing and natural facial expressions for fluid conversations.
Asynchronous Tool Calling
Enables the avatar to trigger background tool calls and fetch data silently without disrupting the ongoing live dialogue.
97-Language Dynamic Adaptation
Dynamically adapts lip-sync and expressions across 97 languages mid-conversation without degrading video fidelity.
SynthID Watermarking
Integrates imperceptible SynthID watermarks directly into audio and video outputs to ensure AI-generated content is detectable.
How it works
Why it matters
Traditional virtual assistants suffer from pauses when executing tasks. By implementing asynchronous tool calling, Live Avatar can perform complex backend workflows (such as reservation checks) while maintaining an empathetic, uninterrupted conversation. This transforms AI agents from basic voice assistants into expressive, functional visual representatives for customer service and interactive guides.
Who it affects
- Enterprise Leader
- AI Developer
- Product Manager
How to use it
- 1Smart reception & customer service: e.g., handling hotel check-ins in the background while maintaining polite live conversation with guests.
- 2Brand customized avatars: developers can generate responsive, expressive animated avatars reflecting their brand's identity using a single reference image.
Limitations & caveats
- Custom avatar generation via reference images is currently restricted to enterprise allowlisted organizations.
- Real-time, synchronized high-fidelity video generation requires robust bandwidth and may experience latency in poor network environments.
Related
Agensh: Scaling Multi-Agent Collaboration to 1,024 Agents Without a Central Orchestrator
突破中心化瓶頸!Agensh 框架將多 Agent 協作無縫擴展至 1,024 個智慧體
Agensh is a decentralized, self-organized multi-agent harness that scales up to 1,024 agents using an asynchronous cooperation loop, significantly boosting efficiency in complex coding tasks.
Grow the Harness, Not the Context: Building Low-Cost Specialist Agents via Failure-Guided Code Synthesis
別撐大脈絡,改建構程式框架:Growing Harness 透過「失敗導向學習」自動生成高效能 Agent
Growing Harness is a novel paradigm that turns recurring agent control decisions into reusable, executable code via failure-guided synthesis, slashing inference costs by up to 98.6% while keeping small models highly capable.