EmoRES-TTS: Training-Free Residual-Enhanced Vector Steering for Emotional Speech Generation
免訓練提升語音情緒!EmoRES-TTS 透過「殘差向量分解」實現精準可控的語音合成
Adjusting emotions in TTS models typically requires expensive retraining and labeled data. While training-free vector steering (such as CoCoEmo) exists, treating the emotion vector as a single direction limits performance. EmoRES-TTS addresses this by decomposing the emotion vector into a shared component (moving away from neutral) and a residual component (directing toward the target emotion). Evaluated on IndexTTS-2 and CosyVoice2, EmoRES significantly outperforms CoCoEmo in both objective metrics and human preference for naturalness.
Key points
Emotion Vector Decomposition
Discovers that an emotion steering vector comprises a shared component (moving away from neutral) and a residual component (directing to the target emotion).
Training-Free Vector Steering
EmoRES modifies the internal representations of a frozen TTS model directly, avoiding expensive retraining costs and the need for emotion-labeled data.
Substantial Objective Gains
On the IEMOCAP dataset, EmoRES improves rank correlation by up to 118.8% and emotion hit rate by up to 20.1% over the prior state-of-the-art.
Stronger Human Preference
Human evaluators preferred EmoRES for naturalness in up to 63.8% of pairwise comparisons, with up to a 35% relative improvement in target emotion identification.
How it works
Why it matters
This research addresses the trade-off between high-quality emotion control and expensive computational retraining in TTS. By proving that emotion steering vectors can be decomposed and manipulated without retraining, EmoRES-TTS offers an efficient, training-free plug-and-play solution. Its compatibility with modern backbones like CosyVoice2 paves the way for cost-effective, expressive, and highly controllable voice assistants and digital avatars.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Dynamic emotion adjustment in real-time voice assistants without retraining dedicated models.
- 2Fine-tuning dramatic tension and specific emotions for virtual anchors and audiobook narration.
- 3Low-cost generation of multi-intensity emotional speech training data for other audio models.
Limitations & caveats
- Reliant on the representation space of the frozen backbone model; if the backbone lacks certain phonetic capabilities, steering might be less effective.
- Requires precise estimation and manual tuning of the weights for both the shared and residual components during inference.
Related
Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking
聽見真假:VeriSpeak 揭示語音語言模型在事實查核中的「模態差距」
This study introduces VeriSpeak, a benchmark for speech-based fact-checking, revealing a critical text-speech modality gap in Large Audio Language Models (LALMs) and demonstrating how explicit reasoning with retrieval achieves 86.1% accuracy.

Google Introduces Gemini 3.8 TTS: Transforming Text-to-Speech into a Creative Voice Studio
Google 發表 Gemini 3.8 TTS:從語音合成邁向客製化「聲音工作室」的革命性升級
Google DeepMind launches Gemini 3.8 Flash and Flash-Lite TTS, introducing prompt-driven voice creation and granular, line-by-line performance controls to redefine text-to-speech.