Aivora
arXivAudio AIIntermediate

EmoRES-TTS: Training-Free Residual-Enhanced Vector Steering for Emotional Speech Generation

免訓練提升語音情緒!EmoRES-TTS 透過「殘差向量分解」實現精準可控的語音合成

2 min read
EmoRES-TTS: Training-Free Residual-Enhanced Vector Steering for Emotional Speech Generation
The 30-second version

Adjusting emotions in TTS models typically requires expensive retraining and labeled data. While training-free vector steering (such as CoCoEmo) exists, treating the emotion vector as a single direction limits performance. EmoRES-TTS addresses this by decomposing the emotion vector into a shared component (moving away from neutral) and a residual component (directing toward the target emotion). Evaluated on IndexTTS-2 and CosyVoice2, EmoRES significantly outperforms CoCoEmo in both objective metrics and human preference for naturalness.

Key points

01

Emotion Vector Decomposition

Discovers that an emotion steering vector comprises a shared component (moving away from neutral) and a residual component (directing to the target emotion).

02

Training-Free Vector Steering

EmoRES modifies the internal representations of a frozen TTS model directly, avoiding expensive retraining costs and the need for emotion-labeled data.

03

Substantial Objective Gains

On the IEMOCAP dataset, EmoRES improves rank correlation by up to 118.8% and emotion hit rate by up to 20.1% over the prior state-of-the-art.

04

Stronger Human Preference

Human evaluators preferred EmoRES for naturalness in up to 63.8% of pairwise comparisons, with up to a 35% relative improvement in target emotion identification.

How it works

EmoRES Vector Steering Process
Target Text & EmotionEmotion VectorEmoRES DecompositionShared (Away fromNeutral)Residual (TargetEmotion)Steering ControlFrozen Backbone ModelNatural EmotionalSpeech

Why it matters

This research addresses the trade-off between high-quality emotion control and expensive computational retraining in TTS. By proving that emotion steering vectors can be decomposed and manipulated without retraining, EmoRES-TTS offers an efficient, training-free plug-and-play solution. Its compatibility with modern backbones like CosyVoice2 paves the way for cost-effective, expressive, and highly controllable voice assistants and digital avatars.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Dynamic emotion adjustment in real-time voice assistants without retraining dedicated models.
  2. 2Fine-tuning dramatic tension and specific emotions for virtual anchors and audiobook narration.
  3. 3Low-cost generation of multi-intensity emotional speech training data for other audio models.

Limitations & caveats

  • Reliant on the representation space of the frozen backbone model; if the backbone lacks certain phonetic capabilities, steering might be less effective.
  • Requires precise estimation and manual tuning of the weights for both the shared and residual components during inference.

Related

Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking
arXivAudio AI

Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking

聽見真假:VeriSpeak 揭示語音語言模型在事實查核中的「模態差距」

This study introduces VeriSpeak, a benchmark for speech-based fact-checking, revealing a critical text-speech modality gap in Large Audio Language Models (LALMs) and demonstrating how explicit reasoning with retrieval achieves 86.1% accuracy.

2 min read