Hugging Face Launches Open TTS Leaderboard for Multilingual TTS and Voice Cloning
Hugging Face 推出 Open TTS Leaderboard:多語音合成與聲音複製的開源評測基準
With over 8,000 TTS models on Hugging Face, human-voting arenas are too slow and costly to scale, often leaving open-weights models underrepresented. The new Open TTS Leaderboard solves this by using objective, automated metrics—such as ASR-based WER/CER for intelligibility, SIM for speaker similarity, and TTFA for streaming latency—to evaluate models in hours instead of weeks. It highlights multilingual support, voice cloning, and streaming capabilities, bridging the gap with an interactive 'Listen' tab for direct output comparisons.
Key points
Fast and Objective Evaluation
Reduces evaluation time from weeks of human voting to hours, enabling scalable and consistent benchmarks for open-source TTS models.
Three Core Metric Pillars
Measures intelligibility via ASR-based WER/CER, voice cloning similarity via WavLM speaker embeddings, and inference speed via RTFx.
Streaming Latency Benchmark
Benchmarks Time-to-First-Audio (TTFA) on H200 GPU and CPU, providing critical latency insights for real-time interactive voice agents.
Focus on Open Source & Multilingualism
Addresses the underrepresentation of open-weights models in traditional arenas, offering multilingual benchmarks and voice cloning evaluations.
How it works
| 傳統語音競技場 (Voice Arenas) | Open TTS Leaderboard | |
|---|---|---|
| Evaluation Time | 數週 (Weeks) | 數小時 (Hours) |
| Evaluation Method | 人類主觀投票 (Human voting) | 客觀自動化指標 (Objective metrics: WER, SIM, TTFA) |
| Open-Source Representation | 較低 (多數為商業 API 託管) | 極高 (主打並優化開源權重模型) |
| Consistency | 波動大 (投票者主觀標準易變) | 穩定一致 (可重複執行的標準化流程) |
Why it matters
Traditional TTS evaluation relies on subjective human opinion, which suffers from voter inconsistency and high hosting costs, leaving open-weights models sidelined. By establishing objective, automated pipelines, this leaderboard cuts evaluation latency to hours. It democratizes benchmark participation for the 8k+ open-source models on Hugging Face and provides crucial metrics like TTFA for real-time voice agent development.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Enterprise Leader
How to use it
- 1Evaluating and selecting cost-effective open-source multilingual TTS models for local deployment
- 2Comparing streaming latency (TTFA) across models on GPU or CPU to develop low-latency voice agents
- 3Analyzing character/word error rates and speaker similarity of models in non-English languages
Limitations & caveats
- Automated objective metrics are proxies and cannot fully capture human subjective preferences regarding naturalness and expressiveness.
- Multilingual evaluation heavily relies on specific datasets (like Seed TTS Eval and CV3 Eval), which may not represent all real-world scenarios.
Related

Fine-Tuning NVIDIA Nemotron ASR for Regional Dialects: A Guide to Dialect Adaptation
如何微調 NVIDIA Nemotron 語音辨識模型?以沙烏地阿拉伯方言為例的跨語言適應指南
This guide details how to fine-tune NVIDIA Nemotron 3.5 ASR using the NeMo framework, leveraging minimal curation, replay mixing, and bucketed batches to drastically improve regional dialect recognition while preserving baseline language accuracy.
EmoRES-TTS: Training-Free Residual-Enhanced Vector Steering for Emotional Speech Generation
免訓練提升語音情緒!EmoRES-TTS 透過「殘差向量分解」實現精準可控的語音合成
EmoRES-TTS improves emotional speech generation by decomposing steering vectors into shared and residual components, achieving superior training-free emotion control without retraining the backbone models.
Hearing the Truth: VeriSpeak Exposes the Text-Speech Modality Gap in Fact-Checking
聽見真假:VeriSpeak 揭示語音語言模型在事實查核中的「模態差距」
This study introduces VeriSpeak, a benchmark for speech-based fact-checking, revealing a critical text-speech modality gap in Large Audio Language Models (LALMs) and demonstrating how explicit reasoning with retrieval achieves 86.1% accuracy.