Research

Open TTS Leaderboard Launches Scalable, Objective Evaluation for Multilingual Text-to-Speech and Voice Cloning

5 days ago

Image via huggingface.co

A new Open TTS Leaderboard has launched on Hugging Face to address a growing gap between the pace of open-source TTS model releases — now exceeding 8,000 models on the Hub — and the capacity of existing evaluation frameworks to keep up. Unlike arena-style leaderboards such as TTS Arena v2 and Artificial Analysis Voice Arena, which rely on human preference voting and Elo rankings, the new leaderboard uses objective metrics: word/character error rate (WER/CER) via Qwen3 ASR for intelligibility, inverse real-time factor and time-to-first-audio for speed, and WavLM cosine similarity for voice cloning fidelity. The system reduces evaluation time from weeks to hours, and currently benchmarks models across English and multilingual splits including Seed TTS Eval and CV3 Eval.

The leaderboard is positioned as complementary to, not a replacement for, human preference ranking — acknowledging that objective metrics cannot capture naturalness or expressiveness. Notably, open-source models are underrepresented on existing arenas (only 16 of 92 models on Artificial Analysis are open-weights), a skew the new leaderboard aims to correct by lowering the barrier to evaluation. Top performers on English WER at launch include Kokoro-82M, Supertonic-3, and Fish Audio's s2-pro. The project is community-driven and open to feedback.