Hume AI Launches Real World VoiceEQ, a Benchmark of 40+ Voice Models Rated by 1M+ Humans
Image via huggingface.co
Hume AI has released Real World VoiceEQ, a large-scale benchmark that evaluates more than 40 proprietary and open-source voice AI models across 15+ dimensions and 60+ metrics covering ASR, TTS, Speech-to-Speech, and Speech Understanding. Built on over one million individual human ratings spanning diverse demographics, speaking styles, and acoustic environments, it is designed to surface the gaps that traditional benchmarks — focused on word error rate and latency — routinely miss, including emotion recognition, accent robustness, and speaker consistency.
Key findings challenge the notion of a single "best" voice model: no system ranked in the top five across all eight TTS capability groups, and the research finds that voice models have generally become better at generating speech than at understanding it. Evaluations were conducted using Kairos, Hume AI's voice-native evaluation platform, which the company also offers to enterprises and AI labs for custom benchmarking, failure-mode analysis, and human preference data collection for RLHF pipelines.