Hugging Face has introduced the Open TTS Leaderboard, a new evaluation framework designed to assess open-source text-to-speech (TTS) models using objective metrics. The initiative aims to address the fragmentation in TTS evaluation and the slow pace of human-preference-based arenas, which have struggled to keep up with the rapid release of new models.
What Happened
As of September 30, 2026, more than 8,000 TTS models are available on the Hugging Face Hub. Despite this abundance, evaluation standards have remained inconsistent. Existing arena-based leaderboards, such as TTS Arena v2, Artificial Analysis, and Voice Arena, rely on human preference votes to compute Elo scores. However, these methods are time-consuming and often underrepresent open-weight models. For instance, the source notes that only 16 of the 92 models listed on Artificial Analysis were open-weight as of the same date, likely due to the operational overhead of hosting open models compared to integrating commercial APIs.
The Open TTS Leaderboard bypasses these bottlenecks by using automated, objective metrics. It evaluates models on three primary dimensions: intelligibility, measured by Word Error Rate (WER) and Character Error Rate (CER) using Qwen3 ASR; speed, measured by inverse real-time factor (RTFx) for batched inference and time-to-first-audio (TTFA) for streaming; and speaker similarity, calculated via cosine similarity of WavLM speaker embeddings. This approach allows for model evaluation to drop from weeks to a couple of hours.
The leaderboard features a default view ranking models by macro-average WER on English splits of Seed TTS Eval and CV3 Eval. Top performers in this category include hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro. For multilingual performance, the leaderboard allows users to toggle between languages, noting that English performance does not necessarily translate to other languages. Strong multilingual models identified include k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. Additionally, the platform includes a 'Voice Cloning' view that compares speaker similarity and intelligibility when a reference audio is provided, showing improved WER for models like bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 under these conditions.
Why It Matters
The launch of the Open TTS Leaderboard provides a scalable alternative to human-preference arenas, which are limited by voter consistency and the logistical challenges of hosting numerous open-source models. By focusing on objective metrics, the leaderboard enables faster assessment of new releases, potentially democratizing access to high-quality evaluation data for the open-source community. This is particularly relevant for developers building voice agents and interactive applications, where streaming performance is critical. The 'Streaming' tab ranks models by TTFA on both H200 GPUs and CPUs, highlighting kyutai/pocket-tts as a strong performer in both environments.
While the leaderboard does not replace human judgment on naturalness or expressiveness, it offers a complementary layer of analysis. The metrics provide proxies for intelligibility and voice identity preservation, which can inform which models should be included in future voting-based evaluations. Hugging Face emphasizes that the goal is to shape evaluations with community feedback, and the platform includes a 'Listen' tab for users to compare outputs and vote on preferences, bridging the gap between automated metrics and human perception.
The Bottom Line
Hugging Face's Open TTS Leaderboard introduces a standardized, objective framework for evaluating multilingual and voice-cloning TTS models. By leveraging automated metrics for intelligibility, speed, and speaker similarity, the platform addresses the scalability limitations of existing arena-style leaderboards. The company plans to open-source the evaluation scripts soon, inviting community contributions via GitHub to further refine the metrics and datasets.