The number of open-source text-to-speech (TTS) models has grown rapidly. As of September 2026, there are more than 8,000 models on the Hugging Face Hub. However, evaluation methods have not kept pace with this growth.
Traditional arena-based leaderboards compare models by asking users to choose which output sounds better. This process is slow and hard to scale. It also depends on voters' changing preferences over time.
To address this, the Open TTS Leaderboard uses objective metrics. These include word error rate (WER) and character error rate (CER) to measure intelligibility. It also uses cosine similarity to compare speaker identity. These metrics evaluate models quickly and consistently.
The leaderboard ranks models by their average WER on English datasets. It also supports multilingual evaluation, including Chinese, Japanese, and Korean. Models supporting voice cloning can be compared using reference audio. The system shows tradeoffs between accuracy, speed, and size with Pareto plots.
Evaluation with metrics takes only a few hours, compared to weeks for human-based tests. This allows faster development and comparison of models. The leaderboard also offers a 'Listen' tab to explore actual generated outputs, helping users judge naturalness and quality.
The 'Streaming' tab measures how quickly models produce audio after being prompted. It ranks models by time-to-first-audio, which is important for voice agents and interactive applications. Users can provide feedback on outputs to improve future evaluations.
This new approach helps the community evaluate models more efficiently and reliably. It supports better decision-making for deploying speech models in real-world applications.