Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

2026-10-02 · Hugging Face

Background and Challenges

As of September 30, 2026, there are over 8,000 text-to-speech (TTS) models available on the Hugging Face Hub. However, evaluation has not kept pace, remaining fragmented and unstandardized. While human preference scores like MOS or MUSHRA are the gold standard, arena-based leaderboards (such as TTS Arena v2, Artificial Analysis, and Voice Arena) use Elo scores computed from user votes. These arenas face several limitations:

  • Scalability: Human voting cannot keep up with the rapid release of TTS models. Collecting sufficient votes takes weeks.
  • Underrepresentation of Open-Source Models: On Artificial Analysis, only 16 of 92 models are open-weights. Adding API models requires just an API key, whereas open models must be hosted by the arena operator. Commercial providers also have more incentive to seek placement.
  • Voter Consistency: No arena can ensure that voters maintain the same criteria for "better" over time, as individual preferences naturally change.

Objective Metrics of the Open TTS Leaderboard

To address these issues, the Open TTS Leaderboard was created. It uses objective metrics to evaluate models on complementary aspects of performance, reducing evaluation time from weeks to hours:

  • Intelligibility: Measures word/character error rate (WER and CER) between the prompt and the generated audio's transcript using Qwen3 ASR, a top-ranking open-source ASR model.
  • Speed: Evaluates the inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for streaming batch size 1 latency on an H200 GPU and CPU.
  • Speaker Similarity: Computes the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.

Importantly, this leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, and speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference, but they can inform voting-based leaderboards on which models to include.

Multilingual and Voice Cloning Evaluation

By default, models are ranked by macro-average WER on the English splits of Seed TTS Eval and CV3 Eval. Models like Kokoro-82M, supertonic-3, and s2-pro lead in English WER.

The leaderboard supports toggling multiple languages to assess multilingual performance. Since Seed TTS Eval only contains English and Chinese audio, other languages are scored solely on CV3 Eval. For character-based languages like Chinese, Japanese, and Korean, character error rate (CER) is reported. Top multilingual models include OmniVoice, s2-pro, and Fun-CosyVoice3-0.5B-2512.

Toggling "Voice cloning" allows users to compare models supporting this functionality. A SIM column appears in the table, along with Pareto plots visualizing tradeoffs between SIM, batched inference, and size. Some models, such as higgs-tts-3-4b and VoxCPM2, show improved average WER under voice cloning when a reference audio is provided.

Listen and Streaming Performance

  • Listen: Numbers only tell part of the story. The "Listen" tab lets users compare generated outputs behind the metrics. Users can pick a language/dataset, toggle voice cloning, and optionally select models or listen to random outputs. Users can also vote on the outputs, helping the community filter out spam and bots.
  • Streaming: The "Streaming" tab ranks models by TTFA, quantifying how long a user waits to receive playable audio after probing a model. This metric is critical for voice agents and interactive applications.

Source