Trending: On-device modelsSearch
iHeartGeek
iTECH

Hugging Face scores open speech models without asking for votes

Hugging Face has published an Open TTS Leaderboard that scores text-to-speech models on intelligibility, speed and speaker similarity instead of collecting human votes.

A graphic reading Open TTS Leaderboard against a purple background, ringed by national flags and a yellow singing emoji holding a microphone

Hugging Face has published the Open TTS Leaderboard, a ranking for text-to-speech models that scores them on objective measurements rather than on the preferences of whoever happens to be clicking. It went up on 30 September, at a point where the company's own Hub carries more than 8,000 TTS models.

What the leaderboard actually measures

Three things, deliberately kept apart. Intelligibility is measured as word or character error rate between the prompt and a transcript of the generated audio, using Qwen3 ASR, the top open model on Hugging Face's own Open ASR Leaderboard, as the transcriber. Speed is measured two ways: inverse real-time factor for batched offline inference on an H200 GPU, and time-to-first-audio for streaming at batch size one on both GPU and CPU. Speaker similarity uses cosine similarity between WavLM speaker embeddings of the generated audio and the reference clip.

Evaluating a model this way drops from a couple of weeks of vote collection to a couple of hours, and that gap is the whole argument for doing it.

Why the existing arenas were not enough

The obvious comparison is TTS Arena v2 and Artificial Analysis' Voice Arena, which rank models by asking listeners to choose between two outputs and turning the results into an Elo score. Hugging Face is careful with those: human preference is still the final word, and it says the arenas remain useful reference points. The problem is scale. Arenas cannot keep up with the release rate, and open-weight models are under-represented on them, partly because adding a commercial API model takes little more than a key while hosting somebody else's open model takes infrastructure. Of the 92 models on Artificial Analysis, only 16 are open-weights. Voter consistency is the other weakness: no arena can guarantee that the same listeners apply the same standard over time.

Who comes out on top

On English, Kokoro-82M, Supertone's supertonic-3 and Fish Audio's s2-pro lead when word error rate is averaged across the Seed TTS Eval and CV3 Eval splits, and the leaderboard's Pareto plots show which models balance quality against size and batched throughput. Multilingual performance is a different picture: k2-fsa's OmniVoice, s2-pro and FunAudioLLM's Fun-CosyVoice3-0.5B-2512 lead once more languages are toggled on. Switching on voice cloning adds a speaker-similarity column, and some models, including bosonai's higgs-tts-3-4b and OpenBMB's VoxCPM2, score better on average error rate once a reference clip is supplied.

Our opinion

This is the right correction to arena culture, and it is also a clear-eyed reminder of what it cannot do. Word error rate is measured by an ASR model, so a synthesised voice that is flat but enunciates cleanly can beat one that sounds alive and swallows the odd syllable; the leaderboard's own authors say as much, and the Listen tab, where you can hear the outputs behind the numbers and vote, is their admission that the table is not the whole story. The more interesting claim is about openness. Only 16 of 92 models being open-weights is not a quality signal but a hosting-cost signal, and an objective leaderboard sidesteps that by making a model cheap to add. The risk runs the other way once this works: publish metrics as a ranking and speech teams will tune against the metric, and an ASR-scored error rate is exactly the number a research group can chase. Speed and speaker similarity are harder to game, which is why splitting the score into three rather than fusing it into one composite is the best decision in the design.