Open TTS Leaderboard splits English rank from the license
The Open TTS Leaderboard is Hugging Face's automatic ranking for open speech models, and a second newsroom read the licenses the scoreboard does not settle. English rank and permission to ship are not the same column.

in this block
The Open TTS Leaderboard went up on Hugging Face on Sept. 30, 2026, as an automatic scoreboard for open text-to-speech models. The launch post by Eric Bezzam, Steven Zheng, and Eustache Le Bihan says the Hub held more than 8,000 TTS models that day, and that human arenas cannot keep up. Creative AI News, also dated Sept. 30, read the board's tables and the license file on each model. The company post explains the metrics. The newsroom is where the rank numbers and the commercial-use column come from.
TL;DR - Hugging Face ranks models on intelligibility, speed, and speaker similarity, and it says those scores do not replace a human preference for naturalness. Creative AI News says the board it pulled covered 31 models. - The launch post names hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro as the English word-error leaders, without printing the scores. Creative AI News prints Kokoro at 1.46 WER, Supertonic 3 at 1.54, and Fish Audio S2 Pro at 1.58. - On the nine-language views Creative AI News tabulated, the first and second models are noncommercial. It says the first model licensed for commercial use in those views is Qwen3-TTS at third.
What actually happened
The launch post is an argument about evaluation, not a model release. Arenas such as TTS Arena v2, Artificial Analysis, and Voice Arena ask people to pick one clip over another and then compute an Elo-style score, often with a Bradley-Terry model. Hugging Face says that takes long enough that arenas fall behind the release pace. It says only 16 of 92 models on Artificial Analysis were open weights that day, with a similar skew on Voice Arena, because an API model needs a key and an open model must be hosted.
The Open TTS Leaderboard answers with three automatic measures. Intelligibility is word error rate, or character error rate for Chinese, Japanese, and Korean, between the prompt and a transcript of the audio made with Qwen3 ASR, which the post calls the top open model on the Open ASR Leaderboard. Speed is inverse real-time factor for batched offline inference on an H200, plus time-to-first-audio for streaming at batch size 1 on an H200 and on CPU. Speaker similarity is cosine similarity of WavLM embeddings between the generated audio and a reference clip. The Open TTS Leaderboard post says a model can be scored in hours rather than the weeks a vote takes.
The default rank is macro-average word error rate on the English splits of Seed TTS Eval and CV3 Eval in a zero-shot setting. The post says Kokoro-82M, Supertonic-3, and Fish Audio S2 Pro lead that English pack, and that English does not carry over to other languages. Seed TTS Eval has audio for English and Chinese, so other languages are the CV3 Eval zero-shot score. The cross-language "average WER" is a macro-average, with CER used for Chinese, Japanese, and Korean. The post names k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. Voice cloning is a toggle. A similarity column and extra plots appear when it is on. The post says some models, including bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improve on average error when a reference clip is provided.
A Listen tab plays the clips and can collect votes later. A Streaming tab ranks time-to-first-audio: the first chunk for a streaming model, or the whole utterance if playback cannot start sooner. The post says each model runs one audio at a time on the same 50 English prompts, drops three warm-up runs, and reports the median. It calls out kyutai/pocket-tts on GPU and CPU. Scripts were not open yet.
What the license read changes
Creative AI News says it pulled tables through the Space API for results version 30-09-2026, then opened license files. The board's license column says "custom" for 8 of 29 English models. The same article's methods note says it read the license field on 32 model cards, so 31 and 32 both appear on that page. Its English top three, from that version, are Kokoro-82M at 1.46 WER and 71.15 RTFx under Apache 2.0, Supertonic 3 at 1.54 and 47.32 RTFx under Open RAIL-M, and Fish Audio S2 Pro at 1.58 and 0.4 RTFx, at 4.95B parameters, with no commercial use unless a separate license is signed. Tenth on that list, VibeVoice-Realtime 0.5B, is 1.82. It also says the board listed Fun-CosyVoice3 as CC-BY-4.0 while the model card says Apache 2.0.
Creative AI News says the English top ten spans only 1.46 to 1.82, so the ASR judge's own mistakes sit inside the gap. Kokoro does not clone a voice. The newsroom, not the launch post, says the Kokoro repo was last updated in April 2025 and downloaded about 11.5 million times in the past month.
The multilingual table is the one that changes the headline. With all nine languages ticked, Creative AI News says only 9 of 31 models have a score in every language, so only those nine are ranked. On nine-language default voice it prints OmniVoice at 3.25, noncommercial; Fish Audio S2 Pro at 3.68, noncommercial; Qwen3-TTS 1.7B CustomVoice at 4.11, Apache 2.0. On nine-language cloning it prints OmniVoice at 3.52 error and 74.77 similarity, Higgs TTS 3 at 3.74 and 69.08, and Qwen3-TTS 1.7B Base at 3.92 and 72.92. On English cloning, Higgs TTS 3 leads error at 1.59 with similarity 64.69, then OmniVoice, then Qwen3-TTS. In each of those views the newsroom says the first commercially usable model is Qwen3-TTS at third.
It sorts the English models into commercial, restricted, and noncommercial buckets. The Fish Audio Research License needs a written agreement for commercial use. OmniVoice's code is Apache 2.0 and its weights are CC-BY-NC, which the newsroom ties to training data including Emilia. Higgs TTS 3's creator grant, as it reads section II-A, allows audio on a channel you own if you credit Boson AI, and does not allow hosting the model for other people. VoxCPM2, Apache 2.0, had the highest nine-language cloning similarity it recorded, 75.19.
What a speed test does not prove
Creative AI News reran pocket-tts and Kokoro on a 6-vCPU Ryzen 7 9700X container, PyTorch 2.14.1 CPU, batch size 1, three warm-up sentences and 47 timed ones. Its 6-thread medians were 69.3 milliseconds and 5.02 RTFx for pocket-tts, against 366.5 milliseconds and 11.36 RTFx for Kokoro. The board's CPU column, as it prints it, was 272.7 milliseconds and 1.19 RTFx for pocket-tts, and 913.4 milliseconds and 6.21 for Kokoro. The order matched. The times did not. It says a desktop-class Ryzen is faster than a shared cloud vCPU, so trust the ratio more than the milliseconds. The same Qwen3-TTS weights, on the board's Streaming tab, hit first audio in 18.3 milliseconds with Moondream's Photon engine on an H200 and in 309.9 milliseconds with the faster-qwen3-tts package. The launch post does not print those rows.
The Open TTS Leaderboard is not a listening test. Both pages say so. Word error and speaker similarity do not measure whether a voice sounds good. Scripts were still planned, not published, when Creative AI News checked, and it says the results dataset returned 401, so it used the Space API instead. For a different speech release aimed at agents, see the MAI voice note. For a small model split that is not speech, see the Solar Mini 4 note.
What to do as a reader (not a trade)
This is not investment advice. There is no token in either page. If you use the Open TTS Leaderboard to pick a checkpoint, write two columns that the newsroom already had to build by hand. One is the view: English, nine languages, or cloning. The other is the license on that repository, not the word "custom" in the table. A 1.46 and a 1.58 are close enough, on the newsroom's own warning, that the ASR judge can move the order.
Pin the results version. Creative AI News used 30-09-2026 and said later updates will move rows. Pin the server if you care about time-to-first-audio. Then listen. The launch post's Listen tab is the piece that still requires a person.
The primary page is the Hugging Face launch post. The second page is Creative AI News' license and CPU read, which is the one that separates a noncommercial weight file from an Apache sibling and a board number from a license.
Not financial advice. DYOR, ser.