Listeners hear two AI voices read the same sentence and pick one, never seeing who made either. Two things make this board different from the usual arena: who votes, and what they were asked.
95% rangescorebottom of leader’s rangeclearly behind the leaderaverage voice (1000)Latestthe company’s newest model here — older generations sit unmarked
01 Who votes
A vetted panel, not the open web
Not whoever shows up. All 224 listeners are paid contributors from joinvoicedata.com who had already cleared onboarding — an approved recording, or an approved application — before any listening work was offered to them. Each one judged only their native language. There is no vote button on this page, so no ranking here moves because a link went around. The usual format is open by design: Hugging Face’s TTS Arena, for one, takes a vote from any account thirty days old.
Native speaker onlyOnboarding clearedPaid per comparisonNo public sign-up
9.3%of the hidden same-clip controls still came back with a confident winner picked — 149 of 1,597. Vetting is not a guarantee, and a vetted panel is still not a representative cross-section of either language.
02 Better at what?
Four questions, not “which is better”
“Better” never says better at what, so a one-question board hands you a single number with the trade-off already averaged out of it. We ask four separate questions and never mix them inside a sitting: overall preference, sounds human, clear and correct, rhythm and expression. In aggregate the four broadly agree. Per voice they do not, and per voice is what you ship — 33 of the 44 English voices move ten places or more depending on which one you read.
Fish Audio s2-pro · one voice, two questions
1st44th
Sounds human#5
Clear and correct#40
Same clips, same listeners, same week. Switch the Question filter above to move between these boards.
Scores come from a standard head-to-head method called Bradley–Terry, centred so the average voice in each language sits at 1000. A listener answers one of the four questions for a whole sitting of 25 comparisons and the questions are never mixed inside a sitting, so the 95% ranges are worked out by resampling listeners rather than individual clicks, because one person voting 25 times is not 25 independent opinions. Results are frozen once voting closes; we do not publish a moving scoreboard while people are still voting.