01 Who votes
A vetted panel, not the open web
Not whoever shows up. All 224 listeners are paid contributors from joinvoicedata.com who had already cleared onboarding — an approved recording, or an approved application — before any listening work was offered to them. Each one judged only their native language. There is no vote button on this page, so no ranking here moves because a link went around. The usual format is open by design: Hugging Face’s TTS Arena, for one, takes a vote from any account thirty days old.
Native speaker only
Onboarding cleared
Paid per comparison
No public sign-up
9.3%
of the hidden same-clip controls still came back with a confident winner picked — 149 of 1,597. Vetting is not a guarantee, and a vetted panel is still not a representative cross-section of either language.
02 Better at what?
Four questions, not “which is better”
“Better” never says better at what, so a one-question board hands you a single number with the trade-off already averaged out of it. We ask four separate questions and never mix them inside a sitting: overall preference, sounds human, clear and correct, rhythm and expression. In aggregate the four broadly agree. Per voice they do not, and per voice is what you ship — 33 of the 44 English voices move ten places or more depending on which one you read.
Fish Audio s2-pro · one voice, two questions
1st44th
Sounds human
#5
Clear and correct
#40