Voice Arena Q2 2026 Model Assessment
Summary of findings
Voice Arena conducted an independent third-party assessment of gpt-4o-mini-tts within the Q2 2026 multi-language TTS evaluation. This document presents the seven findings Voice Arena considers most material to the OpenAI TTS roadmap, ranked by the magnitude of the signal in the data rather than by their alignment with any prior hypothesis. Each finding is supported by (a) the relevant aggregate chart, (b) one or two audio samples Voice Arena recommends the OpenAI research team listen to directly, and (c) verbatim rater commentary translated where necessary. Only rater quotes where the underlying pattern is corroborated by at least three independent raters across at least two sub-buckets are included; single-rater observations have been excluded.
The good news is that the failure modes are concentrated, structural, and addressable. Five of the seven findings below trace to the same two root causes: (a) a frontend that treats disfluencies, numbers, and mixed-script tokens as lexical content, and (b) acoustic-model coverage that is shallow outside Latin-script European languages. Fixing them does not require a new architecture — it requires the training-data and evaluation rigor already applied to English to be replicated for Hindi, Vietnamese, Arabic, and the conversational domain.
Voice Arena recommends prioritizing the work in the order shown. Findings 2 and 6 are P0 — they are the largest blockers to any real-time-agent product surface (Voice Mode, AVM, the upcoming agent SDK). Finding 4 is the largest blocker to non-English market expansion. Everything else can be sequenced behind these.
OpenAI gpt-4o-mini-tts is mid-pack at best, and last in two languages
Elo rankings put gpt-4o-mini-tts at #6/7 in English, #8/8 in Hindi, #7/7 in Vietnamese, and #4 in Arabic and Portuguese. Its single bright spot is Japanese (#3). The gap to Gemini 3.1 Flash TTS — the #1 model in every single language — ranges from −123 to −248 Elo points. For context, that is a larger gap than separates GPT-4 from GPT-3.5 on MMLU.
Audio anchor
“robotic voice”· en-US rater — Billing lookup — vs. Gemini 3.1 Flash
“missed a word and the 'hmm' breaks the flow”· en-US rater — Concert recap — vs. Gemini 3.1 Flash
Robotic-voice is OpenAI's single largest defect — and it is causally tied to losing
Across all 12,692 vote-appearances, the unnatural_robotic_voice tag fires on gpt-4o-mini-tts at 31.8% — 4× Gemini (8.0%), 4.3× ElevenLabs v3 (7.4%). Only Azure Dragon HD is worse, and only by 0.3 points. The outcome-conditioned cut is the smoking gun: 14% on wins, 22% on ties, 53.9% on losses. When raters chose OpenAI, they did not hear it; when they rejected it, they almost always did. This is not a perception artifact — it is causal.
What raters actually said (losing battles)
“robotic voice”· en-US rater — Hospital PA
“Turned robot-like voice when it got to 'umm, please reply' and did not recover”· en-US rater — Ringback / IVR opt-in
“The audio is clear but flat and monotone making it sound robotic”· en-US rater — News commodities
“A sounds like a robot”· en-US rater — Refund initiated
Audio anchor
“robotic voice”· vi-VN rater — Commercial license PA — vs. ElevenLabs v3
“voice sounds a bit robotic and lacks natural expression”· vi-VN rater — Documentary intro — vs. MiniMax 2.8 HD
“Giọng không tự nhiên”(EN: “Voice not natural.”)· vi-VN rater — AI navigation directive — vs. Cartesia Sonic-3
“B audio AI hai or robotic voice hai accent bhaut galat hai.”(EN: “Audio B is AI and has a robotic voice; the accent is very wrong.”)· hi-IN rater — Regulatory compliance PA — vs. Gemini 3.1 Flash
“Sounds like English person speaking Hindi”· hi-IN rater — Commercial license PA — vs. Cartesia Sonic-3
“robotic voice”· en-US rater — Hospital PA — vs. Fish Audio S2 Pro
“audio clear but flat and monotone making it sound robotic”· en-US rater — News commodities — vs. xAI Grok
“audio is clear but monotone and flat so it sounds unnatural. The pace is slow”· en-US rater — Content creation brand-voice promotion — vs. Cartesia Sonic-3
“A sounds like a robot”· en-US rater — Refund initiated — vs. Azure Dragon HD
gpt-4o-mini-tts mishandles disfluencies in every language tested
When the input sentence contains a disfluency token (umm, uh, éto, áh, ahn, &c.) OpenAI's win-rate drops in every one of the six languages. The gap is largest in Arabic (−12.7 pts), Japanese (−9.3), and English (−5.3). Raters are unambiguous about why — gpt-4o-mini-tts renders fillers as lexical content (stretched, over-emphasized, articulated) rather than as paralinguistic noise. Competing models drop fillers into prosodic shadow; OpenAI says them.
What raters actually said
“Everything is good except for the 'umm' being over-emphasized.”· en-US rater — Gen-conv meeting-up coordination — vs. Gemini
“The audio stretches out the 'Uhhhhh' at the beginning a little too long which is throwing off the pacing.”· en-US rater — Telecom prepaid plan expiring — vs. Azure Dragon HD
“The pacing is a little off at the beginning when saying the filler words 'umm' and 'uh'. There are awkward pauses between them.”· en-US rater — Generic agent stall (banking) — vs. Fish Audio S2 Pro
Audio anchor
“The filler words are not natural”· vi-VN rater — Refund initiated — vs. Cartesia Sonic-3
“pacing is not good. The filler words are not natural”· vi-VN rater — Brand-voice promotion — vs. xAI Grok
“Artificial e pausas desnecessárias”(EN: “Artificial, and unnecessary pauses.”)· pt-BR rater — Utilities service disruption — vs. xAI Grok
“The interjection 'hum' is mispronunciated”· pt-BR rater — Music play directive — vs. Cartesia Sonic-3
“pacing of the audio is a little off at the beginning when saying the filler words 'umm' and 'uh'. There are awkward pauses in between them”· en-US rater — Generic agent stall (banking) — vs. Fish Audio S2 Pro
“audio stretches out the 'Uhhhhh' at the beginning a little to long”· en-US rater — Telecom prepaid expiring — vs. Azure Dragon HD
“Everything is good except for teh 'umm' being over-emphasized”· en-US rater — Gen-conv meeting-up coordination — vs. Gemini 3.1 Flash
“Turned robot like voice when it got to 'umm, please reply' and di not recover”· en-US rater — Ringback / IVR opt-in — vs. Gemini 3.1 Flash
Non-English mispronunciation is a frontend problem, not an acoustic one
Severe-mispronunciation rates: Arabic 21.3% (Gemini 4.9%, 4.4×), Hindi 12.5% (Gemini 3.7%, 3.4×), Japanese 10.7% (Gemini 5.1%, 2.1×), Vietnamese 9.5% (Gemini 1.5%, 6.3×). The errors cluster around the same five things, regardless of language: numbers, alphanumeric IDs (HVAC-TKY-4351), mixed-script tokens (Hindi-English code-mix in hi_cs_30: "account number JIO-44X2"), foreign proper nouns, and currency. The same frontend pattern surfaces in US English on multi-digit reservation numbers and brand-name prosody — see Sample D4 below. This is a tokenizer / G2P / text-normalization problem in the frontend — not an acoustic-model problem.
Audio anchor
“at timestamp 0:06, '1.660.000 đồng' is read incorrectly to '1 nghìn .660'”· vi-VN rater — Refund initiated — vs. Gemini 3.1 Flash
“27.999.000 đọc thành 27 nghìn 999 đồng là sai nghĩa hoàn toàn”(EN: “Reads 27,999,000 as 27 thousand 999, completely wrong meaning.”)· vi-VN rater — Gen-Z YouTube tech review — vs. Gemini 3.1 Flash
“Bot should pronounce 2.500.000 like 'hai triệu năm trăm nghìn đồng'”· vi-VN rater — Financial transfer — vs. Azure Dragon HD
“added extra numbers in: 4471-8829-3052-0967”· ar-MSA rater — Banking duplicate-charge dispute — vs. Gemini 3.1 Flash
“Extra Content / Missing Word / Unnatural Robotic Voice”· ar-MSA rater — Card activation confirmation (card 4147-9938-2210-6655) — vs. Cartesia Sonic-3
“Unnatural mild mispronounciation”· hi-IN rater — Card activation confirmation — vs. Azure Dragon HD
“Incorrect card number said card 4147-9938-2210-6655”· hi-IN rater — Card activation confirmation — vs. Sarvam Bulbul-V3
“Audio B is Severe Mispronunciation / Unnatural / Robotic Voice / Hallucination / Extra Content”· hi-IN rater — Card activation confirmation — vs. Cartesia Sonic-3
“'Microsoft' is pronounced with what sounds to be a speech impediment. There is also a noticeable lo-fi type static”· en-US rater — News equity quote — vs. Fish Audio S2 Pro
“Missing a digit when saying the reservation number”· en-US rater — AI assistant reservation status — vs. Azure Dragon HD
“Mispronounced the reservation number. 7714 was read as 7 1 4”· en-US rater — AI assistant reservation status — vs. xAI Grok (same sentence, second independent rater, different opponent)
“Mispronounced 11:30”· hi-IN rater — AI assistant scheduling confirmation — vs. MiniMax Speech 2.8 HD
“Audio A Mispronunce 11:30”· hi-IN rater — AI assistant scheduling confirmation — vs. ElevenLabs Eleven v3
“A has issue with mild mispronciation like 11:30”· hi-IN rater — AI assistant scheduling confirmation — vs. xAI Grok TTS
“The audio mispronounces the service ID number HVAC-CHI-4351. For the letter CHI, the model makes it the word "chi" instead of saying the letters individually.”· en-US rater — Customer-support service booking ID — vs. Microsoft Azure Dragon HD Omni
“At 0:05, the voice says, "Chi- I-I" instead of "C-H-I" like printed.”· en-US rater — Customer-support service booking ID — vs. Cartesia Sonic-3
Four failure clusters explain most of the lost ground
When Voice Arena bucketed OpenAI's worst sub-categories by content type, the same four clusters appeared across every language: (1) refund / payment / card-activation confirmations, (2) phone numbers, account IDs, and regulatory PA reads, (3) disfluent agent dialog (call transfer, recharge upsell, banking stall), and (4) code-mixed education (Hindi STEM, English-Hindi psychology). These four clusters are the entire conversational-AI product surface — and they are precisely where gpt-4o-mini-tts systematically underperforms.
Audio anchor
“Artificial e pausas desnecessárias”(EN: “Artificial, and unnecessary pauses.”)· pt-BR rater — Utilities service disruption — vs. xAI Grok
“audio is clear but monotone and flat so it sounds unnatural. The pace is slow”· en-US rater — Content creation brand-voice promotion (en-US cluster overlap) — vs. Cartesia Sonic-3
“स्व-शासन is mispronounced. Pacing is slow.”(EN: “'Self-governance' is mispronounced. Pacing is slow.”)· hi-IN rater — Education humanities history — vs. Azure Dragon HD
Hindi and Vietnamese show the largest gaps versus Gemini
In Hindi and Vietnamese, Gemini wins 80–85% of head-to-head matchups against gpt-4o-mini-tts. OpenAI's robotic-voice tag rate is above 44% in both languages, and the severe-mispronunciation rate in Hindi is 3.4× Gemini's. Hindi raters note that English code-mixed tokens are read in a disconnected register, and Vietnamese raters flag tonal errors on lexical tone pairs. Both are large speaker bases (Hindi ~600M, Vietnamese ~85M), so the gap has material reach.
Audio anchor
“Sounds like English person speaking Hindi”· hi-IN rater — Commercial license PA — vs. Cartesia Sonic-3
“voice is robotic not indian foreign voice”· hi-IN rater — Education STEM chemistry (HCl / NaCl) — vs. xAI Grok
“Sounds like foreign accent”· hi-IN rater — Flight reschedule — vs. ElevenLabs v3
“Giọng đọc ở đoạn này nghe nhiều chỗ phát âm không tự nhiên và cũng không có độ truyền cảm”(EN: “The narrating voice in this passage sounds unnatural in many places and lacks expressiveness.”)· vi-VN rater — Audiobook mythological — vs. Cartesia Sonic-3
“voice from the audio is not as a native speaker. It looks like coming from a foreigner who's learning Vietnamese”· vi-VN rater — Audiobook mythological (same sentence, different opponent) — vs. Gemini 3.1 Flash
“Ngữ điệu không giống người bản xứ”(EN: “Intonation does not sound like a native speaker.”)· vi-VN rater — News commodities — vs. ElevenLabs v3
Where OpenAI wins reveals what it optimized for — narrator, not agent
The top 12 sub-buckets for gpt-4o-mini-tts are almost entirely long-form Japanese formal speech (weather warnings 87%, physics-law recitation 85%, civics long-form 79%, regulatory PA 79%, legal/court comparison 80%) plus structured Portuguese AI-assistant transcripts (transit status, reservation lookup, number-portability). These are the audio domains where the input is clean, the register is formal, the cadence is slow, and there is no rater expectation of spontaneity. OpenAI is exceptional at narrator mode and mediocre at agent mode — and the product roadmap is moving in the opposite direction.
Audio anchor
“大変素晴らしい”(EN: “Very impressive / excellent.”)· ja-JP rater — News severe weather warning — vs. ElevenLabs v3
“問題なし”(EN: “No issues.”)· ja-JP rater — News severe weather warning — vs. Azure Dragon HD
“افضل”(EN: “Better.”)· ar-MSA rater — Education STEM chemistry — vs. Cartesia Sonic-3
“Good”· ar-MSA rater — Education STEM physics (Newton's law F=ma) — vs. MiniMax 2.8 HD
“audio is clear and natural sounding. The pacing is good and pauses are in appropriate places. It pronounces Avogadro correctly”· en-US rater — Education STEM physics (Avogadro constant) — vs. Azure Dragon HD
“Very good, professional voice, good pacing, no pronunciation errors”· en-US rater — News newspaper coverage — vs. Cartesia Sonic-3
“This is good and consistent. It sounds like a textbook”· en-US rater — Civics long-form — vs. Cartesia Sonic-3
Audio sample index
All 23 unique sample cards used above, one place to scan the entire audio evidence base.
Recommendations
The following is a roadmap-ready prioritization. Each item references the finding(s) it addresses and an honest read on effort, ordered by ratio of strategic-impact to engineering-cost.
Disfluency-aware acoustic modeling
Multilingual frontend overhaul
Hindi & Vietnamese training-data acquisition
Conversational-agent eval set, internal
Closing note from Voice Arena Research
Voice Arena wants to be precise about what this assessment is and is not. It is not a claim that pairwise human preference is the only valid evaluation — the methodology has known biases (length, formality, novelty), and 100 sentences per language is a small sample. It is a claim that the patterns in this dataset are too consistent, too cross-lingual, and too tightly coupled to specific frontend and acoustic-modeling decisions to be dismissed as noise. Gemini 3.1 Flash TTS is shipping on a measurably different curve.
Prepared for OpenAI · Q2 2026 · Confidential