Voice ArenaVoiceArena.comThe evals layer of Voice AI
Prepared for OpenAIQ2 2026
Voice Arena · Prepared for OpenAI · Confidential · 31 May 2026

Voice Arena Q2 2026 Model Assessment

Findings, evidence, and prioritized recommendations for gpt-4o-mini-tts, based on 44,387 approved pairwise human votes collected across six languages in May 2026.

Summary of findings

Voice Arena conducted an independent third-party assessment of gpt-4o-mini-tts within the Q2 2026 multi-language TTS evaluation. This document presents the seven findings Voice Arena considers most material to the OpenAI TTS roadmap, ranked by the magnitude of the signal in the data rather than by their alignment with any prior hypothesis. Each finding is supported by (a) the relevant aggregate chart, (b) one or two audio samples Voice Arena recommends the OpenAI research team listen to directly, and (c) verbatim rater commentary translated where necessary. Only rater quotes where the underlying pattern is corroborated by at least three independent raters across at least two sub-buckets are included; single-rater observations have been excluded.

#6/7
English Elo rank
975 — behind Grok & Eleven
#8/8
Hindi Elo rank
−238 vs. Gemini
#7/7
Vietnamese Elo rank
−248 vs. Gemini
31.8%
Robotic-voice defect rate
vs. Gemini 8.0%
53.9%
Robotic rate on lost battles
vs. 14.0% on wins

The good news is that the failure modes are concentrated, structural, and addressable. Five of the seven findings below trace to the same two root causes: (a) a frontend that treats disfluencies, numbers, and mixed-script tokens as lexical content, and (b) acoustic-model coverage that is shallow outside Latin-script European languages. Fixing them does not require a new architecture — it requires the training-data and evaluation rigor already applied to English to be replicated for Hindi, Vietnamese, Arabic, and the conversational domain.

Voice Arena recommends prioritizing the work in the order shown. Findings 2 and 6 are P0 — they are the largest blockers to any real-time-agent product surface (Voice Mode, AVM, the upcoming agent SDK). Finding 4 is the largest blocker to non-English market expansion. Everything else can be sequenced behind these.

Finding 1

OpenAI gpt-4o-mini-tts is mid-pack at best, and last in two languages

Elo rankings put gpt-4o-mini-tts at #6/7 in English, #8/8 in Hindi, #7/7 in Vietnamese, and #4 in Arabic and Portuguese. Its single bright spot is Japanese (#3). The gap to Gemini 3.1 Flash TTS — the #1 model in every single language — ranges from −123 to −248 Elo points. For context, that is a larger gap than separates GPT-4 from GPT-3.5 on MMLU.

English
Hindi
Vietnamese
Arabic
Portuguese (BR)
Japanese
Source: Voice Arena Q2 2026 leaderboard. Elo computed from 44,387 approved pairwise votes; Bradley–Terry with 200-resample bootstrap CIs, anchored at 1000 (per-language median). Methodology: voicearena.com/tts-methodology.

Audio anchor

HindiGen-conv — music play directive[A1]
हम्म, 'Heeramandi' का जो song viral हुआ था वो था, उम्म.. Raja Hasan का 'सकल बन फूल रही सरसों'।
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: Hindi/English code-switch and proper-noun handling.
VietnameseAI assistant — transit status[A2]
Tàu Thống Nhất SE3 đang chạy chậm 30 phút và sẽ đến ga Hà Nội lúc 9 giờ 15 tối.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: Tonal accuracy on Vietnamese proper nouns and route codes.
US EnglishCustomer support — billing lookup (disfluent)[A3]
Hmm, I'm really sorry, ma'am, but, uh, I actually can't find any electricity bill payment history for account number CG-554-3320 for January.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: OpenAI's worst-performing US English sentence in the dataset: lost 11 of 20 head-to-heads, and lost every one of its 4 matchups against Gemini 3.1 Flash TTS (the #1 US English model). Compare the robotic, flat delivery of the disfluencies ("hmm", "uh") and the alphanumeric account number against Gemini's natural agent-on-the-phone register.
robotic voice
· en-US rater — Billing lookup — vs. Gemini 3.1 Flash
US EnglishGen-conv — concert recap (long, disfluent)[A4]
Bro, Tame Impala's Coachella set this year was absolutely insane — the entire crowd was glowing with phone lights the second 'The Less I Know the Better' dropped, and then, hmm, Kevin Parker just stood there grinning while like sixty thousand people screamed every single word right back at him.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: Second-worst US English sentence: 10 losses in 15 battles, with all 3 head-to-heads against Gemini 3.1 Flash TTS lost. Listen for the missing-word defect and the over-articulated "hmm" mid-sentence in the OpenAI render versus Gemini's casual, in-character storytelling cadence.
missed a word and the 'hmm' breaks the flow
· en-US rater — Concert recap — vs. Gemini 3.1 Flash
Finding 2
P0

Robotic-voice is OpenAI's single largest defect — and it is causally tied to losing

Across all 12,692 vote-appearances, the unnatural_robotic_voice tag fires on gpt-4o-mini-tts at 31.8% — 4× Gemini (8.0%), 4.3× ElevenLabs v3 (7.4%). Only Azure Dragon HD is worse, and only by 0.3 points. The outcome-conditioned cut is the smoking gun: 14% on wins, 22% on ties, 53.9% on losses. When raters chose OpenAI, they did not hear it; when they rejected it, they almost always did. This is not a perception artifact — it is causal.

Robotic-voice tag rate by model (pooled)
OpenAI rate, conditioned on outcome
Left: pooled across all six languages. Right: OpenAI defect rate conditioned on rater verdict.

What raters actually said (losing battles)

robotic voice
· en-US rater — Hospital PA
Turned robot-like voice when it got to 'umm, please reply' and did not recover
· en-US rater — Ringback / IVR opt-in
The audio is clear but flat and monotone making it sound robotic
· en-US rater — News commodities
A sounds like a robot
· en-US rater — Refund initiated

Audio anchor

VietnameseCommercial license PA[B1]
Các chủ doanh nghiệp có giấy phép kinh doanh số BIZ-HAN-2026-W4-882 phải gia hạn giấy phép trước ngày 31 tháng 3 năm 2026, để tránh bị phạt chậm và gián đoạn hoạt động. Việc gia hạn có thể thực hiện trực tuyến qua portal Sở Kế Hoạch Đầu Tư, hoặc trực tiếp tại bất kỳ trung tâm hành chính công nào được ủy quyền. Việc không gia hạn không chỉ dừng hoạt động kinh doanh, mà còn có thể dẫn đến các chế tài tài chính bổ sung và các đợt kiểm toán bắt buộc trong tương lai.
OpenAI gpt-4o-mini-tts
—:—
ElevenLabs Eleven v3
—:—
Listen for: Whether the voice sounds like a person or a recording. The unnatural_robotic_voice defect class fires heavily on this card.
robotic voice
· vi-VN rater — Commercial license PA — vs. ElevenLabs v3
voice sounds a bit robotic and lacks natural expression
· vi-VN rater — Documentary intro — vs. MiniMax 2.8 HD
Giọng không tự nhiên(EN: “Voice not natural.”)
· vi-VN rater — AI navigation directive — vs. Cartesia Sonic-3
HindiRegulatory compliance PA[B2]
Jio recharge subscribers जिनके account number JIO-44X2 है, उन्हें service suspension से बचने के लिए 30 जून तक identity re-verification complete करनी होगी।
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: How JIO-44X2 is read and how the overall PA register lands relative to Gemini's rendering.
B audio AI hai or robotic voice hai accent bhaut galat hai.(EN: “Audio B is AI and has a robotic voice; the accent is very wrong.”)
· hi-IN rater — Regulatory compliance PA — vs. Gemini 3.1 Flash
Sounds like English person speaking Hindi
· hi-IN rater — Commercial license PA — vs. Cartesia Sonic-3
US EnglishHospital PA[B3]
Patient Eleanor Whitaker, please report to Tower B, Room 314 for your Cardiac MRI; please ensure you have the referral form from your primary care provider at Atlas Health Group.
OpenAI gpt-4o-mini-tts
—:—
Fish Audio S2 Pro
—:—
Listen for: How the formal PA cadence lands relative to a peer model. En-US raters describe the gpt-4o-mini-tts rendering as flat, monotone, or distant on cards that should sit inside the model's narrator-mode strength zone.
robotic voice
· en-US rater — Hospital PA — vs. Fish Audio S2 Pro
audio clear but flat and monotone making it sound robotic
· en-US rater — News commodities — vs. xAI Grok
audio is clear but monotone and flat so it sounds unnatural. The pace is slow
· en-US rater — Content creation brand-voice promotion — vs. Cartesia Sonic-3
A sounds like a robot
· en-US rater — Refund initiated — vs. Azure Dragon HD
Finding 3

gpt-4o-mini-tts mishandles disfluencies in every language tested

When the input sentence contains a disfluency token (umm, uh, éto, áh, ahn, &c.) OpenAI's win-rate drops in every one of the six languages. The gap is largest in Arabic (−12.7 pts), Japanese (−9.3), and English (−5.3). Raters are unambiguous about why — gpt-4o-mini-tts renders fillers as lexical content (stretched, over-emphasized, articulated) rather than as paralinguistic noise. Competing models drop fillers into prosodic shadow; OpenAI says them.

OpenAI win-rate on clean vs. disfluent sentences. Disfluency detected via regex over language-specific filler vocabulary in sentence_text.

What raters actually said

Everything is good except for the 'umm' being over-emphasized.
· en-US rater — Gen-conv meeting-up coordination — vs. Gemini
The audio stretches out the 'Uhhhhh' at the beginning a little too long which is throwing off the pacing.
· en-US rater — Telecom prepaid plan expiring — vs. Azure Dragon HD
The pacing is a little off at the beginning when saying the filler words 'umm' and 'uh'. There are awkward pauses between them.
· en-US rater — Generic agent stall (banking) — vs. Fish Audio S2 Pro

Audio anchor

VietnameseRefund initiated[C1]
Dạ, ờ, vâng, khoản hoàn tiền 1.660.000 đồng sẽ được hoàn lại vào tài khoản của anh trong vòng 3 ngày làm việc.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: The opening filler stack Dạ, ờ, vâng and the rendering of the refund amount 1.660.000.
The filler words are not natural
· vi-VN rater — Refund initiated — vs. Cartesia Sonic-3
pacing is not good. The filler words are not natural
· vi-VN rater — Brand-voice promotion — vs. xAI Grok
Portuguese (BR)Utilities — service disruption[C2]
Ah, infelizmente, hum, o fornecimento de água no bairro de Pinheiros vai ser interrompido no dia 31 de março de 2026, principalmente devido a, ah, obras de reparo no encanamento principal perto do Largo da Batata.
OpenAI gpt-4o-mini-tts
—:—
xAI Grok TTS
—:—
Listen for: Phrasing around interjections and the overall naturalness of the agent register.
Artificial e pausas desnecessárias(EN: “Artificial, and unnecessary pauses.”)
· pt-BR rater — Utilities service disruption — vs. xAI Grok
The interjection 'hum' is mispronunciated
· pt-BR rater — Music play directive — vs. Cartesia Sonic-3
US EnglishGeneric agent stall (banking)[C3]
Umm, uh, give me just a second here, yeah, I'm actually pulling up your latest transaction details right now.
OpenAI gpt-4o-mini-tts
—:—
Fish Audio S2 Pro
—:—
Listen for: The opening Umm, uh and the pacing across the stalled clause. gpt-4o-mini-tts stretches the fillers into prosodic dead air; the comparison renderings drop the same tokens into sub-second breaths.
pacing of the audio is a little off at the beginning when saying the filler words 'umm' and 'uh'. There are awkward pauses in between them
· en-US rater — Generic agent stall (banking) — vs. Fish Audio S2 Pro
audio stretches out the 'Uhhhhh' at the beginning a little to long
· en-US rater — Telecom prepaid expiring — vs. Azure Dragon HD
Everything is good except for teh 'umm' being over-emphasized
· en-US rater — Gen-conv meeting-up coordination — vs. Gemini 3.1 Flash
Turned robot like voice when it got to 'umm, please reply' and di not recover
· en-US rater — Ringback / IVR opt-in — vs. Gemini 3.1 Flash
Finding 4
P1

Non-English mispronunciation is a frontend problem, not an acoustic one

Severe-mispronunciation rates: Arabic 21.3% (Gemini 4.9%, 4.4×), Hindi 12.5% (Gemini 3.7%, 3.4×), Japanese 10.7% (Gemini 5.1%, 2.1×), Vietnamese 9.5% (Gemini 1.5%, 6.3×). The errors cluster around the same five things, regardless of language: numbers, alphanumeric IDs (HVAC-TKY-4351), mixed-script tokens (Hindi-English code-mix in hi_cs_30: "account number JIO-44X2"), foreign proper nouns, and currency. The same frontend pattern surfaces in US English on multi-digit reservation numbers and brand-name prosody — see Sample D4 below. This is a tokenizer / G2P / text-normalization problem in the frontend — not an acoustic-model problem.

Tag rate = severe_mispronunciation occurrences per model appearance. n ≈ 1700–2700 per cell.

Audio anchor

VietnameseRefund initiated (numeric)[D1]
Dạ, ờ, vâng, khoản hoàn tiền 1.660.000 đồng sẽ được hoàn lại vào tài khoản của anh trong vòng 3 ngày làm việc.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: The number 1.660.000. Vietnamese uses period as a thousands separator; gpt-4o-mini-tts reads 1.660.000 đồng as '1 nghìn .660', changing a 1.66 M VND refund into ~1,660 VND. High-severity production bug for any payments or refund agent.
at timestamp 0:06, '1.660.000 đồng' is read incorrectly to '1 nghìn .660'
· vi-VN rater — Refund initiated — vs. Gemini 3.1 Flash
27.999.000 đọc thành 27 nghìn 999 đồng là sai nghĩa hoàn toàn(EN: “Reads 27,999,000 as 27 thousand 999, completely wrong meaning.”)
· vi-VN rater — Gen-Z YouTube tech review — vs. Gemini 3.1 Flash
Bot should pronounce 2.500.000 like 'hai triệu năm trăm nghìn đồng'
· vi-VN rater — Financial transfer — vs. Azure Dragon HD
ArabicBanking duplicate-charge dispute[D2]
حسناً، أه، دعوني أراجع هذا للحظة… أم، نعم، أرى هنا أن قسط القرض السكني البالغ 6,250 درهماً على الحساب رقم 4471-8829-3052-0967 قد تم خصمه مرتين بتاريخ 14 يناير 2026.
OpenAI gpt-4o-mini-tts
—:—
Gemini 3.1 Flash TTS
—:—
Listen for: The card-number readout. gpt-4o-mini-tts inserts an extra digit not present in the input — a hallucination_extra_content event combined with severe mispronunciation.
added extra numbers in: 4471-8829-3052-0967
· ar-MSA rater — Banking duplicate-charge dispute — vs. Gemini 3.1 Flash
Extra Content / Missing Word / Unnatural Robotic Voice
· ar-MSA rater — Card activation confirmation (card 4147-9938-2210-6655) — vs. Cartesia Sonic-3
HindiCard activation confirmation[D3]
ठीक है, उम्म, हाँ, आपका card 4147-9938-2210-6655 अब active हो गया है।
OpenAI gpt-4o-mini-tts
—:—
Microsoft Azure Dragon HD Omni
—:—
Listen for: The 16-digit card number rendered inside Hindi prosody, plus the corroborating Hindi number / formula failures across other sub-buckets.
Unnatural mild mispronounciation
· hi-IN rater — Card activation confirmation — vs. Azure Dragon HD
Incorrect card number said card 4147-9938-2210-6655
· hi-IN rater — Card activation confirmation — vs. Sarvam Bulbul-V3
Audio B is Severe Mispronunciation / Unnatural / Robotic Voice / Hallucination / Extra Content
· hi-IN rater — Card activation confirmation — vs. Cartesia Sonic-3
US EnglishNews equity quote[D4]
Microsoft shares fell 1.4 percent to $412 on Tuesday after the company reported weaker-than-expected cloud growth.
OpenAI gpt-4o-mini-tts
—:—
Fish Audio S2 Pro
—:—
Listen for: The rendering of 'Microsoft' at the start of the sentence, and the same frontend failure surfacing on numeric reservation strings (XCQ-7714).
'Microsoft' is pronounced with what sounds to be a speech impediment. There is also a noticeable lo-fi type static
· en-US rater — News equity quote — vs. Fish Audio S2 Pro
Missing a digit when saying the reservation number
· en-US rater — AI assistant reservation status — vs. Azure Dragon HD
Mispronounced the reservation number. 7714 was read as 7 1 4
· en-US rater — AI assistant reservation status — vs. xAI Grok (same sentence, second independent rater, different opponent)
HindiAI assistant scheduling confirmation[D5]
मैंने कौमुदी ज्ञानेश्वरी के साथ Project Atlas पर meeting कल सुबह 11:30 बजे schedule कर दी है।
OpenAI gpt-4o-mini-tts
—:—
MiniMax Speech 2.8 HD
—:—
Listen for: How the model renders the time token '11:30' inside a Hindi-English code-mixed scheduling utterance — raters across multiple opponents independently flagged the same number as mispronounced.
Mispronounced 11:30
· hi-IN rater — AI assistant scheduling confirmation — vs. MiniMax Speech 2.8 HD
Audio A Mispronunce 11:30
· hi-IN rater — AI assistant scheduling confirmation — vs. ElevenLabs Eleven v3
A has issue with mild mispronciation like 11:30
· hi-IN rater — AI assistant scheduling confirmation — vs. xAI Grok TTS
US EnglishCustomer-support service booking ID[D6]
Okay, uh, yeah, your service booking ID HVAC-CHI-4351 is now, umm, confirmed.
OpenAI gpt-4o-mini-tts
—:—
Microsoft Azure Dragon HD Omni
—:—
Listen for: How the model reads the alphanumeric service ID 'HVAC-CHI-4351'. Four independent en-US raters, across four different opponents, all flagged the same failure: the 'CHI' segment is read as the word 'chi' instead of the letters C-H-I.
The audio mispronounces the service ID number HVAC-CHI-4351. For the letter CHI, the model makes it the word "chi" instead of saying the letters individually.
· en-US rater — Customer-support service booking ID — vs. Microsoft Azure Dragon HD Omni
At 0:05, the voice says, "Chi- I-I" instead of "C-H-I" like printed.
· en-US rater — Customer-support service booking ID — vs. Cartesia Sonic-3
Finding 5

Four failure clusters explain most of the lost ground

When Voice Arena bucketed OpenAI's worst sub-categories by content type, the same four clusters appeared across every language: (1) refund / payment / card-activation confirmations, (2) phone numbers, account IDs, and regulatory PA reads, (3) disfluent agent dialog (call transfer, recharge upsell, banking stall), and (4) code-mixed education (Hindi STEM, English-Hindi psychology). These four clusters are the entire conversational-AI product surface — and they are precisely where gpt-4o-mini-tts systematically underperforms.

C1 Refund/Payment C2 Numbers/IDs/PA C3 Disfluent agent C4 Code-mixed eduDashed line: parity 50%
Sub-buckets with n ≥ 10 and OpenAI win-rate < 35%. Dashed line at 50% denotes parity.

Audio anchor

Portuguese (BR)Utilities — service disruption[E1]
Ah, infelizmente, hum, o fornecimento de água no bairro de Pinheiros vai ser interrompido no dia 31 de março de 2026, principalmente devido a, ah, obras de reparo no encanamento principal perto do Largo da Batata.
OpenAI gpt-4o-mini-tts
—:—
xAI Grok TTS
—:—
Listen for: Cluster representative for the Portuguese disfluency / over-articulation pattern; cross-related to C2.
Artificial e pausas desnecessárias(EN: “Artificial, and unnecessary pauses.”)
· pt-BR rater — Utilities service disruption — vs. xAI Grok
audio is clear but monotone and flat so it sounds unnatural. The pace is slow
· en-US rater — Content creation brand-voice promotion (en-US cluster overlap) — vs. Cartesia Sonic-3
↑ See sample [D1] above — Cluster 2 representative (Numbers/IDs/Regulatory PA)↑ See sample [B1] above — Cluster 3 representative (Disfluent agent dialog)
HindiEducation humanities history[E4]
भारतीय स्वतंत्रता संग्राम 1947 में अपने चरम पर पहुँचा और दक्षिण एशिया की राजनीतिक व्यवस्था को रूपांतरित किया, लोकतंत्र और स्व-शासन के लिए भविष्य के आंदोलनों को प्रभावित करते हुए।
OpenAI gpt-4o-mini-tts
—:—
Microsoft Azure Dragon HD Omni
—:—
Listen for: Long-form Hindi narrator-register card with explicit pacing + pronunciation feedback.
स्व-शासन is mispronounced. Pacing is slow.(EN: “'Self-governance' is mispronounced. Pacing is slow.”)
· hi-IN rater — Education humanities history — vs. Azure Dragon HD
Finding 6
P0

Hindi and Vietnamese show the largest gaps versus Gemini

In Hindi and Vietnamese, Gemini wins 80–85% of head-to-head matchups against gpt-4o-mini-tts. OpenAI's robotic-voice tag rate is above 44% in both languages, and the severe-mispronunciation rate in Hindi is 3.4× Gemini's. Hindi raters note that English code-mixed tokens are read in a disconnected register, and Vietnamese raters flag tonal errors on lexical tone pairs. Both are large speaker bases (Hindi ~600M, Vietnamese ~85M), so the gap has material reach.

Gemini head-to-head win share vs. OpenAI
Hindi
80.2% Gemini
Vietnamese
85.1% Gemini
Robotic-voice rate
Severe mispronunciation
Three views of the HI/VI gap. Win-rate is the share of matchups won; robotic and severe-mispronunciation are per-appearance tag rates.

Audio anchor

HindiCommercial license PA[F1]
License number BIZ-MUM-2026-W4-882 वाले व्यवसाय मालिकों को शुक्रवार तक नगर निगम कार्यालय Window 5-B पर अपने permits renew कराने हैं।
OpenAI gpt-4o-mini-tts
—:—
Cartesia Sonic-3
—:—
Listen for: How the formal Hindi register lands when the model is asked to deliver a civic-announcement cadence — native Hindi raters consistently flag the prosody as English-accented.
Sounds like English person speaking Hindi
· hi-IN rater — Commercial license PA — vs. Cartesia Sonic-3
voice is robotic not indian foreign voice
· hi-IN rater — Education STEM chemistry (HCl / NaCl) — vs. xAI Grok
Sounds like foreign accent
· hi-IN rater — Flight reschedule — vs. ElevenLabs v3
VietnameseAudiobook mythological[F2]
An Dương Vương đứng trên mép bờ đá lởm chởm, nhìn về phía bắc qua mặt sông Hồng sẫm màu. Số phận đang gọi vua trở về quê hương, về với thành Cổ Loa xa xôi, và về với công chúa Mỵ Châu đã chờ đợi cha từ rất lâu. Cô chờ một cái ôm cuối cùng để khép lại bi kịch và hàn gắn những vết thương của sự phản bội đã đè nặng lên dòng dõi.
OpenAI gpt-4o-mini-tts
—:—
Cartesia Sonic-3
—:—
Listen for: Lexical tone realization and narrative expressiveness across sustained Vietnamese prose. The same sentence picks up two parallel native-speaker observations across different opponent models, both centered on lack of native intonation and expressiveness.
Giọng đọc ở đoạn này nghe nhiều chỗ phát âm không tự nhiên và cũng không có độ truyền cảm(EN: “The narrating voice in this passage sounds unnatural in many places and lacks expressiveness.”)
· vi-VN rater — Audiobook mythological — vs. Cartesia Sonic-3
voice from the audio is not as a native speaker. It looks like coming from a foreigner who's learning Vietnamese
· vi-VN rater — Audiobook mythological (same sentence, different opponent) — vs. Gemini 3.1 Flash
Ngữ điệu không giống người bản xứ(EN: “Intonation does not sound like a native speaker.”)
· vi-VN rater — News commodities — vs. ElevenLabs v3
Note: en-US is not directly applicable to this finding by construction — Finding 6 is defined by the Hindi and Vietnamese gap. The en-US evidence for the underlying failure modes (robotic voice, disfluency, mispronunciation) appears in Findings 2, 3, and 4 above.
Finding 7

Where OpenAI wins reveals what it optimized for — narrator, not agent

The top 12 sub-buckets for gpt-4o-mini-tts are almost entirely long-form Japanese formal speech (weather warnings 87%, physics-law recitation 85%, civics long-form 79%, regulatory PA 79%, legal/court comparison 80%) plus structured Portuguese AI-assistant transcripts (transit status, reservation lookup, number-portability). These are the audio domains where the input is clean, the register is formal, the cadence is slow, and there is no rater expectation of spontaneity. OpenAI is exceptional at narrator mode and mediocre at agent mode — and the product roadmap is moving in the opposite direction.

Narrator/News STEM education Regulatory/PA AI-assistant transcript
Top sub-buckets where OpenAI wins, color-coded by content type. n is sample size per sub-bucket.

Audio anchor

JapaneseNews severe weather warning[G1]
気象庁は、関東甲信越地方一帯にきわめて激しい雨の警報を発表し、低地での都市型水害への厳重な警戒を呼びかけており、毎時80ミリを超える降水量が、特に夕方の通勤時間帯にピークを迎えると予想されています。
OpenAI gpt-4o-mini-tts
—:—
Microsoft Azure Dragon HD Omni
—:—
Listen for: Formal Japanese weather-warning register. Same sentence picks up two unambiguous Japanese wins against two different opponents.
大変素晴らしい(EN: “Very impressive / excellent.”)
· ja-JP rater — News severe weather warning — vs. ElevenLabs v3
問題なし(EN: “No issues.”)
· ja-JP rater — News severe weather warning — vs. Azure Dragon HD
ArabicEducation STEM chemistry[G2]
في علم الكيمياء، يتفاعل المركّب حمض الهيدروكلوريك ذو الصيغة HCl مع بيكربونات الصوديوم لإنتاج كلوريد الصوديوم بالصيغة NaCl، إلى جانب إطلاق ثاني أكسيد الكربون والماء عبر تفاعل متحكم به.
OpenAI gpt-4o-mini-tts
—:—
Cartesia Sonic-3
—:—
Listen for: Long-form formal MSA on a technical passage — Arabic raters call this the better rendering relative to peer models in the same sub-bucket.
افضل(EN: “Better.”)
· ar-MSA rater — Education STEM chemistry — vs. Cartesia Sonic-3
Good
· ar-MSA rater — Education STEM physics (Newton's law F=ma) — vs. MiniMax 2.8 HD
US EnglishEducation STEM physics (Avogadro)[G3]
According to atomic theory, Avogadro's number, with value 6.022 times 10 to the power 23 per mole, defines how many elementary particles make up one mole of any substance and forms the bridge between the atomic and macroscopic scales.
OpenAI gpt-4o-mini-tts
—:—
Microsoft Azure Dragon HD Omni
—:—
Listen for: Cadence on a long-form formal-register technical passage and explicit correct rendering of the named entity 'Avogadro'. This is the exact pattern where gpt-4o-mini-tts is strongest in every language Voice Arena evaluated.
audio is clear and natural sounding. The pacing is good and pauses are in appropriate places. It pronounces Avogadro correctly
· en-US rater — Education STEM physics (Avogadro constant) — vs. Azure Dragon HD
Very good, professional voice, good pacing, no pronunciation errors
· en-US rater — News newspaper coverage — vs. Cartesia Sonic-3
This is good and consistent. It sounds like a textbook
· en-US rater — Civics long-form — vs. Cartesia Sonic-3

Audio sample index

All 23 unique sample cards used above, one place to scan the entire audio evidence base.

Tag
sentence_id
Language
Sub-bucket
Used in
[A1]
hi_gc_14
Hindi
Gen-conv — music play directive
F1
[A2]
vi_gc_03
Vietnamese
AI assistant — transit status
F1
[A3]
us_cs_17
US English
Customer support — billing lookup (disfluent)
F1
[A4]
us_gc_17
US English
Gen-conv — concert recap (long, disfluent)
F1
[B1]
vi_cs_32
Vietnamese
Commercial license PA
F2
[B2]
hi_cs_30
Hindi
Regulatory compliance PA
F2
[B3]
us_cs_26
US English
Hospital PA
F2
[C1]
vi_cs_08
Vietnamese
Refund initiated
F3
[C2]
br_cs_20
Portuguese (BR)
Utilities — service disruption
F3
[C3]
us_cs_01
US English
Generic agent stall (banking)
F3
[D1]
vi_cs_08
Vietnamese
Refund initiated (numeric)
F4
[D2]
ar_cs_02
Arabic
Banking duplicate-charge dispute
F4
[D3]
hi_cs_05
Hindi
Card activation confirmation
F4
[D4]
us_me_14
US English
News equity quote
F4
[D5]
hi_gc_04
Hindi
AI assistant scheduling confirmation
F4
[D6]
us_cs_16
US English
Customer-support service booking ID
F4
[E1]
br_cs_20
Portuguese (BR)
Utilities — service disruption
F5(C1)
[E4]
hi_ce_17
Hindi
Education humanities history
F5(C4)
[F1]
hi_cs_32
Hindi
Commercial license PA
F6
[F2]
vi_me_06
Vietnamese
Audiobook mythological
F6
[G1]
jp_me_19
Japanese
News severe weather warning
F7
[G2]
ar_ce_12
Arabic
Education STEM chemistry
F7
[G3]
us_ce_13
US English
Education STEM physics (Avogadro)
F7

Recommendations

The following is a roadmap-ready prioritization. Each item references the finding(s) it addresses and an honest read on effort, ordered by ratio of strategic-impact to engineering-cost.

P0

Disfluency-aware acoustic modeling

Treat umm / uh / éto / ah / dá as paralinguistic events, not lexical tokens. Highest ROI fix — it attacks Findings 2, 3, and the conversational sub-buckets in Finding 5 simultaneously. Mechanism: add a filler-token vocabulary to the frontend and route it through a separate prosody head; fine-tune on conversational data where fillers are silent/elided.
Addresses: Findings 2 · 3 · 5
P0

Multilingual frontend overhaul

Rebuild G2P and text normalization for Hindi, Vietnamese, Arabic, and Japanese. Special-case alphanumeric IDs, currency, dates, mixed-script tokens, and foreign proper nouns. Gemini 3.1 Flash TTS handles these without special prompting; gpt-4o-mini-tts does not.
Addresses: Findings 4 · 5
P0

Hindi & Vietnamese training-data acquisition

Both languages need a 10× expansion of native-speaker training audio in conversational and agent registers.
Addresses: Finding 6
P1

Conversational-agent eval set, internal

Mirror Voice Arena's four-cluster taxonomy as an internal CI eval: every gpt-4o-mini-tts checkpoint should report win-rate on refund-confirm, alphanumeric-PA, disfluent-dialog, and code-mixed-education sub-buckets before shipping. This is the eval gap that allowed the regression observed in Finding 1.
Addresses: Findings 1 · 5 · 7

Closing note from Voice Arena Research

Voice Arena wants to be precise about what this assessment is and is not. It is not a claim that pairwise human preference is the only valid evaluation — the methodology has known biases (length, formality, novelty), and 100 sentences per language is a small sample. It is a claim that the patterns in this dataset are too consistent, too cross-lingual, and too tightly coupled to specific frontend and acoustic-modeling decisions to be dismissed as noise. Gemini 3.1 Flash TTS is shipping on a measurably different curve.

Voice Arena Research
Prepared for OpenAI · Q2 2026 · Confidential