ElevenLabs shipped two new speech models on September 28, 2026 — and the most instructive thing in the announcement is a footnote.
Every latency figure below is ElevenLabs’ own, measured in September 2026, with network latency measured and removed for all systems — so none of them is a time a user waits, and no real-world latency is estimated here. The two speed figures measure different quantities and are not compared below: about 100ms is a median inference latency for Eleven v4, while about 150ms is a median time to first speech for Eleven v4 Turbo over WebSocket streaming. The #1 ranking is attributed to Artificial Analysis; the roughly 75 per cent blind preference carries no external attribution and is the company’s own test, with no sample size disclosed. The announcement gives no previous language count.
What Happened
On September 28, 2026, ElevenLabs announced Eleven v4 and Eleven v4 Turbo in a post titled “Introducing Eleven v4, our most emotive model.” Both models are available immediately in ElevenAgents, ElevenCreative, and via ElevenAPI. Both support more than 90 languages. The company also describes Instant Voice Clones as capturing voices with high fidelity from just 10 seconds of audio.
The two headline performance figures are a median inference latency of about 100ms for Eleven v4, and a median time to first speech of about 150ms for Eleven v4 Turbo over WebSocket streaming. The announcement attributes a number-one ranking to third-party evaluator Artificial Analysis. Separately — and this distinction matters — ElevenLabs reports that Eleven v4 was preferred by approximately 75% of listeners in blind head-to-head tests over competing models; that preference share is the company’s own test, with no sample size or confidence interval disclosed in the announcement.
The benchmark footnote names the systems tested against: Cartesia Sonic 3.6, xAI TTS, Google Gemini 3.8 Flash-Lite TTS, and OpenAI GPT-4o mini TTS. It specifies identical scripts and default settings, and dates the measurement to September 2026. This publication has not reproduced that test, and takes no position on which system is fastest.
The key insight: ElevenLabs disclosed the benchmark’s most important limit — network latency removed — in the same sentence as the number. That transparency is what makes the figure usable for model comparison and unusable for experience promises. The same number is routinely used for both, and the tension between those two uses is the structural story here.

The Structural Read
The ~100ms inference figure for Eleven v4 measures the model. It does not measure what a user waits for, because every real request crosses a network and the company’s own footnote confirms that network latency was measured and removed for all systems. Nobody experiences a network-removed latency. That is not a criticism — it is a description of what the number is.
The important structural property is that removing the network is the right method for comparing models and the wrong one for sizing live-agent experience — and the same figure tends to do duty for both. A buyer evaluating vendors wants the network out, because it isolates the thing being purchased. An engineer sizing the turn-taking budget for a live voice agent needs the network in, because users hear it. One number, two legitimate uses, one of which it cannot serve.
The ~150ms figure for Eleven v4 Turbo measures something different again — median time to first speech over WebSocket streaming, also with the network removed. The two figures are not the same quantity and are not compared here. Subtracting one from the other or reading one as faster than the other would misrepresent what was actually disclosed.
Map of AI — Enabler Layer
The Infrastructure Disclosure Gap
Most vendor latency claims name no comparator and no date. Naming four systems, a methodology, and a month turns a number into something falsifiable. That is a higher bar than the industry norm — and it also creates the conditions under which limits become visible. The footnote that enables scrutiny is the same one that reveals what the number cannot tell you. That is not a contradiction. It is how credible infrastructure claims should work.
The “default settings” framing in the benchmark is worth noting plainly: defaults differ between vendors, so what is controlled is identical scripts rather than identical configurations. Comparing under each vendor’s own defaults is a defensible method. It is also a consequential choice, and readers evaluating the benchmark should hold both things at once.
ElevenLabs — September 2026
“Median time from request to audible speech. Measured September 2026 with identical scripts and default settings against Cartesia Sonic 3.6, xAI TTS, Google Gemini 3.8 Flash-Lite TTS, and OpenAI GPT-4o mini TTS; network latency measured and removed for all systems.”
The Bottom Line
Eleven v4 and v4 Turbo are real product releases with real availability and a notably transparent benchmark — the company disclosed the network subtraction, named the comparators, and dated the test, which is a higher evidentiary bar than most voice AI claims clear. What the announcement cannot tell you is what a user actually waits for, because that depends on a network figure that was not disclosed. The figure that measures the model and the figure that would govern an experience are not the same figure, and the space between them is where most buyer decisions go wrong.
Source: ElevenLabs — “Introducing Eleven v4, our most emotive model” (September 28, 2026)
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Every latency figure above is ElevenLabs’ own, measured in September 2026, with network latency measured and removed for all systems. None of them is a time a user waits, and no network-inclusive or real-world latency is estimated above because none is disclosed. The two speed figures measure different quantities and are not compared above: about 100ms is a median inference latency for Eleven v4, while about 150ms is a median time to first speech for Eleven v4 Turbo over WebSocket streaming. Nothing above subtracts one from the other or describes either model as slower than the other. The #1 ranking is attributed in the announcement to Artificial Analysis, a third party whose methodology was not seen here. The claim of roughly 75 per cent blind preference carries no external attribution and is therefore the company’s own test, and the number of listeners, how they were recruited and any confidence interval are not disclosed — so the preference share is reported without being weighed. The announcement gives no previous language count, so nothing above describes more than 90 languages as an increase from any particular figure. The comparators named in the footnote — Cartesia Sonic 3.6, xAI TTS, Google Gemini 3.8 Flash-Lite TTS and OpenAI GPT-4o mini TTS — are reported as named. This publication has reproduced none of the testing, ranks nobody against them, and takes no view on which system is fastest. Where the use of each vendor’s default settings is noted as a methodological choice, nothing above suggests that choice was self-serving. Pricing, explicit API model identifiers, the language list, any accuracy or word-error metric, and the hardware or region behind the latency test are not established and do not appear — a limit of one announcement and of this reporting rather than evidence that none exist. Nothing above calls any figure misleading or inflated; disclosing the network subtraction alongside the number is to the company’s credit. Nothing above predicts anything.









