Aller au contenu
Updated after each benchmark capture

The TTS benchmark for real-time voice agents Real-time text-to-speech (TTS) APIs, ranked on measured latency, accuracy, voice quality and price.

Every text-to-speech API measured by three independent benchmarks (Coval, Speko and Artificial Analysis), ranked on P50 and P99 time-to-first-audio, latency consistency, accuracy, voice quality and price.

Data checked on September 30, 2026; Coval 30-day window ending October 1, 2026.

How models are scored

100 points, 4 criteria

  • Real-time speed 60
  • Voice quality 20
  • Accuracy 15
  • Price 5
Read the methodology

112

text-to-speech models benchmarked

44

providers

44

with measured time-to-first-audio

3

independent benchmarks: Coval, Speko, Artificial Analysis

TTS benchmark 2026

Top 5 TTS APIs for voice agents

See all 112 models →

Best TTS by use case

Top rated in each Artificial Analysis listening category, plus voice agents: models whose P90 time-to-first-audio on Coval is 200 ms or less.

Multilingual TTS: languages tested

Languages in which Artificial Analysis or Speko actually tested each model. A model missing from a language was not tested in it.

Frequently asked questions

Choosing a TTS API for a voice agent

What is the best TTS for voice agents in 2026?
In this benchmark, Gradium TTS Beta ranks first of 112 with 76.8/100. On Coval's 30-day window ending October 1, 2026, its P50 time-to-first-audio is 47.2 ms and its P99 is 108.8 ms. The next models are Alibaba Qwen3 TTS Fast (73.7), Soniox TTS Real-Time v2 (73.7), Inworld Realtime TTS-2 (68.8).
How is this text-to-speech benchmark built?
It combines three independent benchmarks: Coval (continuous latency and word error rate measurements over 30 days), Speko (latency, naturalness, robustness and cost) and Artificial Analysis (blind listener preference in its Speech Arena, and price). Each model is scored on real-time speed (60), voice quality (20), accuracy (15), price (5), out of 100. A measurement a benchmark does not publish counts zero, and the model page says why.
What is time-to-first-audio (TTFA), and why P50 and P99?
Time-to-first-audio is the wait between sending text to a TTS API and hearing the first syllable. For a voice agent it decides whether a reply feels instant. The P50 (median) is the typical turn; the P99 is the slowest 1% of turns, the ones a caller notices; the spread between runs is latency consistency.
Which TTS API has the lowest latency?
On Coval's 30-day window ending October 1, 2026, the lowest P50 time-to-first-audio is 47.2 ms, for Gradium TTS Beta. The full speed ranking, on P50, P99 and consistency only, is on the real-time page.
Which TTS models are open source?
19 models in the benchmark publish their weights (open weights), according to Artificial Analysis, the Coval registry or the creator's own release. They are listed on the open-weights page, with the same measurements as the proprietary APIs.

Follow the TTS benchmark

An email when the scores change or new TTS models are benchmarked.

Benchmark updates only. Unsubscribe at any time. See our privacy policy.