Text to speech
Models that turn text into spoken audio, with a choice of named voices.
The speech models behind text to speech. Each is priced per million characters and ships a set of named voices you can audition on its page. Some run at real-time latency for live use, others favor quality. Choose a model and a voice to match the tone and speed a moment calls for.
12 models
ElevenLabs TTSText to speechElevenLabs
~$0.20/ 1k characters
—
Voices
Realtime
Latency
Low latency
Cartesia Sonic 3Text to speechCartesia
~$37/ 1M characters
—
Voices
Realtime
Latency
RealtimeLow latency
Gemini 3.1 Flash TTSText to speechGoogle
$20.00/ 1M tokens
30
Voices
Standard
Latency
Audition 30 voices
GPT-4o-mini TTSText to speechOpenAI
$2.40/ 1M tokens
9
Voices
Standard
Latency
Audition 9 voices
Grok Voice TTSText to speechX.ai
$15/ 1M characters
28
Voices
Realtime
Latency
Audition 28 voices
Kokoro 82MText to speechHexgrad· via
DeepInfra
$0.62/ 1M chars
54
Voices
Standard
Latency
Audition 54 voices
Minimax Speech 2.8 HDText to speechMiniMax· via
WaveSpeed
$0.007/ run
17
Voices
Standard
Latency
Audition 17 voices
Orpheus 3BText to speechCanopy Labs· via
DeepInfra
$7.00/ 1M chars
7
Voices
Standard
Latency
Audition 7 voices
Qwen Audio 3.0 TTS PlusText to speechQwen· via
Alibaba Cloud
$27.59/ 1M characters
2
Voices
Standard
Latency
Audition 2 voices
Qwen3 TTSText to speechQwen· via
DeepInfra
$20.00/ 1M chars
9
Voices
Standard
Latency
Audition 9 voices
TTS HDText to speechOpenAI
$30.00/ 1M characters
6
Voices
Standard
Latency
Audition 6 voices
TTS-1Text to speechOpenAI
product: {5.00/ 1M characters
6
Voices
Standard
Latency
Audition 6 voices