Transcription
Models that turn spoken audio into text.
The transcription models behind audio transcription. Each is priced per minute of audio and documents the languages it covers and whether it separates speakers and marks timestamps. Name one directly when a job needs a particular language, diarization, or a lower cost.
10 models
Deepgram Nova-3TranscriptionDeepgram
$0.0058/ min of audio
February 2025
Released
RealtimeLow latency
GPT TranscribeTranscriptionOpenAI
$10.00/ 1M tokens
March 2025
Released
Grok STTTranscriptionX.ai
$0.10/ hour
2026
Released
Low cost
Qwen3 ASR 1.7BTranscriptionQwen· via
DeepInfra
$0.00045/ min
2026
Released
Open sourceMultilingual
Scribe v1TranscriptionElevenLabs
$0.0067/ min
July 2025
Released
Scribe v2TranscriptionElevenLabs
$0.0067/ min
2026
Released
Voxtral Mini 3BTranscriptionMistral· via
DeepInfra
$0.0010/ min
July 2025
Released
Open source
Whisper Large v3TranscriptionOpenAI· via
DeepInfra
$0.00045/ min
November 2023
Released
Open sourceLow cost
Whisper Large v3 TurboTranscriptionOpenAI· via
DeepInfra
$0.00020/ min
October 2024
Released
Open sourceFastLow cost
Whisper-1TranscriptionOpenAI
$0.006/ minute
September 2022
Released
Open sourceLow cost