Speech-to-text (STT)

Speech-to-text (STT) for voice AI

Real-time transcription that turns the caller’s speech into text the agent can act on.

Supported providers

Telenow Smart (language routing)
One STT that routes each language to the best recognizer — an always-on multilingual anchor plus per-language specialists, picked live by audio language-ID. Routing (anchor + specialists + margin) is configured centrally under Admin → STT routing, so there are no per-agent settings.
Deepgram
Nova-3 multilingual (real-time code-switching); sub-300 ms partials.
Sarvam (Saaras)
India-native ASR — up to 22 Indic languages + English, Hinglish code-mixing, transliteration & translation, streaming.
ElevenLabs Scribe v2 Realtime
Streaming multilingual STT (90+ languages), ~150 ms first partials.
AWS Transcribe
Streaming Hindi + Indian-English with en/hi code-switch; Mumbai region for data residency.
Soniox
Real-time multilingual STT — strong Hindi + code-switching, token-streaming.
xAI (Grok)
Grok streaming STT — 25 languages, ~300 ms partials; same key as the Grok LLM.
OpenAI Whisper
Higher accuracy, non-streaming — better for batch.
Groq (Whisper-Large)
Whisper-large running on Groq — much faster than OpenAI Whisper.
Google Gemini STT
Gemini 3.5 Transcribe (live) — streaming partials over the Gemini Live API, 85+ locales, automatic language detection that code-switches inside a sentence with no configuration. Indian locales except Tamil. Same key as the Gemini LLM.
Custom STT (bring your own endpoint)
Point at any real-time WebSocket STT endpoint. Describe your vendor’s wire format once as a JSON descriptor and we map it to transcripts — no code. MVP: WebSocket transport.
$1.00 free credit on signup

Build a voice agent with speech-to-text (stt)

Sign up free and get $1.00 in credit — no card required. Connect your number, pick a template, and go live in minutes.