Text-to-speech (TTS)

Custom TTS (bring your own endpoint) for AI voice agents

Custom TTS (bring your own endpoint) is one of the voices you can run a Telenow AI voice agent on. Point at any streaming-HTTP TTS endpoint. Describe the request body + audio format once; we reframe the audio to telephony μ-law. MVP: streaming HTTP (mulaw / pcm16 / mp3). Give your agent a Custom TTS (bring your own endpoint) voice for natural, streaming speech the caller can't tell from a human. Mix and match it with any LLM, STT, TTS and carrier — Telenow is component-level, so you're never locked in.

Last updated 2026-09-16

Custom TTS (bring your own endpoint) voices

Streaming~300ms0 voices

Frequently asked questions

Can I use a Custom TTS (bring your own endpoint) voice for my AI agent?+

Yes — pick a Custom TTS (bring your own endpoint) voice in Telenow and your agent speaks with it on every call, streaming so there's no awkward pause.

What voices does Custom TTS (bring your own endpoint) offer?+

Custom TTS (bring your own endpoint) offers a range of voices selectable in the builder.

Does Custom TTS (bring your own endpoint) stream speech in real time?+

Yes — Custom TTS (bring your own endpoint) streams audio as it's generated, so the agent starts talking immediately instead of waiting for the whole sentence.

How much does Custom TTS (bring your own endpoint) cost for a voice agent?+

You pay Custom TTS (bring your own endpoint)'s usage at cost plus Telenow's transparent platform fee — billed per component (speech, model, telephony) and per minute, with new accounts getting free signup credit to try it.

Can I bring my own Custom TTS (bring your own endpoint) API key?+

Yes. Telenow supports BYOK — paste your own Custom TTS (bring your own endpoint) key to bill Custom TTS (bring your own endpoint) usage to your account, or use the platform key and pay through Telenow.

Other text-to-speech (tts) options

Telenow TTS
Telenow in-house TTS — 22 Indian languages + English, voice cloning & design; μ-law 8k telephony, 24k HD on web calls.
ElevenLabs
Most natural prosody; great for branded voices.
OpenAI TTS
Predictable cost, 6 stock voices.
Cartesia
Lowest-latency streaming TTS (Sonic); near-realtime, μ-law 8k direct.
xAI (Grok)
Grok streaming TTS — expressive voices + your xAI voice clones, speech tags (laugh, sigh, pause) when expression is on; μ-law 8k telephony, 24k HD on web calls.
Soniox
Soniox real-time TTS — 60+ languages with one voice speaking them all (strong Hindi + code-switching), plus instant cross-lingual voice cloning; μ-law 8k telephony, 24k HD on web, gapless streaming.
Amazon Polly
AWS Polly neural voices — cheap, many languages, μ-law telephony output.
LMNT (Aurora)
Real-time multilingual TTS — low latency, 30+ languages, low cost.
Rime
Conversational US-English with accents/dialects; sub-100ms.
Sarvam (Indian languages)
Bulbul — Hindi + 10 Indian languages & Indian-English accents.
Smallest.ai (Lightning)
Ultra-low-latency TTS — 225+ voices, 30+ languages incl. Indian.
Hume (Octave)
Expressive, emotional TTS. Higher latency — warmth over speed.
UnrealSpeech
Low-cost streaming TTS — Kokoro voices across English (US/UK) & Hindi.
Google Gemini TTS
Gemini 3.1 Flash TTS — 30 voices that each speak 90+ languages, directed in plain English: you write the speaker, the scene and the accent (Indian, American, British…) and the same voice performs it. Inline tags ([whispers], [laughs], [excited]) steer delivery mid-sentence when expression is on. Synthesised a whole turn at a time so the performance never drifts between clauses; μ-law 8k telephony, 24k HD on web.
$1.00 free credit on signup

Build a voice agent with Custom TTS (bring your own endpoint)

Sign up free and get $1.00 in credit — no card required. Connect your number, pick a template, and go live in minutes.