LLM modelsBYOK supported

Google Gemini for AI voice agents

Google Gemini is one of the LLMs you can run a Telenow AI voice agent on. Gemini via the Gemini API (OpenAI-compatible). These are the rolling `*-latest` aliases — Google keeps them pointed at the current generation, so they never break when an older model is retired. Flash and Flash-Lite are fast + cheap with huge context (strong defaults for live calls); Pro for the hardest reasoning. Pick a Google Gemini model as the agent's brain and it drives the whole conversation — understanding the caller, deciding what to say, and calling your tools — fast enough to keep a live call natural. Mix and match it with any LLM, STT, TTS and carrier — Telenow is component-level, so you're never locked in.

Last updated 2026-09-16

Google Gemini models for voice

ModelContextInput /MOutput /MFirst tokenTools
Gemini Flash (latest)1049K$0.3$2.5~280ms
Gemini Flash-Lite (latest)1049K$0.1$0.4~220ms
Gemini Pro (latest)1049K$1.25$10~650ms

Frequently asked questions

Can I use Google Gemini for an AI voice agent?+

Yes — Google Gemini is a first-class LLM option in Telenow. Select a Google Gemini model as your agent's brain and it powers real-time voice and chat conversations.

Which Google Gemini models can I use?+

Gemini Flash (latest), Gemini Flash-Lite (latest), and Gemini Pro (latest) — Telenow lists the low-latency Google Gemini models that keep the first reply fast on a live call.

Does Google Gemini support tools / function calling?+

Yes — the listed Google Gemini models support function calling, so your agent can take actions (book, pay, look up, transfer) mid-conversation.

How much does Google Gemini cost for a voice agent?+

You pay Google Gemini's usage at cost plus Telenow's transparent platform fee — billed per component (speech, model, telephony) and per minute, with new accounts getting free signup credit to try it.

Can I bring my own Google Gemini API key?+

Yes. Telenow supports BYOK — paste your own Google Gemini key to bill Google Gemini usage to your account, or use the platform key and pay through Telenow.

Other llm models options

xAI Grok
Grok via the xAI API (OpenAI-compatible). Grok 4 Fast (non-reasoning) is built for low-latency chat with a 2M context; Grok 3 mini is the budget pick. Always-on reasoning Grok 4 is excluded — too slow to first token for live calls.
Anthropic
Strong reasoning and tool use; great for nuanced personas. For live calls prefer Haiku (fastest/cheapest) or Sonnet (balanced) — Opus is higher quality but pricier and slower to first token.
OpenAI
Industry-leading reasoning, broad tool-call support. Only low-latency chat models are listed — slow reasoning models (o1 / o3) are deliberately excluded, they stall the first reply on a live call.
Sarvam AI
Sarvam 105B Conversations via Sarvam's OpenAI-compatible API — India-built, tuned for real-time voice dialogue in the 10 most-spoken Indian languages + English, native script, romanised and code-mixed (Hinglish). Streams, calls tools, caches prompts, never spends time reasoning. Uses the same SARVAM_API_KEY as Sarvam STT/TTS.
Groq
Fastest tokens-per-second on hosted open models.
OpenRouter
One key, hundreds of models through an OpenAI-compatible gateway. Type or paste ANY model ID from openrouter.ai/models (e.g. anthropic/claude-sonnet-4.5) — the suggestions are just popular call-friendly picks; pricing roughly tracks the underlying provider.
Azure OpenAI
OpenAI models hosted in your Azure tenancy. The deployment you configure below selects the model; the chosen model here is used for the cost/latency estimate. Requires endpoint + deployment + API version.
AWS Bedrock
Claude, Llama and Amazon Nova through AWS Bedrock (unified ConverseStream API). Bring your own AWS credentials, or leave them blank to use the platform AWS account.
Custom API (full agentic workflow)
You already built an agentic workflow on your server — tools, RAG, memory, system prompt all live with you. We just attach voice (STT + TTS) and pipe each user turn through your SSE endpoint.
Custom LLM (open-source / self-hosted)
Point at any OpenAI-compatible endpoint — Ollama, vLLM, llama.cpp, LM Studio, or a self-deployed model. We treat it exactly like OpenAI: same chat-completions wire format, your model + your URL. Tools work too.
$1.00 free credit on signup

Build a voice agent with Google Gemini

Sign up free and get $1.00 in credit — no card required. Connect your number, pick a template, and go live in minutes.