Compare AI voice agents, runtimes, and contact-center platforms by the details that change a launch: telephony, latency, tools, transfer behavior, controls, and cost units.
Programmable Voice component that connects a TwiML ConversationRelay session to an application WebSocket for STT/TTS, language settings, DTMF, and call events.
Telephony voice-agent runtime
Scope and signals
Best fit: Developers building their own voice-agent logic while keeping telephony, media handling, and call control in Twilio.
Voice stack: Cascaded speech interface with developer-supplied LLM
Google's Gemini Live API streams multimodal sessions over WebSockets; its stable native voice-agent models are Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
Native realtime multimodal model API
Scope and signals
Best fit: Developers building low-latency conversational agents that need native audio, multimodal input, and function calling in a persistent WebSocket session.
Voice stack: Native audio agents: Gemini 3.8 Live (gemini-3.8-live) and Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking)., Separate TTS models: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts generate speech; neither is a Live agent., Separate transcription endpoints: gemini-3.5-transcribe handles audio files; gemini-3.5-transcribe-live is a Live API speech-to-text mode, not a conversational agent., Separate translation mode: gemini-3.5-live-translate-preview is a Live API interpreter for 70+ languages, not a conversational agent; it is preview.
Standard paid rates are $0.005/min equivalent for input audio and $0.018/min equivalent for output audio, billed separately plus text tokens; Google also lists a free API tier.
Two agents can sound equally fluent while differing in carrier coverage, interruption behavior, tool permissions, and what a human inherits after transfer. Keep those decisions separate and testable.
Conversation surface
Phone, browser, SIP, and messaging each introduce different audio and routing constraints.
Runtime behavior
Check latency, barge-in, turn-taking, retries, and how the system handles an unknown.
Tool boundary
Trace every lookup or write, with identity, permission, confirmation, and failure recovery.
Evidence trail
Keep recordings, transcripts, traces, retention rules, and cost assumptions in the pilot record.
Native realtime
Keep more of the conversation intact.
Native audio models avoid mandatory speech-to-text and text-to-speech handoffs, preserving more pace, pauses, emphasis, interruption, and nonverbal cues. Compare the complete phone path—including carrier, tools, and transfer—with the same real call scenarios.
Test recognition, speech quality, interruptions, language switching, tool use, and human transfer with your own calls in the required language. Availability can vary by model, plan, and phone route.
DocsBot documents multilingual voice workflows and Japanese support. Verify the exact voice, phone path, and transfer behavior in your pilot. Review DocsBot language information ↗
Other providers’ language claims reflect cited vendor material; confirm Japanese and other required languages directly with each provider.
Before you decide
Choosing an AI voice agent
What is an AI voice agent?
An AI voice agent holds a live spoken conversation, uses approved knowledge, can call permitted tools, and may transfer to a person. Products range from packaged receptionists to enterprise contact-center systems, orchestration APIs, and low-level realtime model runtimes.
How should I choose an AI voice agent?
Start with one call job, the systems it must use, the action it may take, and a safe human fallback. Run the same real call scenarios on each candidate and compare task completion, tool accuracy, interruption recovery, transfer, operational work, and all-in cost.
What is native speech-to-speech?
A native speech-to-speech or realtime audio model accepts and produces audio directly, preserving more tone, pacing, interruption, and nonverbal context. A cascaded stack connects speech recognition, a text language model, and speech synthesis for more modular control.
What should I compare first?
Start with the call job, required channels, business tools, transfer path, safety controls, and all-in cost. Then run the same representative calls with each finalist.
Your business. Your shortlist.
Find your best-fit AI voice system.
Start with your website. Answer a few focused questions, then get a personalized PDF with fit, trade-offs, alternatives and a pilot plan for your call flow.