Voice Stack
Voice runtime & latency
Full-duplex ASR, dialog, tool calling, and TTS stacked under a 300ms P50 budget.
Latency budget
Streaming ASR P50 74ms, intent & decision model P50 110ms, neural TTS first audio P50 92ms. Time to first audio token targets 276ms. Barge-in cuts synthesis within 60ms when the caller speaks over the agent.
Turn taking
The runtime uses endpointing plus acoustic barge-in. Dialog state is retained across 20+ turns including objections, language switches, and partial slot fills. Audio is 24kHz HD with OPUS fallback for 2G/3G.
Book a demo or open the operations console to inspect campaigns, TRAI gates, and CRM actions.