Voice Stack

Voice runtime & latency

Full-duplex ASR, dialog, tool calling, and TTS stacked under a 300ms P50 budget.

Latency budget

Streaming ASR P50 74ms, intent & decision model P50 110ms, neural TTS first audio P50 92ms. Time to first audio token targets 276ms. Barge-in cuts synthesis within 60ms when the caller speaks over the agent.

Turn taking

The runtime uses endpointing plus acoustic barge-in. Dialog state is retained across 20+ turns including objections, language switches, and partial slot fills. Audio is 24kHz HD with OPUS fallback for 2G/3G.

Need a live architecture session?

Book a demo or open the operations console to inspect campaigns, TRAI gates, and CRM actions.

Open Console