Voice Stack
ASR, NLU, dialog & TTS
Streaming recognition, entity extraction, dialog policy, and neural synthesis with Indian dialects.
ASR
Acoustic models are trained for Indian dialects, Hinglish/Tanglish code-switching, and noisy telephony. Partial hypotheses stream over WebSocket for live captions.
NLU & dialog
POST /v1/conversations then POST /v1/conversations/:id/turns runs deterministic NLU in English, Hinglish, Hindi (Devanagari), Tamil, and Telugu for six domains. POST /v1/conversations/:id/audio accepts JSON { wavBase64 } or multipart audio/file/wav; WHISPER_URL or WHISPER_CMD transcribes when set, otherwise fixture ASR, then the same NLU/dialog path. If OPENAI_API_KEY is set the same path may enrich slots; missing key stays deterministic. Each turn response includes processingMs. First agent turn prepends the tenant agent disclosure from GET/PUT /v1/agents. TTS remains Web Speech in the browser or env-gated vendor TTS — not a live neural clone without keys.
TTS
POST /v1/conversations/:id/tts returns audio/wav. On Darwin the runtime uses `say` (WAVE LEI16 16 kHz) unless SHRIS_TTS_BACKEND=fixture. espeak-ng/espeak is used when present. Otherwise a short sine WAV is returned (not a neural clone). The marketing widget still uses Web Speech in the browser as a fallback. Vendor TTS_API_KEY remains optional live mode — never invent a production voice-clone key.
Book a demo or open the operations console to inspect campaigns, TRAI gates, and CRM actions.