Platform
HA operations runbook
Two-replica API/web, probes, PDBs, HPA, optional Redis, and the NOC health strip.
Topology
Kubernetes manifests in deploy/k8s/shris.yaml run api and web at 2 replicas with HTTP readiness/liveness on /v1/health and /. PodDisruptionBudgets keep minAvailable 1 during drains. HorizontalPodAutoscalers scale api 2–8 and web 2–6 on 70% CPU. This is a real HA scaffold, not a 99.9% SLA claim until a production cluster and carrier keys exist.
Local compose
From the workspace root: docker compose up --build. Shared Redis rate-limit: REDIS_URL=redis://redis:6379 docker compose --profile ha up. Local Asterisk ARI lab: docker compose --profile sip up --build (lab credentials in lab/asterisk/README.md — not TRAI CPaaS). Leave DATABASE_URL empty in a local .env to stay on the JSON store; compose sets Postgres for the api service.
NOC checks
GET /v1/health is public and returns status, mode, traiOpen, istHour, tenant, queueDepth, and store kind. GET /v1/integrations is admin JWT and lists provider modes (simulator|live|unconfigured|lab) from env presence — never secret values. Probe local uptime with: cd Shris-ai-app && npm run slo-probe (writes evidence/slo-probe.json; not a 99.9% claim).
Incident steps
1) Confirm /v1/health. 2) If TRAI window is closed, outbound commercial traffic is expected to 422. 3) If JWT fails in production, SHRIS_JWT_SECRET must not be empty or the local default. 4) Scale: kubectl -n shris get hpa,pdb. 5) Live PSTN originate still requires human-held Twilio/Exotel/SIP secrets — never invent them. 6) Recordings stay on RECORDINGS_DIR unless AWS/S3 env is set.
Book a demo or open the operations console to inspect campaigns, TRAI gates, and CRM actions.