The platform

A conversation runtime, not a bot builder.

Most voice vendors sell you a flowchart editor and hope. Nivākya runs a governed runtime: every call is admitted, scheduled, policy-checked, measured and written to an audit spine you can query afterwards. You get four surfaces on one model — operations, design, simulation and analytics — and nothing is retyped between them.

0.6s median response, published Six layers, one version number Every turn hashed and retained
Runtime · prod · uk-south
Active calls3,184
Agent versionretail-inbound · v41
Policy denials, last hour6
Audit spine lag41 ms
The conversation loop

Every call is three moves. We publish the budget for all three.

A runtime is only as good as the worst millisecond inside it. Nivākya splits the turn into hear → decide → act, gives each stage its own latency budget, and alarms on the stage — not on the average — when the budget is blown.

Stage 01 — Hear
180ms

Streaming ASR with true barge-in. The agent yields the floor 40 ms after the caller starts a word, so interruptions survive instead of being talked over.

Stage 02 — Decide
240ms

Intent, entity and policy resolution against your price list, calendars, DNC register and knowledge base. Retrieval is pre-warmed per active campaign, so nothing cold-starts mid-sentence.

Stage 03 — Act
190ms

Speech synthesis, tool dispatch and CRM write issued in parallel. The deal record usually lands 300 ms before the caller finishes saying goodbye.

Turn budget, p50 / p95 / p99
0.6s / 0.94s / 1.4s

Measured at the carrier edge, caller-to-agent, across 4.2M production calls in the trailing 90 days. Regional splits sit on the status page.

Why a budget and not a headline number. Median response times hide the tail, and the tail is where callers hang up. Each stage above has a hard ceiling that trips a paging alert; the regional p50s in the last 30 days were 561 ms (Mumbai), 608 ms (Virginia), 634 ms (Frankfurt) and 671 ms (Singapore). A regression of more than 80 ms on any stage blocks the next agent publish until it is explained.
Architecture

Six layers, each one independently observable.

Nothing in Nivākya is a monolith you have to trust. Every layer emits metrics, has its own failure mode, and can be replaced — including the speech models, the LLM routing and the telephony carrier underneath.

LAYER 01 / TELEPHONY EDGE

Carrier-grade ingress

SIP trunking, PSTN failover and number portability across Mumbai, Frankfurt, Ashburn and Singapore edges. Calls are admitted, tagged with consent state, and pinned to a region before a single model sees audio.

LAYER 02 / SPEECH STACK

Hear and speak, swappable

Streaming ASR, diarisation, endpointing and neural TTS, with per-language voice routing. Bring your own model or use ours; the runtime treats either as a pluggable provider with the same telemetry.

LAYER 03 / POLICY ENGINE

Rules that cannot be talked around

Every proposed response passes a deterministic guardrail pass before synthesis. Disclosure requirements, price floors, DNC lists and escalation triggers are compiled to fast predicates — evaluated in 9 ms, not prompt-engineered.

LAYER 04 / TOOL RUNTIME

Actions with a paper trail

Typed tool calls to your CRM, calendar, ERP and payment rails, with schema validation, idempotency keys and a per-tool retry ladder. A failed booking retries three times and then hands off — it never silently disappears.

LAYER 05 / MEMORY STORE

Context that survives the hangup

Caller profiles, open tickets, prior objections and grant-state are resolved before the first word and cached in a per-region store. Memory is versioned alongside the agent, so a rollback restores the facts the old version expected.

LAYER 06 / AUDIT SPINE

Append-only, queryable, hash-chained

Every turn, tool call and policy decision is appended with a SHA-256 chain and shipped to your own storage. Compliance queries 90 days of conversations in console time; we hold the chain, you hold the keys.

0M
Calls measured / 90 days
Caller-to-agent, at the carrier edge
0s
Median turn, p50
Published per stage, alerted per stage
0ms
Audit spine lag
Hash-chained, near-real-time
0
Rolling availability
12-month window, all regions
0ms
Guardrail evaluation
Deterministic pass, before synthesis
Turn trace · retail-inbound v41 · call 8e41-9917
edge: mumbai-2 · 3 ms vad: speech-end 142 ms asr: nk-stream-hi-en · 176 ms intent: object_exchange · 0.91 policy: 4 rules · 9 ms tools: order.lookup (41 ms) · slot.hold (88 ms) tts: voice-kavya-v3 · first byte 74 ms audit: chained · seq 91,442,118

Rendered from the same trace object the console, the rehearsal report and the warehouse export all read from. There is no second copy of the truth.


Versioning & governance

An agent is an asset with a change history.

Prompt, tools, guardrails, voice, memory schema and escalation policy travel together in one immutable version. Nothing reaches production without a named reviewer, a passing rehearsal suite and a rollback path that has been tested at least once.

Draft

An editor branches from production and changes anything — a sentence of persona, a price floor, a calendar route. Every edit is captured with the author and the timestamp, never overwritten.

Review

A second person reviews a semantic diff: which behaviours changed, which guardrails moved, which tools were added. Sign-off is a named approval, not a thumbs-up in a chat thread.

Canary · 5%

The version takes one in twenty live calls for 45 minutes. Task completion, escalation rate and policy denials are compared against the incumbent with automatic halt thresholds.

Full rollout

Traffic moves in 25% steps with a 15-minute soak between each. Any halt condition returns traffic to the incumbent without dropping a live call.

Rollback

One click, under 20 seconds, mid-call. The runtime re-reads the prior version's memory schema on the fly, so a caller never notices the change of brain.

Release gate · v41 → v42
Rehearsal suite · 10,000 simulated callsPASSED 9,841
Policy rule coverage100% of 214
Reviewer sign-offR. Iyer · 18 Sep
Latency delta, p95+31 ms (within budget)
Rollback drill18.4s · 12 Sep
Data-residency checkeu-west pin honoured

A failing gate does not block the edit — it blocks the deploy. Engineers keep shipping drafts; only the runtime is opinionated about which one customers meet.

Read the governance model
Deployment models

Same runtime, four ways to run it.

Regulated buyers rarely want the same thing twice. Pick the floor you are comfortable with — the API, the agent semantics and the audit format do not change when you move.

Model Where it runs Data residency Typical p50 turn Who it suits Ops burden
Managed cloud
Shared multi-tenant
Nivākya control plane, 4 regions Region-pinned on request 0.60s Teams shipping a first production agent in under two weeks None — we page ourselves
Dedicated VPC
Single-tenant
Your AWS, GCP or Azure account Your cloud footprint 0.64s Enterprises with a network team and a security review it must pass Terraform module, we own upgrades
In-country residency
Sovereign region
India (Mumbai), EU (Frankfurt), UAE (Dubai), UK (London) Audio, transcripts and derivatives never leave the border 0.61 – 0.71s Banking, insurance, health and public sector under DPDP or GDPR scrutiny We operate, you audit
On-prem enclave
Air-gap capable
Your data centre, GPU nodes you own No egress at all 0.78 – 1.10s Defence, payments infrastructure, telcos with strict interconnect rules Joint runbook, quarterly model refresh

Latency ranges measured on comparable 8-vCPU / 2-GPU node classes. Moving between managed cloud and dedicated VPC is a configuration change, not a re-implementation — the agent version hash stays identical.


The honest comparison

Buy, build, or keep the queue.

We lose deals to in-house builds and we lose deals to incumbents. Here is where each option actually wins, written by the team that has to support the choice afterwards.

Dimension Nivākya Build-your-own stack Legacy CCaaS + IVR
Time to first live call 9–14 days 4–9 months, plus hiring 6–10 weeks of flow configuration
Latency ownership Published per stage, alerted on Yours to discover, usually in production Menu trees, not conversational turn-taking
What you can prove after a call Hash-chained transcript, policy decisions, tool calls, redaction receipt Whatever you remembered to log CDR fields and a recording
Guardrails Deterministic rule pass, 9 ms, mapped to policy clauses Prompt instructions that can be argued out of Script branching decided in advance
Pre-release testing 10,000 simulated calls per publish, free A QA team on a headset Regression scripts replayed quarterly
Multilingual behaviour Mid-call code-switching across 42 languages One model, one locale, one integration each Separate IVR per language
Where it wins Boring, observable, defensible conversations at volume Full control, full cost, full responsibility Deep telephony estate and existing contracts
Cost shape Per connected conversation + pass-through carrier minutes Engineering headcount, forever Per-seat, per-month, whether calls come or not
Read the build-vs-buy model

The spreadsheet we hand to CTOs, including the parts where building is cheaper.

Open the model →
Component gap analysis

What we were missing against LivePerson, and what shipped in response.

Read the analysis →
Benchmark, Q3 2026

Latency, completion and hallucination rates across six vendors on one test set.

See the numbers →
Start

Put the runtime on one real number.

Bring three recordings of calls you are proud of and three you are not. We will build the agent against them, rehearse it overnight, and show you the report before it rings anyone new.

Managed cloud. 1,000 free conversations in the first 30 days. No credit card, no ring light.