Benchmarks
Six vendors, one test set, one carrier.
Every vendor was given the same 1,000 recorded calls across three regions and asked to complete the same tasks: book a visit, confirm an address, take a payment promise, escalate to a human.
| Vendor |
Median turn latency |
p95 latency |
Task completion |
Entity accuracy, noisy line |
Fabrication rate |
WER, accented English |
| Nivākya v4 |
0.61s | 1.38s | 92.4% | 96.1% | 0.4% | 6.2% |
| Kestrel CX Voice |
0.98s | 2.44s | 85.3% | 91.2% | 0.3% | 9.4% |
| Aspen Dialog |
1.12s | 2.90s | 83.1% | 89.4% | 1.9% | 10.6% |
| LivePerson Conversational AI |
1.24s | 3.10s | 81.0% | 88.7% | 2.6% | 11.9% |
| Northvoice Cloud |
1.41s | 3.86s | 78.6% | 85.9% | 3.1% | 13.2% |
| Ensemble Voice |
1.77s | 4.20s | 74.2% | 83.5% | 4.2% | 15.1% |
Methodology, in full. Calls were replayed from a consented corpus of 1,000 English and Hindi-English conversations recorded between February and April 2026. All six vendors were reached over the same carrier on the same day-parts, in uk-south, eu-west and ap-south. Latency is measured at the carrier edge as the interval between the end of the customer's utterance and the first audio byte of the reply. Fabrication means a claim made by the agent that no tool result, knowledge-base article or customer utterance supports.
What we did not control. Task completion is judged by two human raters with a 0.91 agreement rate; disagreements were resolved by a third rater who did not know which vendor produced the call. Aspen and Ensemble ran on their default US voices for the Hindi-English segment, which is not their strongest configuration and we say so in the appendix rather than quietly removing them.
Raw transcripts, per-turn timings and the rater sheets are available to customers and to any prospect who has signed an NDA. Request them through the demo form.
Post-mortems
Three calls we got wrong, written up in public.
We publish these because a vendor with no failure record is either very young or not looking. Each one lists what broke, how long it lasted, and the specific change that followed.
INCIDENT 2026-05-11 · 34 CALLS
The postcode we misheard
What broke. On a noisy line, ASR returned “ZR-4” for “ZR-14” on a delivery-rescheduling agent. The agent read the address back as a single string, the caller said “yes” to the wrong thing, and 34 reschedules went out with an undeliverable address.
What changed. All alphanumeric fields now require digit-by-digit readback with explicit confirmation per character, and a keypad fallback fires after the second failed readback. Median handling time rose 4.1 seconds. We accepted that.
INCIDENT 2026-02-02 · 3 CALLS
The retry ladder that kept dialling
What broke. A caller asked us to stop calling, in the words “please don't ring this number again”. The opt-out classifier only matched a fixed set of phrases and missed it. The ladder dialled twice more over four days.
What changed. Twenty-two opt-out phrasings added, suppression is now immediate and irreversible for 90 days, and every suppression event goes to a human queue for next-day review. The classifier's job changed from “is this an opt-out” to “could this plausibly be an opt-out”.
INCIDENT 2026-03-27 · 18 MIN, 1,240 CALLS
The policy engine stall in ap-south
What broke. A rule-set publish evicted a warm cache faster than it could refill, and policy evaluation stalled for 18 minutes in one region. Calls degraded to a static script that completed bookings but lost every upsell branch.
What changed. Publishes are staged to 5% of traffic for ten minutes, a p95 regression on that slice triggers automatic rollback, and cache prewarm is now a required step in the release checklist rather than an optimisation.
How we write these. Within 72 hours of a customer-impacting incident we publish a summary to affected workspace owners, with the call count, the affected region and the change list. The three you just read are the ones where the fix was expensive enough that we wanted other operators to learn from it. Status history lives on the
status page.
Podcast — Signal Path
Conversations with the people who rebuilt their voice channel.
Forty minutes, one operator, no vendor pitch. Transcripts are published within a week, and guests get to correct their own words before we ship them.
| Ep | Title | Guest | Length | Published |
| 14 | Nobody misses a queue if nobody notices one | VP Customer Experience, national property developer | 38 min | 02 Sep 2026 |
| 13 | The ninety seconds that decide a collections call | Head of Collections, regulated NBFC | 44 min | 12 Aug 2026 |
| 12 | Buying voice AI without a sandbox | Procurement Lead, six-hospital network | 31 min | 22 Jul 2026 |
| 11 | Transcripts as evidence | Compliance Director, general insurer | 46 min | 01 Jul 2026 |
| 10 | Your IVR tree is the best documentation you have | Director of Service Operations, telecom | 29 min | 10 Jun 2026 |
| 09 | Hindi, Tamil, English: one number | Head of CX, insurance distributor | 41 min | 20 May 2026 |
| 08 | The rehearsal report that killed a project | Programme Manager, retail bank | 36 min | 29 Apr 2026 |
| 07 | Agent handover without a sigh | Team Lead, outsourced contact centre | 33 min | 08 Apr 2026 |
Signal Path is also available as a written digest — every episode ships with a two-page summary and the three timestamps worth hearing. Ask for the digest through the demo form.