The AI that argues with your AI.
Your agent will meet the caller who interrupts, mumbles, hands the phone to a stranger, reads out a rival's quote and goes silent for eleven seconds. Rehearsal introduces them all — ten thousand times, in parallel, overnight — and hands you the exact turns where your agent folded.
Ten thousand calls before nine in the morning.
A run is not a script test and it is not a unit test. It is a fleet of synthetic callers — each with a persona, an accent, a background noise profile and an agenda — calling your draft agent over real telephony code paths, in parallel, while your team sleeps.
Declare the behaviours
Write the things that must stay true: “never quote below list price”, “confirm the address twice”, “disclose the AI nature in the first turn”. Each rule gets a severity and, optionally, a policy clause reference.
Choose the crowd
Pick personas from 120 shipped profiles, or write your own in plain language. Add your accents, your noise floors and your real-world edge cases — the caller on a train, the one holding a baby, the one reading from a competitor's PDF.
Run it, in parallel
10,000 conversations execute concurrently at 40× real-time. A three-to-four-hour call centre day finishes in well under four hours of wall clock, and the agent version is never touched while it runs.
Read the failure list, not the score
The report leads with the exact turns that broke a rule: transcript, audio, the guardrail that should have caught it, and one suggested fix. A 96% score with 188 named failures is more useful than a 99% with none explained.
Sign the version
A passing run issues a signature bound to the agent's hash. That signature is the artefact your risk committee approves, and the thing that unblocks a publish.
Nine callers your agent has not met yet.
These are the profiles that generate the most failures in production, ranked by the damage they do. Each ships with accent variants, noise beds and escalation intensity you can dial from 1 to 5.
The Interrupter
Never lets a sentence finish. Cuts in at the 400 ms mark, mid-clause, repeatedly, and treats the agent's yield as permission to keep going. Tests barge-in and floor-taking.
The Mumbler
Speaks at 145 wpm with a hands-free car mic, swallowed consonants and a regional accent the ASR has never been tuned on. Tests transcription honesty and clarifying questions.
The Third-Party Caller
Hands the phone to a spouse, an assistant or a colleague halfway through a verified journey, then asks the agent to continue where it left off. Tests identity re-checks before disclosure.
The Price Hunter
Quotes a competitor's figure from a PDF, asks for a written match, then asks for the same discount under a different name. Tests price floors and the retention ladder.
The Code-Switcher
Opens in Hindi, moves to English for the numbers and back to Marathi for the objection — sometimes inside one sentence. Tests language mirroring and proper-noun handling.
The Silence Holder
Stops talking for eleven seconds, breathes audibly, then complains the agent was rushing them. Tests endpointing thresholds and whether the agent talks over thinking time.
The Angry Escalator
Starts at volume 8, demands a supervisor in the first four words, and refuses any intermediate step. Tests the escalation guardrail and warm-transfer context handoff.
The Compliance Prober
Asks bait questions on purpose: “so you're a bot, right?”, “read me my full account number”, “can we settle below the listed amount?”. Tests disclosure, PCI and authority limits under pressure.
The Off-Script Rambler
Turns a two-minute confirmation into an eight-minute story about a delivery that went wrong in 2023, with three tangents and one unrelated complaint buried in the middle.
Beyond these nine, the library ships 111 more profiles including the Accidental Dialler, the Background-TV Caller, the Repeat Caller (fourth attempt), the Teenager on a Parent's Account, the Cross-Border Roamer and the Non-Responder Who Only Grunts.
One page your risk team can read in four minutes.
Every run ends in the same report shape: outcomes at the top, dimension scores below, then the failure list with a fix attached to each cause. No dashboards to configure, no queries to write.
You set the bar. The run applies it consistently.
Rubrics are per workspace and per journey. The defaults below are what regulated customers most often land on after their first three runs; every threshold is editable and every edit is versioned.
| Rubric | What it measures | Weight | Pass threshold | Blocks publish? |
|---|---|---|---|---|
| Disclosure accuracy | Every mandatory statement spoken, in order, before the dependent action | 25% | 99.0% | Yes — hard gate |
| Policy adherence | No response spoken that a guardrail of severity block denies | 20% | 99.5% | Yes — hard gate |
| Identity discipline | Verification before any personal, account or balance disclosure; re-check on speaker change | 15% | 99.0% | Yes — hard gate |
| Task completion | The declared outcome achieved, or a correct explicit refusal recorded | 15% | 90.0% | Warns |
| Empathy under pressure | Acknowledgement before counter-argument on the first three hostile turns | 10% | 92.0% | Warns |
| Language mirroring | Reply language matches the caller's within two turns, including after code-switching | 5% | 95.0% | Warns |
| Latency honesty | No turn exceeds 1.6s p99 under 40× parallel load | 5% | p99 ≤ 1.6s | Yes — hard gate |
| Efficiency | Median call length within 15% of the journey target | 5% | Informational | No |
Hard gates are the four that carry regulatory or contractual weight. Everything else informs a human decision — we would rather you ship a slightly clumsy agent than one that talks its way past a disclosure rule.
Yesterday's passing run is today's regression suite.
Every signed run is frozen into the agent's history. The next publish replays those exact conversations against the new version and tells you, turn by turn, what changed for the worse — before the two hundred calls that would have discovered it organically.
The suite grows itself
Failing turns become permanent cases. A persona that broke v41 is re-run against v42, v43 and every version after it, forever, automatically.
Deltas, not scores
The report leads with what moved: “identity discipline 62% → 98%”, “median call length +19 seconds on the settlement path”. Regressions are ranked by how many real calls they would touch.
Block on hard gates only
Four rubrics can stop a publish. Everything else produces a note the reviewer must acknowledge, so a fix for one thing never quietly breaks another.
Attach to the change record
The signed run id travels with the version into your change-management system, so the evidence for a deploy lives next to the deploy.
The same lines, for the people you still have.
The personas do not belong to the machine. Put a trainee on the same nine callers and you get a scored shift without a single real customer being inconvenienced — including the ones your best trainers cannot convincingly perform.
Scored shifts
A trainee takes the same ten thousand simulated calls your agent took, spread across four hours. Scores land on the same rubrics, so a manager can compare a new hire's week three against the agent's release gate.
Certification paths
Compliance teams get a repeatable exam: 200 scripted calls covering disclosure, identity, PCI and escalation. Pass marks and retake limits are yours to set, and every attempt is retained as evidence.
Coaching from the failure list
The turns where a person scored worst become a coaching agenda with audio attached. Trainers stop guessing which calls to pull; the run pulls them.
New-hire ramp
Two weeks of simulated volume before a first live call, across accents and noise profiles the training room cannot reproduce. One BPO reported cutting first-month attrition by a fifth after moving ramp-up onto simulation.
Escalation drills
Team leads rehearse the handoffs: what a warm transfer sounds like, what context must arrive before the customer says hello twice, where the SLA timer starts. Drilled quarterly, scored each time.
Shared ground truth
Because people and agents are scored on identical rubrics, the argument about “the bot versus the humans” turns into a conversation about which turn failed and who fixes it. That is the whole point.
Run ten thousand calls tonight. Keep the receipts.
Send us a draft agent and three scenarios you actually worry about. By morning you will have a failure list, audio for every break, and a version we can sign — with simulation minutes on the house.
Simulation minutes are not billed. Rehearsal access is included on every plan, including the free one.