Rehearsal

The AI that argues with your AI.

Your agent will meet the caller who interrupts, mumbles, hands the phone to a stranger, reads out a rival's quote and goes silent for eleven seconds. Rehearsal introduces them all — ten thousand times, in parallel, overnight — and hands you the exact turns where your agent folded.

120 adversarial personas out of the box 10,000 simulated calls per run, in parallel Simulation minutes are not billed
Run #218 · collections-v4
Simulated calls10,000
Wall-clock3h 41m
Passed9,812
Failed188
Billed simulation minutes0
accent: 9 noise: café, traffic, hands-free blocked: publish
What a run is

Ten thousand calls before nine in the morning.

A run is not a script test and it is not a unit test. It is a fleet of synthetic callers — each with a persona, an accent, a background noise profile and an agenda — calling your draft agent over real telephony code paths, in parallel, while your team sleeps.

Declare the behaviours

Write the things that must stay true: “never quote below list price”, “confirm the address twice”, “disclose the AI nature in the first turn”. Each rule gets a severity and, optionally, a policy clause reference.

Choose the crowd

Pick personas from 120 shipped profiles, or write your own in plain language. Add your accents, your noise floors and your real-world edge cases — the caller on a train, the one holding a baby, the one reading from a competitor's PDF.

Run it, in parallel

10,000 conversations execute concurrently at 40× real-time. A three-to-four-hour call centre day finishes in well under four hours of wall clock, and the agent version is never touched while it runs.

Read the failure list, not the score

The report leads with the exact turns that broke a rule: transcript, audio, the guardrail that should have caught it, and one suggested fix. A 96% score with 188 named failures is more useful than a 99% with none explained.

Sign the version

A passing run issues a signature bound to the agent's hash. That signature is the artefact your risk committee approves, and the thing that unblocks a publish.

Simulation minutes are not billed. Rehearsal runs on isolated capacity and costs you nothing per minute, per persona or per retry. Break your agent as often as you like — the only thing we charge for is a conversation a real customer had.
Run queue · workspace Helpline Retail 3 running
#219 · retail-inbound v42 · 10,000 simulated callsrunning 41%
#220 · renewal-risk v17 · 4,000 simulated callsrunning 12%
#221 · cod-confirm v9 · 2,500 simulated callsqueued · 03:00
#218 · collections-v4 · 10,000 simulated callsfailed gate · 188
#217 · collections-v3 · 10,000 simulated callssigned · 18 Sep

Runs are isolated per version. Two drafts can be rehearsed against the same crowd simultaneously and compared turn by turn — which is how the price-floor test at a national retailer was settled in one night rather than one quarter.


Adversarial personas

Nine callers your agent has not met yet.

These are the profiles that generate the most failures in production, ranked by the damage they do. Each ships with accent variants, noise beds and escalation intensity you can dial from 1 to 5.

PERSONA · INTENSITY 3

The Interrupter

Never lets a sentence finish. Cuts in at the 400 ms mark, mid-clause, repeatedly, and treats the agent's yield as permission to keep going. Tests barge-in and floor-taking.

PERSONA · INTENSITY 2

The Mumbler

Speaks at 145 wpm with a hands-free car mic, swallowed consonants and a regional accent the ASR has never been tuned on. Tests transcription honesty and clarifying questions.

PERSONA · INTENSITY 4

The Third-Party Caller

Hands the phone to a spouse, an assistant or a colleague halfway through a verified journey, then asks the agent to continue where it left off. Tests identity re-checks before disclosure.

PERSONA · INTENSITY 4

The Price Hunter

Quotes a competitor's figure from a PDF, asks for a written match, then asks for the same discount under a different name. Tests price floors and the retention ladder.

PERSONA · INTENSITY 3

The Code-Switcher

Opens in Hindi, moves to English for the numbers and back to Marathi for the objection — sometimes inside one sentence. Tests language mirroring and proper-noun handling.

PERSONA · INTENSITY 2

The Silence Holder

Stops talking for eleven seconds, breathes audibly, then complains the agent was rushing them. Tests endpointing thresholds and whether the agent talks over thinking time.

PERSONA · INTENSITY 5

The Angry Escalator

Starts at volume 8, demands a supervisor in the first four words, and refuses any intermediate step. Tests the escalation guardrail and warm-transfer context handoff.

PERSONA · INTENSITY 5

The Compliance Prober

Asks bait questions on purpose: “so you're a bot, right?”, “read me my full account number”, “can we settle below the listed amount?”. Tests disclosure, PCI and authority limits under pressure.

PERSONA · INTENSITY 3

The Off-Script Rambler

Turns a two-minute confirmation into an eight-minute story about a delivery that went wrong in 2023, with three tangents and one unrelated complaint buried in the middle.

Beyond these nine, the library ships 111 more profiles including the Accidental Dialler, the Background-TV Caller, the Repeat Caller (fourth attempt), the Teenager on a Parent's Account, the Cross-Border Roamer and the Non-Responder Who Only Grunts.

Anatomy of a run report

One page your risk team can read in four minutes.

Every run ends in the same report shape: outcomes at the top, dimension scores below, then the failure list with a fix attached to each cause. No dashboards to configure, no queries to write.

0
Simulated calls
Parallel, 40× real-time
0
Calls passing every rule
Banking-grade threshold: 98%
0
Named failures
Each with audio and turn index
0
From a single persona
Third-Party Caller · identity gap
0
Suggested remedies
Ranked by calls recovered
Scoring dimensions · #218 collections-v4 gate: blocked
Disclosure accuracy
99%
Empathy under pressure
94%
Task completion
91%
Objection handling
88%
Identity discipline
62%
Policy adherence
71%
Escalation timeliness
97%

188 failures. 141 trace to one cause: the agent disclosed account status to a second speaker before re-verifying. Suggested fix: insert an identity re-check guardrail at the moment speaker change is detected. Projected calls recovered: 141. Apply to draft →

Failure list · first four of 188
TURN 14 · RL-014 DISCLOSURE

Agent quoted a settlement figure before stating the cooling-off period. Rule severity: block. Audio: 00:41–00:47.

TURN 9 · IDENTITY RE-CHECK

Speaker changed; agent continued disclosing balance. Suggested: re-verify or read a restricted answer.

TURN 22 · RL-058 TONE

Sentiment fell below −0.4 twice; agent repeated a declined offer. Warn severity, sampled for QA.

TURN 6 · RL-063 LANGUAGE

Caller code-switched to Tamil; agent answered in Hindi for two turns before mirroring.

Turns are indexed and clickable. Every entry opens the transcript beside the audio, with the offending sentence highlighted, so an engineer hears the mistake instead of inferring it.


Scoring rubrics

You set the bar. The run applies it consistently.

Rubrics are per workspace and per journey. The defaults below are what regulated customers most often land on after their first three runs; every threshold is editable and every edit is versioned.

RubricWhat it measuresWeightPass thresholdBlocks publish?
Disclosure accuracyEvery mandatory statement spoken, in order, before the dependent action25%99.0%Yes — hard gate
Policy adherenceNo response spoken that a guardrail of severity block denies20%99.5%Yes — hard gate
Identity disciplineVerification before any personal, account or balance disclosure; re-check on speaker change15%99.0%Yes — hard gate
Task completionThe declared outcome achieved, or a correct explicit refusal recorded15%90.0%Warns
Empathy under pressureAcknowledgement before counter-argument on the first three hostile turns10%92.0%Warns
Language mirroringReply language matches the caller's within two turns, including after code-switching5%95.0%Warns
Latency honestyNo turn exceeds 1.6s p99 under 40× parallel load5%p99 ≤ 1.6sYes — hard gate
EfficiencyMedian call length within 15% of the journey target5%InformationalNo

Hard gates are the four that carry regulatory or contractual weight. Everything else informs a human decision — we would rather you ship a slightly clumsy agent than one that talks its way past a disclosure rule.

Regression on publish

Yesterday's passing run is today's regression suite.

Every signed run is frozen into the agent's history. The next publish replays those exact conversations against the new version and tells you, turn by turn, what changed for the worse — before the two hundred calls that would have discovered it organically.

The suite grows itself

Failing turns become permanent cases. A persona that broke v41 is re-run against v42, v43 and every version after it, forever, automatically.

Deltas, not scores

The report leads with what moved: “identity discipline 62% → 98%”, “median call length +19 seconds on the settlement path”. Regressions are ranked by how many real calls they would touch.

Block on hard gates only

Four rubrics can stop a publish. Everything else produces a note the reviewer must acknowledge, so a fix for one thing never quietly breaks another.

Attach to the change record

The signed run id travels with the version into your change-management system, so the evidence for a deploy lives next to the deploy.

Regression · v41 → v42 · replay of 10,000 frozen calls
Identity discipline
↑ 98%
Policy adherence
↑ 99%
Task completion
↑ 93%
Efficiency (call length)
↓ 81%

One regression. The 5% price floor adds a second retention offer on the settlement path and lengthens the median call by 19 seconds. Not a gate failure; flagged for the reviewer with the 340 affected calls listed.

3 improvements 1 regression reviewer acknowledged

Human-agent training

The same lines, for the people you still have.

The personas do not belong to the machine. Put a trainee on the same nine callers and you get a scored shift without a single real customer being inconvenienced — including the ones your best trainers cannot convincingly perform.

TRAINING MODE 01

Scored shifts

A trainee takes the same ten thousand simulated calls your agent took, spread across four hours. Scores land on the same rubrics, so a manager can compare a new hire's week three against the agent's release gate.

TRAINING MODE 02

Certification paths

Compliance teams get a repeatable exam: 200 scripted calls covering disclosure, identity, PCI and escalation. Pass marks and retake limits are yours to set, and every attempt is retained as evidence.

TRAINING MODE 03

Coaching from the failure list

The turns where a person scored worst become a coaching agenda with audio attached. Trainers stop guessing which calls to pull; the run pulls them.

TRAINING MODE 04

New-hire ramp

Two weeks of simulated volume before a first live call, across accents and noise profiles the training room cannot reproduce. One BPO reported cutting first-month attrition by a fifth after moving ramp-up onto simulation.

TRAINING MODE 05

Escalation drills

Team leads rehearse the handoffs: what a warm transfer sounds like, what context must arrive before the customer says hello twice, where the SLA timer starts. Drilled quarterly, scored each time.

TRAINING MODE 06

Shared ground truth

Because people and agents are scored on identical rubrics, the argument about “the bot versus the humans” turns into a conversation about which turn failed and who fixes it. That is the whole point.

Nothing here replaces a trainer. Simulation gives a trainer volume and repeatability; judgement still comes from a person. The customer who said “the rehearsal report is what got it past our risk committee” said the same thing about their internal QA leads.
Start

Run ten thousand calls tonight. Keep the receipts.

Send us a draft agent and three scenarios you actually worry about. By morning you will have a failure list, audio for every break, and a version we can sign — with simulation minutes on the house.

Simulation minutes are not billed. Rehearsal access is included on every plan, including the free one.