Runs this month
62
▲ 9 vs August
Simulations executed
1.24M
▲ 38% on a wider persona set
Pass rate (all runs)
97.1%
▲ 0.6 pts this week
Worst rubric
92.4%
Empathy under pressure
Mean time to fix
3.4 days
worst: billing-dispute at 9 days

Run #218 — results

collections-v4 · 10,000 simulated calls Agent library
PASSED
9,812
FAILED
188
PASS RATE
98.12%
publish gate cleared
RUBRIC BREAKDOWN
Empathy under pressure
92.4%
Disclosure accuracy
99.1%
Objection handling
96.8%
Policy adherence
99.6%
Identity verification
99.9%
Language switching
97.3%
Off-script recovery
94.2%
Wrap-up completeness
98.7%
RUN TIMELINE · 18M 42S
Scored turns Failure clusters Idle / setup

Failure drill-down

188 failures
The Interrupter54

“…no wait — just tell me whether it's covered.”

Fix: raise interruption tolerance to 35% and cap the first response at twelve words.
The Angry Escalator47

“I've been a customer for nine years and this is how you treat me?”

Fix: the first empathy marker lands on turn three; move it to the opening turn.
The Compliance Prober31

“So if I'm recording this call, what exactly are you allowed to promise me?”

Fix: add a scripted recording-disclosure line before any monetary offer is made.
The Code-Switcher26

“Mira, I need the plan in Spanish but the invoice in English.”

Fix: hold the chosen language across the invoice sub-flow instead of resetting to en-GB.
The Silence Holder22

“(fourteen seconds pass with no answer, then a single word) …maybe.”

Fix: drop the silence threshold from 8s to 5s before the agent re-prompts.

Persona library

adversarial set · v9
PersonaDescriptionDifficultyTimes usedCatch rate
The InterrupterCuts across the agent mid-sentence and refuses to let a full answer land.hard1,20454%
The MumblerLow volume, trailing consonants and a regional accent under steady background hum.medium98638%
The Third-Party CallerCalls about someone else's account and pushes for details without the account PIN.hard74271%
The Price HunterQuotes a competitor's price and demands the difference in writing within the call.medium1,53044%
The Code-SwitcherAlternates between English and Spanish inside a single sentence.hard87862%
The Silence HolderAnswers with long pauses and single words, testing how long the agent waits.easy61121%
The Angry EscalatorMoves from mild irritation to shouting within four turns and demands a manager.hard1,41768%
The Compliance ProberAsks, on the record, what the agent is and is not allowed to promise.medium80349%
The Off-Script RamblerIgnores every question and narrates unrelated personal history for minutes.medium1,12241%
The Background-Noise CallerCalls from a moving vehicle or a crowded street with layered noise.easy1,68827%
The Fast TalkerSustains 210+ words per minute with almost no natural pauses.medium95446%
The Repeat CallerHas called four times about the same unresolved ticket and says so repeatedly.hard68957%

Run history

last 10 runs
RunAgent versionSimsPass rateDurationTriggered byStatus
#218collections-v410,00098.1%18m 42sPublish gatecomplete
#217qualify-buyer v1410,00099.4%16m 05sPublish gatecomplete
#216renewal-nudge v610,00096.2%21m 18sNightly regressionfailed
#215cod-confirm v910,00099.1%15m 44sNightly regressioncomplete
#214qualify-buyer v1310,00097.0%19m 02sPublish gatecomplete
#213return-pickup v410,00095.8%22m 36sManual · R. Davidblocked
#212nurse-triage v25,00094.4%11m 09sManual · clinical QAblocked
#211qualify-buyer v1210,00098.8%17m 51sPublish gatecomplete
#210order-status v1210,00099.6%14m 58sNightly regressioncomplete
#209billing-dispute v410,00091.3%24m 12sPublish gaterejected

Schedule & regression

workspace defaults
UPCOMING SCHEDULED RUNS
Tonight 02:00Nightly · 12 production agents
Tonight 02:00Drift re-check · renewal-nudge v6
Sun 23:00Persona refresh · adversarial library
Mon 09:00Full sweep · all 38 agents
next in 1h 48m

Human agent training

shared rubric

The same twelve adversarial personas that stress the AI run live drills with new hires. Trainers replay the recording and score it against the identical rubric, so a human and an agent are measured on the same evidence — one scorecard, two cohorts. Cohorts clear a persona when three consecutive drills score above 90% on the empathy and policy rubrics together.

Cohort 24C graduated
COHORT PROGRESS
Cohort 24A · 18 trainees9 / 12 personas
Cohort 24B · 14 trainees7 / 12 personas
Cohort 24C · 11 trainees12 / 12 · 94.8%
Escalation pod · 9 agents11 / 12 · 89.6%
Cleared 3 consecutive drills Needs a re-drill