All systems operational.
Checked every thirty seconds from probes in each region. This page is generated from the same monitoring that wakes our on-call engineers, so it cannot be edited into a better story after the fact.
Measured at the carrier edge over the last 60 minutes, across 41,290 turns. Regional detail is in the table below.
Ninety days of uptime, component by component.
Uptime is measured from external probes on the customer-facing surface, not from internal health checks. A component that answers its own health endpoint while failing real calls counts as down here.
| Component | State | 90-day uptime | Trend | Last degradation |
|---|---|---|---|---|
| Telephony edge — US | Operational | 99.99% | 19 Aug 2026 | |
| Telephony edge — EU | Operational | 99.98% | 14 Jun 2026 | |
| Telephony edge — APAC | Operational | 99.96% | 27 Mar 2026 | |
| Speech recognition (ASR) | Operational | 99.995% | 11 May 2026 | |
| Speech synthesis (TTS) | Operational | 99.997% | 08 Jul 2026 | |
| Policy engine | Operational | 99.99% | 27 Mar 2026 | |
| Tool runtime | Operational | 99.97% | 11 May 2026 | |
| Console | Operational | 99.95% | 19 Aug 2026 | |
| Webhooks | Operational | 99.93% | 19 Aug 2026 | |
| Analytics pipeline | Operational | 99.90% | 02 Apr 2026 |
The analytics pipeline is the lowest number on this page and has been for three quarters. Dashboards can lag behind live calls by up to ten minutes under heavy backfill; no call is affected, and we would rather publish the gap than round it away.
Distance still costs milliseconds.
Numbers below are a sixty-minute window on a normal Tuesday, measured from the carrier edge to the first audible reply. Your own figures will sit inside these ranges unless your carrier adds a leg.
| Region | p50 turn | p99 turn | ASR first token | TTS first byte | Turns sampled |
|---|---|---|---|---|---|
| uk-south · London | 0.58s | 1.31s | 166 ms | 138 ms | 14,204 |
| us-west · Oregon | 0.60s | 1.38s | 171 ms | 142 ms | 9,118 |
| eu-central · Frankfurt | 0.62s | 1.44s | 174 ms | 146 ms | 7,640 |
| eu-west · London | 0.63s | 1.49s | 177 ms | 149 ms | 5,025 |
| ap-south · Mumbai | 0.66s | 1.58s | 182 ms | 154 ms | 3,901 |
| ap-southeast · Sydney | 0.71s | 1.72s | 194 ms | 163 ms | 1,180 |
| me-central · Dubai | 0.74s | 1.84s | 201 ms | 169 ms | 222 |
Eight entries since February.
Maintenance windows are listed alongside incidents, because a planned change that misbehaves is still worth reading about. Root causes are stated in one line; the three customer-impacting incidents link to full write-ups.
| Date | Type | What happened | Duration | Root cause |
|---|---|---|---|---|
| 04 Sep 2026 | Maintenance | EU media gateway version upgrade; no customer impact | 22 min | Planned change, drained cleanly |
| 19 Aug 2026 | Degraded | Webhook deliveries in uk-south arrived up to 12 minutes late | 46 min | One customer endpoint responded in 400 ms, tripping our own consumer rate limiter |
| 08 Jul 2026 | Partial outage | One en-GB voice unavailable; 3,410 calls used the fallback voice | 1 h 12 min | Model activation file failed a checksum after a routine deploy |
| 14 Jun 2026 | Maintenance | ap-southeast failover drill; synthetic traffic only | 30 min | Planned drill, no anomaly |
| 11 May 2026 | Incident | ASR misread alphanumeric strings on noisy lines; 34 delivery reschedules went out with a bad postcode | 4 days to full fix | Language-model bias on digit groups at low bitrate — post-mortem |
| 02 Apr 2026 | Degraded | Analytics dashboards stale by up to 40 minutes | 2 h 04 min | An overnight backfill job competed with live aggregation for warehouse slots |
| 27 Mar 2026 | Partial outage | Policy evaluation stalled in ap-south; 1,240 calls ran on a static script | 18 min | Cache eviction storm after a rule-set publish — post-mortem |
| 02 Feb 2026 | Incident | Three outbound calls continued after a caller asked us to stop | Remediated same week | Opt-out classifier matched a fixed phrase list — post-mortem |
The policy engine stall, in detail
At 11:42 IST a rule-set publish to the Mumbai region evicted the warm policy cache across all 14 cells at once. Rebuilding the cache required a full read of the rule set from storage, and with every cell doing it simultaneously, evaluation latency passed the 900 ms fail-safe threshold. The runtime did exactly what it was built to do: it stopped evaluating and served a static script that could still complete a booking, dropping every upsell branch. One thousand two hundred and forty calls were affected over eighteen minutes.
Detection was automatic — the p95 alert fired 94 seconds after the first stalled call, and the on-call engineer rolled back the publish at 12:00 IST. The failure was not in the fail-safe; it was in the release process that assumed a cache prewarm would happen on its own.
- Duration
- 18 minutes of degraded evaluation, 41 minutes total incident
- Blast radius
- 1,240 calls in one region; no data loss; no dropped connections
- Root cause
- Simultaneous cache eviction during a rule-set publish
- Fix shipped
- Staged rollout to 5% of traffic for ten minutes, automatic rollback on p95 regression, prewarm as a mandatory release step
- Verified by
- Rehearsal run 1,904 — 10,000 simulated calls against the new publish path, 10,000 passed
Three windows between now and November.
EU media gateway upgrade
Rolling restart of the Frankfurt and Amsterdam media gateways. Calls in progress continue uninterrupted; new calls re-register within 30 seconds and may see one extra ring.
Expected impact: none. Rollback window: 30 minutes.
Analytics warehouse maintenance
Storage compaction and index rebuilding in both US and EU analytics clusters. Live calls are unaffected; dashboards will read up to 30 minutes behind for the duration.
Expected impact: stale dashboards only. No API change.
Annual uk-south failover drill
We move synthetic traffic from London to Oregon and back to prove the failover path works when nobody is panicking. Customer calls stay on their primary region unless they opt into the drill.
Expected impact: none for customers who do not opt in.
Three ways to hear about it before your customers do.
Pick the one that fits your operations. All three fire on the same state change, so there is no lag between the page and the notification.
Email digest
One plain-text message when an incident opens and one when it closes. No marketing, no product news, no unsubscribe tricks — the list exists only for this. Ask support to add the addresses you want, or set them per team in the Console.
Webhook
Subscribe to incident.opened, incident.updated and incident.resolved from your own endpoint. Delivery follows the same ten-attempt, at-least-once semantics as call webhooks, described on the developers page.
Feed reader
A feed at the status page path, updated on every state change, readable without an account. Operations teams usually wire this into the same screen they keep open for carrier status.
Per-component alerts
Subscribe to a single component rather than everything — useful when only telephony in one region matters to your escalation path. Configured in the Console, applied per person.
- First notice
- Under 15 minutes from confirmed impact
- Update cadence
- Every 30 minutes while an incident is open
- Resolution note
- Within 2 hours of service restoration
- Root cause
- Written summary within 72 hours
- Credit claims
- Automatic, no ticket required
- Language
- English only; regional summaries on request
If we miss the fifteen-minute first notice, the incident is escalated to the engineering director on call and we say so in the entry.
What we owe you, by plan.
Credits are issued against the monthly invoice automatically when a commitment is missed. You do not have to notice, and you do not have to ask.
| Commitment | Launch | Scale | Sovereign |
|---|---|---|---|
| Monthly uptime | 99.9% | 99.95% | 99.99% |
| Service credits | 10% of monthly fee | 10% / 25% / 50% by severity | 15% / 30% / 60% by severity |
| Support response | Business hours, email | 1 hour, priority queue | 30 minutes, named architect |
| Incident notification | 60 minutes, email | 15 minutes, email and webhook | 10 minutes, bridge call opened |
| Maintenance notice | 72 hours | 7 days | 14 days |
| Latency commitment | None | p99 under 2.0s per region | p99 under 1.6s, contractually measured |
| Post-incident review | On request | Within 72 hours | Within 72 hours plus a review call |
Rolling twelve-month availability across all regions is 99.98%. The latency commitment is measured the same way as the figures at the top of this page, from the carrier edge, and is not adjusted for a customer's own carrier legs. Plan limits are on the pricing page.
Ask us about the incident you are worried about.
If a competitor has told you our history is too short to trust, bring it up on a call. We will walk through any entry on this page, including the three post-mortems we published on purpose.
This page is generated from live monitoring and archived monthly. Previous months are available on request.