Overview
- When the AAA server stops answering, nothing new connects, and sessions already online drop as they re-authenticate or time out.
- For a telecom operator, the cost of network downtime climbs through the hour and lands mostly in SLA credits, churn, and unbilled usage; lost usage revenue is a small share.
- Below: a cost framework, a formula with telecom inputs, and an illustrative example to run on your own numbers.
When the AAA (Authentication, Authorization, and Accounting) server stops answering, nothing new connects, and sessions already online drop as they re-authenticate or time out. For a telecom operator, the cost of network downtime climbs through the hour and lands mostly in service-level agreement (SLA) credits, churn, and unbilled usage. Lost usage revenue is a small share.
Most downtime statistics assume a system goes dark and the meter runs at a steady rate. Authentication downtime does not behave that way. At minute five the network looks mostly fine. By minute thirty the call center is flooded, prepaid customers cannot top up, and the regulator’s clock may be running.
What Counts as Authentication Downtime?
Authentication downtime is any period when a network’s AAA function cannot validate subscribers or devices fast enough to admit them. New connections fail immediately. Existing sessions survive only until their next re-authentication, interim accounting update, or session timeout, which is why an AAA outage escalates instead of staying flat.
That covers a hard outage, where the server stops responding and the NAS (network access server) marks it dead, and degraded authentication, where it answers slowly or rejects valid requests. Treat sustained latency above your admission timeout as downtime too.
AAA sits in the control plane, so a dead server drops no packets by itself. The gateway carrying the session, whether a broadband gateway, a 4G packet gateway, or a 5G user plane function, keeps forwarding until its timers expire. That is why the first fifteen minutes are so often misdiagnosed: the dashboards still show traffic.
The 60-Minute Anatomy of an AAA Outage
Exact timings depend on your session timeout, interim-accounting interval, and re-authentication policy; the shape holds.
| Time | What fails | What still works | Cost category triggered |
| T+0–5 min | New attaches, Wi-Fi joins, router reboots, SIM activations | Established sessions keep forwarding | Direct revenue; incident response |
| T+5–15 min | Interim accounting updates unanswered; NAS retry queues build | Most existing sessions; dashboards look normal | Accounting leakage; call volume |
| T+15–30 min | Short-timer sessions (Wi-Fi, 802.1X, prepaid data) drop; Wi-Fi calling registrations fail | Long-timeout fixed broadband | SLA breaches; churn risk |
| T+30–45 min | Session timeouts cascade; re-authentication storm on recovery | Little outside fixed broadband | Regulatory thresholds; public visibility |
| T+45–60 min | Prepaid top-ups and roaming-partner authentication fail; call center at peak | Recovery depends on MTTR (mean time to recovery) | All six; credits and churn dominate |
The NAS sends a RADIUS Interim-Update at the interval the server set in Acct-Interim-Interval (RFC 2869), often configured between 5 and 15 minutes. Unanswered updates are buffered or dropped, and that usage becomes hard to bill. Short re-authentication timers expire and fail. Wi-Fi calling registrations go with them, because the ePDG (evolved packet data gateway) authenticates through the 3GPP AAA server over SWm (3GPP TS 29.273).
By minute thirty the outage is visible from outside. In the US, covered providers must file an outage of at least 30 minutes that potentially affects 900,000 or more user-minutes in the FCC’s Network Outage Reporting System (47 CFR 4.9).
Six Cost Categories Most Downtime Estimates Miss
Most downtime calculators give you one line: revenue per hour times hours down. For an operator that is usually the smallest line, because flat-rate postpaid customers still pay whether or not they could connect.
| Category | What drives it |
| Direct revenue loss | Prepaid and usage-based revenue in the window; activations and top-ups lost outright |
| SLA credits | Enterprise, wholesale, and MVNO (Mobile Virtual Network Operator) host availability commitments, breached across every contract at once |
| Accounting-record leakage | Interim-Update and Stop records lost or duplicated; usage unbilled |
| Incident response | Engineers on the bridge, vendor support, call-center overtime, post-incident review |
| Churn replacement | Subscribers who leave, times the cost to replace them |
| Regulatory and brand | Outage reports, remediation plans, fines, press coverage |
One hour against a 99.99% monthly commitment is roughly fourteen times the allowance of about 4.4 minutes. Churn is the most uncertain input and the most expensive: one hour rarely makes a customer leave on its own, but it tips customers who were already marginal, especially prepaid and Wi-Fi users.
How to Calculate the Cost of One Hour of Authentication Downtime
To calculate the cost of one hour of network downtime, add direct revenue lost, SLA credits owed, estimated churn replacement cost, incident response labor, and revenue leakage from lost accounting records, then scale by the share of subscribers affected and the peak-hour multiplier.
The inputs
| Input | Symbol | Illustrative value |
| Subscribers | N | 2,000,000 |
| Blended monthly ARPU (average revenue per user) | A | $25 |
| Usage-contingent share of revenue | U | 30% |
| Sessions failing in the hour | S | 50% |
| Peak-hour multiplier | P | 1.5 |
| Sales lost, not deferred | L | $6,000 |
| Monthly revenue under SLAs | E | $2,000,000 |
| Credit percentage triggered | C | 5% |
| Surviving-session usage unrecorded | R | 65% |
| Reconciliation and dispute labor | D | $5,000 |
| Engineering and call-center response | I | $30,000 |
| Incremental churn, affected subscribers | X | 0.1% |
| Customer acquisition cost | K | $150 |
| Regulatory and communications cost | G | $10,000 |
The formula
Cost of one hour = (H × U × S × P + L) + (E × C) + (H × U × (1 − S) × R + D) + I + (N × S × X × K) + G, where H = N × A ÷ 730 is hourly recurring revenue.
Terms, in order: direct revenue, SLA credits, accounting leakage (on surviving sessions), incident response, churn replacement, regulatory cost.
Worked example
Illustrative, not any operator’s data: a 2-million-subscriber broadband and Wi-Fi operator with the inputs above (H = $68,493).
| Category | Result |
| Direct revenue loss | $21,400 |
| SLA credits | $100,000 |
| Accounting leakage | $11,700 |
| Incident response | $30,000 |
| Churn replacement | $150,000 |
| Regulatory and communications | $10,000 |
| Total, one peak hour | ≈ $323,000 |
Direct revenue is under 7% of the total; credits and churn are about three quarters. Off-peak, with the multiplier at 1.0 and the churn and incident lines halved, the same operator lands near $230,000. A partial outage affects fewer sessions but lasts longer, and churn, support, and leakage scale with duration, which is why MTTR belongs in the model.
Network Downtime Cost Benchmarks, and Why They Understate Telecom Exposure
Uptime Institute’s Annual Outage Analysis 2026 reports that 57% of respondents to its 2025 survey said their most recent major outage cost more than $100,000, up from 54% a year earlier. Another survey, which polled more than 1,000 firms worldwide, found that an hour of downtime exceeds $300,000 for over 90% of mid-size and large enterprises, and that 41% put it from $1 million to more than $5 million.
Neither is a telecom number: Uptime’s respondents run data centers, ITIC’s are enterprises across sectors, and neither models prepaid share, SLA exposure, accounting-record loss, or regulatory reporting. They show finance that six-figure hours are normal. The most-quoted figure, $5,600 per minute, is a cross-industry average from a 2014 Gartner analyst blog post, never a telecom number.
What Actually Causes Authentication Outages?
Six recurring causes:
- Active-passive deployment with slow or untested failover
- Expired certificates on EAP-TLS (Extensible Authentication Protocol with Transport Layer Security), RadSec, or Diameter-over-TLS
- Database locks, replication lag, or storage exhaustion
- A configuration push applied to every node at once, with no canary
- Capacity exhaustion during a re-authentication storm
- End-of-life software and upstream dependencies: the LDAP (Lightweight Directory Access Protocol) directory or HSS (Home Subscriber Server)
Active-passive looks resilient on a slide. In practice the standby has not taken production load since the last drill, its configuration has drifted, and failover takes minutes. RFC 3539 describes the watchdog and retransmission timers clients use to decide a server is dead; failover slower than those timers is, from the subscriber’s side, an outage. The quietest risk is upstream: when the directory or HSS holding the credentials stops answering, the AAA server is up and still cannot admit anyone, so upstream identity-store dependencies need their own failover plan.
Reducing the Number: Availability Levers vs MTTR Levers
Every term in the formula is multiplied, implicitly, by how likely the hour is and how long it lasts. Availability levers reduce the first, MTTR levers the second; most mitigation plans fund only the first.
| Cost driver | Availability lever (how often) | MTTR lever (how long) |
| Direct revenue | Geo-redundant active-active | Failover faster than client retry timers |
| SLA credits | Load balancing, no single point of failure | Alert on reject rate and latency |
| Accounting leakage | NAS-side buffering; server-side de-duplication | Replay and reconciliation tooling |
| Incident response | Cached authorizations during upstream failure | Observability that shows which hop failed |
| Churn | Containerized deployment with node failover | Rate limiting so a recovering server can recover |
| Regulatory | Design to 99.999%; monthly budget in minutes | A rehearsed runbook |
Graceful degradation is the least discussed lever: an AAA that keeps admitting known subscribers from cache during an identity-store failure buys recovery time. Alepo AAA Server is a carrier-grade AAA server designed for five nines (99.999%) availability; it runs geo-redundant active-active, including as containerized AAA on Kubernetes. It is one option, and the design matters more than the vendor. Diagnosis is often the longest part of recovery. AI-assisted detection of authentication anomalies shortens it, and rate limiting on recovery keeps the re-authentication storm from becoming the second incident.
Turning the Cost Into a Business Case
Present the formula with your real inputs, a low and a high case, and every assumption in a footnote. Then compare expected annual cost at your current availability tier against the target: 99.9% permits about 8.8 hours of downtime a year, 99.99% about 53 minutes, 99.999% about 5 minutes. At $300,000 an hour, moving from 99.9% to 99.999% removes roughly $2.6 million of expected annual exposure.
That number anchors the build-versus-buy decision for your AAA stack, the business case for consolidating AAA systems, and the security case for AAA modernization. This page prices the alternative.
Conclusion
See how a carrier-grade AAA server handles site failover, accounting buffering, and a re-authentication storm on a live system, with your own session timeouts and traffic profile as the inputs. Book a demo
Frequently Asked Questions
Q1. What happens to subscribers who are already online when the AAA server goes down?
Existing sessions keep working until their next re-authentication, interim accounting update, or session timeout; new connections fail immediately. Short-timeout sessions (Wi-Fi, 802.1X, prepaid data) drop first; fixed broadband lasts longest.
Q2. How much does one hour of network downtime cost a telecom operator?
There is no single number. Add direct revenue lost, SLA credits, unbilled usage, incident response, churn replacement, and regulatory cost, scaled by subscribers affected and a peak multiplier. The illustrative example above lands near $323,000 at peak, three quarters of it credits and churn.
Q3. Can lost RADIUS accounting records really cause revenue loss?
Yes. RADIUS accounting (RFC 2866) relies on Stop and Interim-Update records reaching the server. Records lost while accounting is down leave usage-based plans unbillable or reconstructed, which invites disputes. NAS-side buffering and server-side de-duplication limit it.

