Authentication, Authorization and Accounting (AAA) scalability is seven dimensions, not one throughput number. Most AAA platforms are bought on a transactions-per-second (TPS) figure and fail on something else: fixed wireless access, an IoT fleet or a 5G standalone core arrives, and the platform runs out of session-table memory, protocol coverage, or the ability to add a node without a maintenance window. Here is a vendor-neutral way to find which limit your network hits first, with a sizing method, a checklist and an RFP scorecard you can reuse.
What Makes an AAA Platform Scalable?
A scalable AAA platform sustains authentication, authorization and accounting performance as load grows along seven dimensions: transaction throughput, concurrent sessions, subscriber and device volume, protocol and access-type breadth, policy complexity, geographic distribution and operational scale. Carrier-grade platforms add capacity horizontally, keep latency stable under burst events, and support RADIUS, Diameter and TACACS+ on one system. “Carrier-grade” means behavior under load, not a vendor tier.
The Seven Dimensions of AAA Scalability
| Dimension | What to measure | Growth driver that stresses it first |
| Throughput and latency under load | Sustained and burst TPS with accounting on; p99 latency at peak | Re-authentication storms, FWA CPE re-attach |
| Concurrent sessions and state | Active sessions; session-table memory per node; state rebuild time | FWA, IoT fleets with long-lived sessions |
| Subscriber and device volume | Identity-store lookup time at 2× and 5× current records; provisioning rate | IoT onboarding, MVNO growth |
| Protocol and access-type breadth | Protocols terminated natively; EAP methods; 3GPP interfaces | Wi-Fi offload, 5G SA, enterprise TACACS+ |
| Policy and integration complexity | Backend round trips per request; time to add an integration | Converged offers, private 5G, slicing |
| Geographic and topological distribution | Sites; failover time; cross-site replication lag | Multi-region expansion, data-residency rules |
| Operational scale | Time to add a node; config drift; time to diagnose an auth failure | Any growth program once the team stops growing |
1. Throughput and latency under load
Measure sustained TPS with accounting writes on, at a p99 latency inside your network access server (NAS) retransmit timer. Alepo AAA’s published figures are 36,000+ TPS with sub-millisecond response; validate any vendor’s figure against your own peak profile.
2. Concurrent sessions and session state
Every live session costs memory and replication traffic for duplicate detection, accounting correlation and Change-of-Authorization (CoA), even when transaction rate is flat. Ask how long state takes to rebuild after a restart.
3. Subscriber and device volume
The identity store behind the AAA usually slows first. Test the full chain, including the HSS, UDM, LDAP or SQL backend, at five times today’s record count.
4. Protocol and access-type breadth
Protocol coverage is a scalability dimension because the alternative is a second platform. Alepo’s answer is a single platform for RADIUS, Diameter and TACACS+ with the full EAP family; count what is terminated natively, not what a gateway can reach.
5. Policy and integration complexity
Every backend round trip is latency on every request: a broadband profile might cost one lookup, a converged offer with OCS quota and a PCRF decision over Gx five. Good platforms evaluate policy in memory.
6. Geographic and topological distribution
Losing a site has to be invisible to subscribers. Ask how session and accounting state stay consistent between sites and when failover was last exercised; the patterns are in our high-availability guide.
7. Operational scale
Measure it in time: to add a node, to roll a configuration change, to get from complaint to root cause. Platforms built for it expose metrics and logs to your stack, run configuration through CI/CD, and increasingly ship AI-assisted anomaly detection.
Vertical vs Horizontal Scaling for AAA: Why Stateful Authentication Changes the Math
Vertical scaling means a bigger node and a maintenance window; horizontal means more nodes behind a load balancer, added while the service runs. Every vendor claims the second; it is only true if the platform has solved the state problem.
RADIUS is stateless at the packet level (RFC 2865); the server behind it is not. Four kinds of state live there: the active-session table for duplicate-login detection and CoA; accounting correlation tying Interim-Update and Stop records (RFC 2866) to their Start; EAP conversation state; and Diameter peer state. If that state sits in one process, a second node splits the session table and breaks CoA for half the requests.
Three patterns solve it: sticky routing via the RADIUS State attribute (simple, but a single point of failure for that node’s sessions), an externalized store (any node serves any request, at a round trip per lookup), or replicated in-memory state (local reads, short staleness window). Ask which pattern each kind of state uses and its measured latency cost.
How to Size an AAA Platform: A Four-Step Method
Step 1 – Baseline peak rates
Record the sustained busiest hour from two or more weekly cycles of AAA and NAS logs, keeping authentication and accounting separate. Accounting is usually larger: 400,000 concurrent sessions on a 15-minute interim timer produce roughly 445 accounting transactions per second before a single new login.
Step 2 – Apply event multipliers
Events, not the baseline, break AAA platforms: power restoration or BNG restarts, CPE firmware pushes, mass roaming events, upstream outage recovery and interim-interval changes. Estimate a multiplier for each from your incident history.
Illustrative, invented numbers: an operator at a steady 450 TPS loses a regional BNG pair and 150,000 sessions re-attach over eight minutes. That is roughly 310 Access-Requests and 310 Accounting-Starts per second before retransmits, and about 800 TPS once a quarter of the NAS fleet retransmits. That storm lands on top of the 450 baseline, so the platform sees roughly 1,250 TPS, or 2.8× baseline. A platform sized on the baseline extends the outage past the fix.
Step 3 – Model growth by access type
Access types load the AAA differently, so model each separately, then sum.
| Access type | Authentication | Accounting | Sessions | Protocol added |
| Fixed wireless access (CPE) | Moderate; re-attach heavy | Strong | Long-lived | RADIUS or Diameter from PGW/SMF or BNG |
| IoT fleet on a private APN/DNN | By device count | Modest | Very long idle | RADIUS from PGW/SMF; EAP-TLS |
| Wi-Fi offload / OpenRoaming | EAP-heavy | Grows | Short, churning | RadSec, EAP-SIM/AKA/AKA’, SWx |
| 5G SA data sessions | Secondary only | Per DN-AAA policy | Per PDU session | RADIUS/Diameter from SMF |
| Enterprise device administration | Per-command | Per-command | Low | TACACS+ |
Step 4 – Set headroom and load-test
Apply a 30–50 percent buffer above the forecast peak, then test baseline plus worst storm with accounting on, replication on and the real identity store behind it; repeat before each access-type launch.
Future-Proofing Checklist: 12 Capabilities to Require in Your Next AAA Platform
Require these in writing and verify each in a demo or proof of concept (PoC).
| # | Capability | How to verify |
| 1 | RADIUS, Diameter and TACACS+ terminated natively on one platform | One authentication per protocol against the same subscriber record |
| 2 | RadSec (RADIUS over TLS, RFC 6614) at scale | Load-test RadSec, not just UDP RADIUS |
| 3 | Full EAP family, including EAP-TLS 1.3 (RFC 9190) and EAP-AKA’ (RFC 9048) | A SIM device and a certificate device on the same server |
| 4 | 3GPP non-3GPP access interfaces (SWa, STa, SWm, S6b per TS 29.273) | Live SWx lookup and S6b exchange |
| 5 | DN-AAA secondary authentication for 5G SA (TS 23.501, TS 23.502) | Authorize a PDU session from an SMF against an external identity |
| 6 | Horizontal scaling with externalized or replicated state | Add a node under load; CoA still works for existing sessions |
| 7 | Active-active geo-redundancy with exercised failover | Kill a site under load; measure sessions dropped and failover time |
| 8 | Rate limiting, backpressure and storm protection | Replay a recorded re-attach storm at 3× |
| 9 | Same software on bare metal, VMs, Kubernetes and managed service | Ask for the upgrade path between models |
| 10 | Open REST APIs and configuration as code | Provision and change policy via the API, then roll back |
| 11 | Exportable metrics, traces and logs; AI-assisted anomaly detection | Connect to your observability stack during the PoC |
| 12 | IPv6 throughout, including Framed-IPv6 accounting attributes | Authenticate and account an IPv6-only session |
Deployment Models Compared: Bare Metal, VMs, Kubernetes and Managed Service
| Model | Scaling pattern | Best fit |
| Bare metal | Vertical first; hosts added in a maintenance window | Latency-sensitive packet-core adjacency; regulated on-premises estates |
| Virtual machines | Horizontal by cloning images; hypervisor-level HA | Operator private cloud; brownfield replacements |
| Kubernetes | Pod autoscaling on custom metrics; rolling upgrades with no outage | Cloud native teams; 5G core alignment |
| Managed service | Vendor scales to an SLA; operator sees capacity, not nodes | Lean teams; multi-access consolidation |
A managed service moves the scaling problem rather than removing it, so the contract should state the TPS, session and storm profile the vendor commits to. Alepo AAA ships the same software across all four models, so the choice can change over a contract.
Eight Signs Your Current AAA Is Running Out of Road
- NAS retransmit counters rise while success rate looks fine.
- p99 latency pulls away from the median.
- Accounting writes queue up in the evening and drain overnight.
- Identity-store lookups take longer every quarter.
- Restarts take longer as the session table grows.
- Adding capacity needs a weekend.
- The last storm test was at install.
- Every integration was a project.
AAA Platform Evaluation Scorecard
Score each criterion from 1 to 5 using the checklist’s verification column, multiply each score by its weight, and add the results. Divide that total by 5 to get a score out of 100: all 3s comes to 60, all 5s to 100. The weights sum to 100 and favor growth; adjust them and state them in the RFP.
| # | Criterion | Weight | What a 5 looks like |
| 1 | Throughput and latency under load | 15 | Sustained TPS with accounting on; p99 inside the NAS timer at a 3× storm |
| 2 | Session state and horizontal scaling | 15 | Node added under load with no CoA or accounting loss |
| 3 | Subscriber and device volume | 10 | Tested at 5× current records with the real identity store |
| 4 | Protocol and access-type breadth | 15 | RADIUS, Diameter, TACACS+, full EAP family and TS 29.273 interfaces native |
| 5 | Policy and integration complexity | 10 | Policy as configuration; open APIs; integration added in the PoC |
| 6 | Geographic distribution and resilience | 10 | Active-active across sites; failover exercised in the PoC |
| 7 | Operational scale | 10 | Metrics and logs into your stack; configuration in CI/CD |
| 8 | Deployment flexibility | 5 | Same software on all four models with a documented path between them |
| 9 | Standards compliance and openness | 5 | RFCs and 3GPP specs cited per feature; no proprietary attributes |
| 10 | Five-year cost of ownership | 5 | Transparent scaling increments; no re-platform for growth |
| Total weight | 100 |
Treat a score under 60 as a signal that operational gaps are likely to outweigh any licensing advantage; 60 to 80 is a conditional shortlist with gaps written into the PoC plan.
Book a demo. See how one AAA platform scales across RADIUS, Diameter and TACACS+ for operators.
Frequently Asked Questions
Q1. What is a carrier-grade AAA platform? One defined by behavior under load: authentication latency inside the NAS retransmit timer during storms, no sessions dropped when a site is lost, RADIUS, Diameter and TACACS+ terminated natively, and upgrades without an outage.
Q2. Can a stateful AAA server scale horizontally? Yes, provided session, accounting and EAP conversation state are externalized or replicated so any node can serve any request. The trade-off is latency: an external store adds a round trip per lookup; replicated state adds a short staleness window.
Q3. Do I still need RADIUS and Diameter after moving to 5G SA? Yes. Fixed broadband, Wi-Fi offload, Wi-Fi calling and device administration still run on RADIUS, RadSec and TACACS+, a 4G core alongside the 5G one still needs Diameter, and inside the 5G core the AAA takes on secondary authentication.

