An Authentication, Authorization, and Accounting (AAA) server that is fast enough today can be the reason a network stalls in eighteen months. Subscriber growth, 5G session density, Internet of Things (IoT) device counts, and Wi-Fi offload surges all land on the same authentication path. When that path saturates, nothing downstream matters. This post defines AAA server scalability in measurable terms and walks through three growth scenarios that break undersized platforms. It then shows the capacity math and gives you an evaluation table for comparing how vendors scale.
An AAA server is scalable when it can handle growing authentication, authorization, and accounting demand, meaning more subscribers, more concurrent sessions, or new access technologies like 5G, without added latency or downtime. Scalability is typically achieved through horizontal scaling (adding nodes), containerized cloud-native deployment, and active-active high-availability architecture.
Most buyers’ guides treat scalability as one bullet in a list of twelve. This post treats it as the whole subject, because it is the criterion that separates platforms built for carrier scale from platforms sized for the last network the vendor sold into.
Why scalability is the criterion that separates AAA servers
Feature lists for AAA servers converge. Every serious platform speaks RADIUS (Remote Authentication Dial-In User Service) and Diameter, supports the Extensible Authentication Protocol (EAP) methods you need, and integrates with your subscriber store. What does not converge is what happens to those features when the request rate triples. That is the question an evaluation should be organized around.
The reason is where AAA sits: in the connection path of every subscriber, device, and session on the network. The broadband network gateway (BNG) asks it whether a fiber customer can come online. The packet core asks it whether a roaming device is who it claims to be. The Wi-Fi controller asks it whether a phone can offload. Operators scale each of those systems independently, often without checking that the AAA layer behind them kept pace.
The result is a failure mode that is easy to miss until it is expensive. The AAA process is up and the dashboards are green, but response times have drifted from a few milliseconds to a few seconds. That is past the point where network access servers (NAS) stop waiting. They retransmit, raising the load on a server that was already behind. Success rate drops because capacity has run out, not because anything broke, and the network operations center (NOC) spends the first hour looking for a fault that does not exist.
What AAA server scalability actually means
AAA server scalability is the platform’s ability to sustain a growing rate of authentication, authorization, and accounting transactions, and a growing number of concurrent sessions, while keeping response latency inside the NAS timeout. It is measured in transactions per second (TPS) under a realistic traffic mix. Subscriber count is an input to that number, never a substitute for it.
Three numbers do most of the work. Sustained TPS is the request rate the system can absorb all day, with accounting writes enabled, without latency creeping up. Burst TPS is what it can absorb for a few minutes when a region’s worth of sessions reconnects at once. Concurrent sessions is how many active sessions the system is tracking state for, which drives memory, replication traffic, and the size of the session store. Ask for all three. A vendor who quotes one of them has described a third of the platform.
The traffic mix matters as much as the rate. A simple Password Authentication Protocol (PAP) authentication is one request and one response. An EAP-AKA authentication for a Wi-Fi calling session is several round trips, each holding conversation state on the server between messages. Interim accounting updates (RFC 2866) are writes, and writes cost more than reads. Diameter messages (RFC 6733) run over long-lived connections with their own watchdog traffic. A TPS figure benchmarked on PAP against a cached user store says little about the same platform on EAP with live accounting and a replicated session database behind it.
Then there is how capacity is added when you need more. Vertical scaling means a bigger server: more cores, more memory, a faster disk. It is simple, it works until it does not, and the ceiling arrives without warning because the next step up is a hardware replacement. Horizontal scaling means more servers sharing the load. There is no fixed ceiling, but there is one demand: any node must be able to serve any request, so session state has to live in a shared, replicated store that every node can reach. The second model is harder to build. It is also the one that keeps working as the network grows.
Three growth scenarios that break undersized AAA servers
Platforms that fail on scale rarely fail on average load. They fail on the shape of the growth, which is almost never a smooth line. The three below come up again and again in post-incident reviews.
| Scenario | What changes in the load | What breaks first | What to ask a vendor |
| Subscriber growth spike (FTTH rollout, MVNO launch, acquisition) | Steady-state TPS and concurrent sessions step up over months; accounting write volume rises in proportion | Session store capacity and database write throughput; vertical-only platforms hit a hardware ceiling mid-rollout | How is capacity added, and does adding it require downtime or a re-architecture? |
| 5G network slicing and IoT density | Sessions per device multiply (one device, several slices); device counts grow faster than human subscribers; slice-specific and secondary authentication add EAP exchanges | Memory for session state; EAP state under concurrent slice authentications; anything sized by “subscribers” rather than “sessions” | What is the concurrent session limit, and is it per node or per cluster? |
| Seasonal and event Wi-Fi offload surge | Burst TPS climbs by a large multiple at venues, holidays, and when a mobile cell sheds load; EAP-SIM/AKA round trips hold state under burst | Burst headroom and EAP state handling; retransmit amplification when latency crosses the NAS timeout | What burst multiple over daily peak has the platform sustained in production, and for how long? |
Subscriber growth is the scenario everyone plans for and still underestimates, because the growth that hurts is step-function growth. A fiber operator that acquires a neighboring network, or a Mobile Virtual Network Operator (MVNO) that signs a distribution deal, adds a region in a quarter, not a percentage a year. If the AAA platform scales vertically, that quarter ends with a hardware procurement in the middle of a migration. Consolidation sharpens it. An operator replacing several legacy AAA systems with one platform is asking that platform to carry the sum of all of them from day one.
5G and IoT change the unit of measurement. In 4G, one subscriber was roughly one session. In 5G Standalone (SA), one device can hold multiple protocol data unit (PDU) sessions across network slices. Each slice is identified by its own Single Network Slice Selection Assistance Information (S-NSSAI) value, per 3GPP TS 23.501. Two of the authentication flows defined in TS 33.501 terminate on an AAA server: network slice-specific authentication and authorization (NSSAA), run per slice over EAP, and secondary authentication toward an external data network. IoT multiplies the device count on top of that. An AAA platform sized by subscriber count will meet a session count several times larger, and an EAP exchange for each. We cover the slicing side in our post on AAA’s role in 5G network slicing.
Wi-Fi offload produces the sharpest spikes. A stadium filling up, a holiday peak at an airport, or a macro cell shedding load to carrier Wi-Fi all produce authentication bursts at many times the daily average. Each authentication is an EAP-SIM or EAP-AKA exchange that holds state across several round trips. If the burst pushes latency past the point where the controllers give up waiting, they retransmit, and the platform now faces the original burst plus the retries. We have written before about AAA built for traffic spikes; this is where “we tested it at peak” turns out to have meant last year’s peak.
Capacity planning: sizing an AAA server for the next three years
Capacity planning for AAA comes down to translating subscribers and sessions into transactions per second, then adding headroom for the burst case and for growth. The arithmetic is not complicated, but it has to be done with your traffic mix rather than a vendor’s.
Take a worked example, with round numbers and no claim to represent any particular network. An operator with one million broadband subscribers runs interim accounting every 15 minutes. That alone is 1,000,000 ÷ 900 seconds, or roughly 1,100 accounting transactions per second, all day, all writes. Add a Wi-Fi footprint where sessions re-authenticate hourly, say 300,000 concurrent Wi-Fi sessions, and that is another 80 or so EAP authentications per second, each a multi-round-trip exchange. At steady state, this operator needs comfortable room above about 1,200 TPS with a write-heavy mix.
Now the burst. A BNG serving 200,000 subscribers reboots, and those sessions reconnect within about two minutes. That is 1,700 authentications per second on top of steady state, during an evening peak when accounting load is already highest. Each reconnect also sends an accounting start, so the true transaction count is higher still. The platform that was fine at 1,200 TPS now needs to answer at least 3,000 TPS for two minutes without crossing the NAS timeout, or the retransmits turn two minutes into twenty.
Then growth. Apply the three-year subscriber plan, apply the session multiplier if 5G SA or IoT is on the roadmap, and ask what the platform looks like at that number. If the answer is “add nodes,” the exercise is done. If the answer is “a bigger server,” ask what happens the year after. Capacity planning for a horizontally scalable platform is a budgeting exercise. For a vertically scaled one, it is a countdown.
How to evaluate an AAA server’s scaling architecture
The scaling architecture of an AAA server shows up in four design decisions: how it adds capacity, how it handles failover, where it keeps session state, and whether it runs across regions. Each one can be checked against documentation and demonstrated in a proof of concept. That makes them the most useful questions in an evaluation.
| Dimension | Scales well | Scales poorly | How to verify |
| Capacity model | Horizontal: containerized nodes added under load, orchestrated on Kubernetes or equivalent, no fixed ceiling | Vertical: capacity bound to the largest supported server; expansion means hardware replacement and a maintenance window | Ask to watch a node added to a running cluster during the PoC, with TPS on a graph |
| Failover model | Active-active: every node and site carries live traffic; node loss costs headroom, not availability | Active-passive: a standby promoted on failure; the failover path is the least-exercised code in the system | Ask to watch a node killed under load; our high availability guide covers what to measure |
| Session state | Externalized to a replicated store; any node serves any request; EAP conversation state survives node loss | Held in process memory; sessions pinned to a node; loss of the node means re-authentication for everyone on it | Ask where EAP state lives mid-conversation and what happens to it when that node dies |
| Geographic model | Multi-region, both sites active, state replicated across the inter-site link | Single site, or a disaster recovery site that has never carried production traffic | Ask how long the secondary site has been serving live traffic |
State and failover are where vendor claims most often outrun architecture.
Horizontal scaling is only real if state is externalized. A platform can run as containers and still pin each session to the node that started it, so “add a node” helps new sessions and does nothing for the ones in flight. Ask the specific question: if an EAP exchange is on message three of five when its node disappears, does it complete on another node, or does the subscriber start again? Our post on containerized, cloud-native AAA deployment goes further into what the container model changes and what it does not.
Active-active and horizontal scaling are the same design seen from two angles. A pool of nodes sharing load every day is both the scaling mechanism and the redundancy mechanism, because adding a node for capacity and losing one to failure are the same operation in reverse. An active-passive pair on a vertically scaled server has neither property. Our guide to high availability AAA server design covers the failover side in depth.
What is an AAA server?
An AAA server answers three questions for every connection. Who is this subscriber or device? What may they access, at what speed, under what policy? How much did they use, for billing and audit? It communicates with network access servers over RADIUS (RFC 2865), with mobile cores over Diameter, and with network devices over TACACS+ for administrative access. If you want the protocol-level walkthrough, start with our RADIUS beginner’s guide. This post assumes that ground and concentrates on what happens to the AAA layer as the network grows.
Where Alepo stands on scale
The “scales well” column above describes how the Alepo AAA Server is built, so here is how we answer our own four questions. Capacity is added as containerized nodes under live load, orchestrated on Kubernetes. The same platform runs on virtual machines, bare metal, or private cloud where regulation or latency requires it. Session state sits in stateless storage backed by database persistence, with real-time database replication and N+1 or N+N redundancy, so any node can take any request. Deployments are geo-redundant, with state replicated between sites.
On the numbers: the platform is engineered for 99.999% availability and is rated at 36,000+ transactions per second (AAA Server datasheet). RADIUS, Diameter, and TACACS+ terminate on the same stack, so an operator consolidating fixed, mobile, and enterprise access scales one platform rather than three.
Alepo AAA deployments span Tier-1 and Tier-2 fixed and mobile operators across the Middle East, Europe, Africa, Asia-Pacific, and Latin America, with more than ten carrier AAA deployments globally. Much of that work has been migration, where the new platform had to carry full production load from cutover day.
The criteria in this post are yours to apply to any vendor, including us. Book a demo and we will run your subscriber plan, traffic mix, and burst scenario against the platform live, with TPS and latency on screen. If 5G is the specific scaling question in front of you, the network slicing post is the better next read.
Frequently asked questions
Q1. What is an AAA server?
An AAA server authenticates subscribers and devices, authorizes what they can access, and records their usage for billing and audit. It serves network access servers over RADIUS, mobile cores over Diameter, and network device administration over TACACS+. The section above covers the three functions in more detail.
Q2. Why does scalability matter for an AAA server?
Because RADIUS runs over UDP, a network access server that gets no answer in time does not wait; it retransmits. A saturated AAA server therefore receives extra load precisely because it is slow, and the queue compounds until subscribers cannot get online, however healthy the rest of the network is. Scalability is what keeps response time inside that timeout as subscribers, sessions, and access technologies grow.
Q3. How many transactions per second should a scalable AAA server support?
The right TPS figure is derived from your own traffic: subscribers, accounting interval, re-authentication interval, EAP mix, and the size of the largest NAS whose sessions could reconnect at once. Many operators land in the low thousands of TPS at steady state, with burst headroom several times that. A carrier-grade platform should state its sustained TPS with accounting enabled and the test conditions behind it. Ask for the figure under your mix, not the vendor’s.
Q4. What is the difference between horizontal and vertical scaling for AAA infrastructure?
Vertical scaling adds capacity by moving to a larger server, which is simple but has a fixed ceiling and usually requires downtime to cross it. Horizontal scaling adds capacity by adding nodes that share load, which has no fixed ceiling but requires session state to live in a shared, replicated store so any node can serve any request. Horizontal scaling is the carrier-grade model.
Q5. Can an AAA server scale for 5G network slicing?
Yes, provided it is sized by sessions rather than subscribers and scales horizontally. In 5G SA, one device can hold multiple PDU sessions across slices, and slice-specific authentication (NSSAA) runs a separate EAP exchange per slice, so session count and authentication volume both grow faster than the subscriber count. The platform needs externalized session state and EAP handling that does not degrade as slice count grows.
Q6. What happens if an AAA server cannot scale with subscriber growth?
Response latency rises until it crosses the NAS timeout, at which point network access servers retransmit and multiply the load. Authentication success rate drops, new subscribers cannot connect, existing sessions fail at re-authentication, and accounting records are delayed or lost. Because the server process is still running, the incident often presents as a mystery outage, and the NOC looks for a fault before anyone checks capacity.
Q7. Is a containerized AAA server more scalable than a traditional deployment?
Usually, but only if the architecture behind the container is stateless at the node level. Containers make adding and replacing nodes fast and routine, which is the mechanism horizontal scaling depends on. A platform that runs as containers but pins session state to individual nodes has the packaging of horizontal scaling without the benefit. Ask where session and EAP state lives before crediting the container label.
Q8. How does geo-redundancy relate to AAA server scalability?
Geo-redundancy and scalability share one architecture: a pool of active nodes across sites that share load every day. Adding a site for capacity and surviving the loss of one are two uses of that single design. A disaster recovery site that sits idle contributes nothing to capacity, and it has also never proved it can carry production load.

