Scalable AAA Platform: How to Future-Proof Telecom Authentication

Authentication, Authorization and Accounting (AAA) scalability is seven dimensions, not one throughput number. Most AAA platforms are bought on a transactions-per-second (TPS) figure and fail on something else: fixed wireless access, an IoT fleet or a 5G standalone core arrives, and the platform runs out of session-table memory, protocol coverage, or the ability to add a node without a maintenance window. Here is a vendor-neutral way to find which limit your network hits first, with a sizing method, a checklist and an RFP scorecard you can reuse.

What Makes an AAA Platform Scalable?

A scalable AAA platform sustains authentication, authorization and accounting performance as load grows along seven dimensions: transaction throughput, concurrent sessions, subscriber and device volume, protocol and access-type breadth, policy complexity, geographic distribution and operational scale. Carrier-grade platforms add capacity horizontally, keep latency stable under burst events, and support RADIUS, Diameter and TACACS+ on one system. “Carrier-grade” means behavior under load, not a vendor tier.

The Seven Dimensions of AAA Scalability

Dimension What to measure Growth driver that stresses it first
Throughput and latency under load Sustained and burst TPS with accounting on; p99 latency at peak Re-authentication storms, FWA CPE re-attach
Concurrent sessions and state Active sessions; session-table memory per node; state rebuild time FWA, IoT fleets with long-lived sessions
Subscriber and device volume Identity-store lookup time at 2× and 5× current records; provisioning rate IoT onboarding, MVNO growth
Protocol and access-type breadth Protocols terminated natively; EAP methods; 3GPP interfaces Wi-Fi offload, 5G SA, enterprise TACACS+
Policy and integration complexity Backend round trips per request; time to add an integration Converged offers, private 5G, slicing
Geographic and topological distribution Sites; failover time; cross-site replication lag Multi-region expansion, data-residency rules
Operational scale Time to add a node; config drift; time to diagnose an auth failure Any growth program once the team stops growing

1. Throughput and latency under load

Measure sustained TPS with accounting writes on, at a p99 latency inside your network access server (NAS) retransmit timer. Alepo AAA’s published figures are 36,000+ TPS with sub-millisecond response; validate any vendor’s figure against your own peak profile.

2. Concurrent sessions and session state

Every live session costs memory and replication traffic for duplicate detection, accounting correlation and Change-of-Authorization (CoA), even when transaction rate is flat. Ask how long state takes to rebuild after a restart.

3. Subscriber and device volume

The identity store behind the AAA usually slows first. Test the full chain, including the HSS, UDM, LDAP or SQL backend, at five times today’s record count.

4. Protocol and access-type breadth

Protocol coverage is a scalability dimension because the alternative is a second platform. Alepo’s answer is a single platform for RADIUS, Diameter and TACACS+ with the full EAP family; count what is terminated natively, not what a gateway can reach.

5. Policy and integration complexity

Every backend round trip is latency on every request: a broadband profile might cost one lookup, a converged offer with OCS quota and a PCRF decision over Gx five. Good platforms evaluate policy in memory.

6. Geographic and topological distribution

Losing a site has to be invisible to subscribers. Ask how session and accounting state stay consistent between sites and when failover was last exercised; the patterns are in our high-availability guide.

7. Operational scale

Measure it in time: to add a node, to roll a configuration change, to get from complaint to root cause. Platforms built for it expose metrics and logs to your stack, run configuration through CI/CD, and increasingly ship AI-assisted anomaly detection.

Vertical vs Horizontal Scaling for AAA: Why Stateful Authentication Changes the Math

Vertical scaling means a bigger node and a maintenance window; horizontal means more nodes behind a load balancer, added while the service runs. Every vendor claims the second; it is only true if the platform has solved the state problem.

RADIUS is stateless at the packet level (RFC 2865); the server behind it is not. Four kinds of state live there: the active-session table for duplicate-login detection and CoA; accounting correlation tying Interim-Update and Stop records (RFC 2866) to their Start; EAP conversation state; and Diameter peer state. If that state sits in one process, a second node splits the session table and breaks CoA for half the requests.

Three patterns solve it: sticky routing via the RADIUS State attribute (simple, but a single point of failure for that node’s sessions), an externalized store (any node serves any request, at a round trip per lookup), or replicated in-memory state (local reads, short staleness window). Ask which pattern each kind of state uses and its measured latency cost.

How to Size an AAA Platform: A Four-Step Method

Step 1 – Baseline peak rates

Record the sustained busiest hour from two or more weekly cycles of AAA and NAS logs, keeping authentication and accounting separate. Accounting is usually larger: 400,000 concurrent sessions on a 15-minute interim timer produce roughly 445 accounting transactions per second before a single new login.

Step 2 – Apply event multipliers

Events, not the baseline, break AAA platforms: power restoration or BNG restarts, CPE firmware pushes, mass roaming events, upstream outage recovery and interim-interval changes. Estimate a multiplier for each from your incident history.

Illustrative, invented numbers: an operator at a steady 450 TPS loses a regional BNG pair and 150,000 sessions re-attach over eight minutes. That is roughly 310 Access-Requests and 310 Accounting-Starts per second before retransmits, and about 800 TPS once a quarter of the NAS fleet retransmits. That storm lands on top of the 450 baseline, so the platform sees roughly 1,250 TPS, or 2.8× baseline. A platform sized on the baseline extends the outage past the fix.

Step 3 – Model growth by access type

Access types load the AAA differently, so model each separately, then sum.

Access type Authentication Accounting Sessions Protocol added
Fixed wireless access (CPE) Moderate; re-attach heavy Strong Long-lived RADIUS or Diameter from PGW/SMF or BNG
IoT fleet on a private APN/DNN By device count Modest Very long idle RADIUS from PGW/SMF; EAP-TLS
Wi-Fi offload / OpenRoaming EAP-heavy Grows Short, churning RadSec, EAP-SIM/AKA/AKA’, SWx
5G SA data sessions Secondary only Per DN-AAA policy Per PDU session RADIUS/Diameter from SMF
Enterprise device administration Per-command Per-command Low TACACS+

Step 4 – Set headroom and load-test

Apply a 30–50 percent buffer above the forecast peak, then test baseline plus worst storm with accounting on, replication on and the real identity store behind it; repeat before each access-type launch.

Future-Proofing Checklist: 12 Capabilities to Require in Your Next AAA Platform

Require these in writing and verify each in a demo or proof of concept (PoC).

# Capability How to verify
1 RADIUS, Diameter and TACACS+ terminated natively on one platform One authentication per protocol against the same subscriber record
2 RadSec (RADIUS over TLS, RFC 6614) at scale Load-test RadSec, not just UDP RADIUS
3 Full EAP family, including EAP-TLS 1.3 (RFC 9190) and EAP-AKA’ (RFC 9048) A SIM device and a certificate device on the same server
4 3GPP non-3GPP access interfaces (SWa, STa, SWm, S6b per TS 29.273) Live SWx lookup and S6b exchange
5 DN-AAA secondary authentication for 5G SA (TS 23.501, TS 23.502) Authorize a PDU session from an SMF against an external identity
6 Horizontal scaling with externalized or replicated state Add a node under load; CoA still works for existing sessions
7 Active-active geo-redundancy with exercised failover Kill a site under load; measure sessions dropped and failover time
8 Rate limiting, backpressure and storm protection Replay a recorded re-attach storm at 3×
9 Same software on bare metal, VMs, Kubernetes and managed service Ask for the upgrade path between models
10 Open REST APIs and configuration as code Provision and change policy via the API, then roll back
11 Exportable metrics, traces and logs; AI-assisted anomaly detection Connect to your observability stack during the PoC
12 IPv6 throughout, including Framed-IPv6 accounting attributes Authenticate and account an IPv6-only session

Deployment Models Compared: Bare Metal, VMs, Kubernetes and Managed Service

Model Scaling pattern Best fit
Bare metal Vertical first; hosts added in a maintenance window Latency-sensitive packet-core adjacency; regulated on-premises estates
Virtual machines Horizontal by cloning images; hypervisor-level HA Operator private cloud; brownfield replacements
Kubernetes Pod autoscaling on custom metrics; rolling upgrades with no outage Cloud native teams; 5G core alignment
Managed service Vendor scales to an SLA; operator sees capacity, not nodes Lean teams; multi-access consolidation

A managed service moves the scaling problem rather than removing it, so the contract should state the TPS, session and storm profile the vendor commits to. Alepo AAA ships the same software across all four models, so the choice can change over a contract.

Eight Signs Your Current AAA Is Running Out of Road

  1. NAS retransmit counters rise while success rate looks fine.
  2. p99 latency pulls away from the median.
  3. Accounting writes queue up in the evening and drain overnight.
  4. Identity-store lookups take longer every quarter.
  5. Restarts take longer as the session table grows.
  6. Adding capacity needs a weekend.
  7. The last storm test was at install.
  8. Every integration was a project.

AAA Platform Evaluation Scorecard

Score each criterion from 1 to 5 using the checklist’s verification column, multiply each score by its weight, and add the results. Divide that total by 5 to get a score out of 100: all 3s comes to 60, all 5s to 100. The weights sum to 100 and favor growth; adjust them and state them in the RFP.

# Criterion Weight What a 5 looks like
1 Throughput and latency under load 15 Sustained TPS with accounting on; p99 inside the NAS timer at a 3× storm
2 Session state and horizontal scaling 15 Node added under load with no CoA or accounting loss
3 Subscriber and device volume 10 Tested at 5× current records with the real identity store
4 Protocol and access-type breadth 15 RADIUS, Diameter, TACACS+, full EAP family and TS 29.273 interfaces native
5 Policy and integration complexity 10 Policy as configuration; open APIs; integration added in the PoC
6 Geographic distribution and resilience 10 Active-active across sites; failover exercised in the PoC
7 Operational scale 10 Metrics and logs into your stack; configuration in CI/CD
8 Deployment flexibility 5 Same software on all four models with a documented path between them
9 Standards compliance and openness 5 RFCs and 3GPP specs cited per feature; no proprietary attributes
10 Five-year cost of ownership 5 Transparent scaling increments; no re-platform for growth
  Total weight 100  

Treat a score under 60 as a signal that operational gaps are likely to outweigh any licensing advantage; 60 to 80 is a conditional shortlist with gaps written into the PoC plan.

Book a demo. See how one AAA platform scales across RADIUS, Diameter and TACACS+ for operators.

Frequently Asked Questions

Q1. What is a carrier-grade AAA platform? One defined by behavior under load: authentication latency inside the NAS retransmit timer during storms, no sessions dropped when a site is lost, RADIUS, Diameter and TACACS+ terminated natively, and upgrades without an outage.

Q2. Can a stateful AAA server scale horizontally? Yes, provided session, accounting and EAP conversation state are externalized or replicated so any node can serve any request. The trade-off is latency: an external store adds a round trip per lookup; replicated state adds a short staleness window.

Q3. Do I still need RADIUS and Diameter after moving to 5G SA? Yes. Fixed broadband, Wi-Fi offload, Wi-Fi calling and device administration still run on RADIUS, RadSec and TACACS+, a 4G core alongside the 5G one still needs Diameter, and inside the 5G core the AAA takes on secondary authentication.

Want to see how this applies to your business? Let’s talk.

Share the Post:

Latest Posts

Receive the latest news

Subscribe To Our Newsletter

Subscribe to our Newsletter

Receive the latest news

Subscribe To Our Newsletter