AAA Server Stress Testing – How to Validate Carrier-Grade Performance

AAA server stress testing proves that an authentication, authorization and accounting platform holds your peak traffic, survives failures and recovers without losing sessions or accounting records. Most lab tests send simple Password Authentication Protocol (PAP) Access-Requests and stop there, which tells you little about a live network. A useful test starts from your own traffic profile, runs nine defined scenarios and judges the results against acceptance criteria you write down before the first packet is sent.

Every AAA vendor quotes a transactions-per-second (TPS) figure. That number was measured in the vendor’s lab, with the vendor’s traffic mix and backend. Your network runs a different mix of Extensible Authentication Protocol (EAP) methods and a different accounting interval. Its subscriber databases answer at their own speed, and it has its own failure history.

This guide gives you the full method: how to build a traffic profile, the nine scenarios a carrier-grade AAA platform should pass, starting-point acceptance criteria and a side-by-side comparison of Remote Authentication Dial-In User Service (RADIUS, RFC 2865) and Diameter (RFC 6733) load testing tools. It applies to any AAA platform, including ours.

What Is AAA Server Stress Testing?

AAA server stress testing means pushing an authentication, authorization and accounting platform past its expected peak on purpose, to find where it saturates and what it does when it gets there. Load testing is the gentler version: it checks that performance holds at the volume you planned for. A proper stress test layers authentication bursts, accounting surges, slow backends and node failures on top of each other, because that is what a bad night in production looks like. What you want to see is a platform that slows down predictably, sheds load cleanly and recovers without losing sessions or accounting records.

A complete AAA performance testing program uses five test types. Each one answers a question the other four cannot.

Test type Purpose Typical duration Pass signal
Load Confirm performance at planned peak 1–4 hours Latency and error rate inside budget
Stress Find the saturation point and failure mode Until saturation Predictable degradation, no crash
Soak Expose slow leaks and drift 24–72 hours Flat latency, memory and disk use
Spike / storm Survive a sudden mass re-authentication Minutes Queue drains, no retry spiral
Failover Prove recovery under load Per failure event Zero session and record loss

In practice, they form one network stress test run in five passes. The same harness, traffic profile and dashboards serve all of them.

Why the Vendor’s TPS Is Not Your TPS

TPS is meaningful only when you know what a transaction contains. A PAP Access-Request answered from cache and an EAP Subscriber Identity Module (EAP-SIM) authentication that waits on a Home Subscriber Server (HSS) lookup are both “one authentication,” but they cost the server very different amounts of work.

Authentication method cost: PAP vs. MS-CHAPv2 vs. EAP-TLS vs. EAP-SIM/AKA

PAP completes in one request and one response. Microsoft Challenge Handshake Authentication Protocol version 2 (MS-CHAPv2) adds a challenge-response calculation but stays in a single round trip when carried directly in RADIUS rather than inside Protected EAP (PEAP). EAP-Transport Layer Security (EAP-TLS, RFC 5216, updated for TLS 1.3 by RFC 9190) needs several RADIUS round trips and certificate processing for each authentication. EAP-SIM and EAP Authentication and Key Agreement (EAP-AKA; RFC 4186, RFC 4187) need multiple round trips plus an authentication-vector fetch from the HSS or Home Location Register (HLR).

So a server’s “authentications per second” can shift several-fold depending on the method mix. A fair test states the mix next to every TPS figure.

Accounting usually outnumbers authentication

One subscriber session produces Accounting-Request messages of type Start, Interim-Update (at every interval) and Stop (RFC 2866, RFC 2869). On a broadband network with long-lived sessions and short interim intervals, accounting traffic often exceeds authentication traffic. Accounting is also write-heavy, so it stresses disks, databases and billing mediation in ways that authentication tests never touch.

Backend lookups, caching and TLS/RadSec overhead

Most AAA transactions call another system: a Lightweight Directory Access Protocol (LDAP) directory, an SQL subscriber database, an HSS or 5G Unified Data Management (UDM) function, an online charging system or a policy server. The AAA server’s measured latency is partly the backend’s latency. Caching can hide that cost in a short test and expose it in a long one. Transport adds cost too: RadSec (RADIUS over TLS) and Diameter over TLS spend CPU on encryption that plain UDP RADIUS does not.

Start With Your Traffic Profile, Not the Tool

Before choosing any tool, write down what your AAA platform actually has to carry. That’s your traffic profile, built from your own NAS (network access server) and AAA logs, and it comes first because it decides which tools can even generate your traffic.

Pull at least two full weeks of logs so weekday, weekend and evening peaks all show up. Your design peak is the busiest sustained hour in that data, and every scenario later in this guide is measured against it. If you have not sized the platform yet, start with sizing an AAA platform for growth.

The example values below are illustrative only. Replace every one with your measured figures.

Profile input What to measure Illustrative value
Sustained peak TPS Busiest sustained hour, all message types 4,000
Authentication mix Share of each method 50% EAP-SIM/AKA, 30% PAP, 20% EAP-TLS
Accounting ratio Accounting messages per authentication 3:1
Interim interval Acct-Interim-Interval in use 15 minutes
Change of Authorization (CoA) / Disconnect rate Dynamic authorization messages per second 50
Backend latency p95 response from LDAP, HSS, SQL 5–20 ms
NAS count Gateways and controllers sending traffic (see Topology below) 120
Storm multiplier Peak re-authentication rate vs. sustained peak 5×

Authentication mix and EAP round trips

Count authentications by method and by access type. Then count RADIUS round trips per method, since an EAP conversation of several exchanges loads the server more than a single PAP exchange. Your test harness must reproduce both the mix and the round trips.

Accounting profile: Start, Interim-Update, Stop ratios and interval

Record how Start, Interim-Update and Stop messages split, and which interim interval each NAS uses. A RADIUS accounting load test that skips Interim-Updates will understate the write load from long-lived broadband sessions. Also check whether accounting feeds billing mediation in real time. If it does, that path has its own capacity limit.

CoA and Disconnect volume

Change of Authorization and Disconnect messages (RFC 5176) normally travel the other way, from the AAA server out to the NAS. They tend to spike at awkward moments: plan changes, quota exhaustion, fraud blocks. They compete for the same server resources as inbound traffic, so they belong in the profile.

Backend dependencies and realistic latency

List every system the AAA server calls per transaction and its real response time under load. Configure the test backend to respond with that latency. A stub that answers in 0.1 ms makes any AAA platform look fast.

Topology: NAS count, proxies, load balancers and sites

Match the number of NAS clients, not just the total rate. Every BNG, PGW, ePDG and wireless LAN controller is a separate client, and a thousand of them sending four TPS each will test load balancing and connection handling in ways a single client sending four thousand never will. If production has RADIUS proxies, Diameter routing agents or a second site, the lab needs them too.

The Nine Scenarios Every Carrier-Grade AAA Must Pass

These nine scenarios cover the RADIUS performance testing and Diameter load testing a carrier needs before production. Run them in order, because each one builds on the baseline from the first.

# Scenario What it proves Pass looks like
1 Throughput ramp to saturation Where capacity ends Saturation well above design peak
2 Sustained peak soak Stability over time No drift in 24–72 hours
3 Re-authentication storm Recovery from mass reconnect Backlog clears, no retry spiral
4 Accounting burst Write path holds Zero lost records
5 Node failure under load Local resilience No session loss
6 Site failure Geo-redundancy Surviving site carries full load
7 Slow or failed backend Graceful degradation Defined fallback, no collapse
8 Malformed and rogue traffic Input hardening Dropped and logged, service unaffected
9 Rolling upgrade under load Zero-downtime change No visible dip at the NAS

Each description below covers what you do, what you watch and what passing looks like.

1. Baseline throughput ramp to saturation

Increase traffic in steps from zero, holding each step for a few minutes, and keep going until latency climbs sharply or errors start to appear. At every step, note p99 latency, error rate and CPU per node. The saturation point should sit comfortably above your design peak multiplied by your storm multiplier. Past that point the platform should slow down or start rejecting requests in a controlled way. It should not fall over.

2. Sustained peak soak (24–72 hours)

This one is dull to run, but don’t skip it. Hold design-peak traffic for at least 24 hours, ideally 72, and watch the trend lines for memory, disk use, database size, log rotation and latency. After warm-up, all of them should be flat. Slow memory leaks show up here, and so does the table that grows quietly until a nightly job runs out of room.

3. Re-authentication storm (power-restoration simulation)

Simulate a power cut followed by restoration, when thousands of customer premises equipment (CPE) units, optical network terminals (ONTs) and other devices try to authenticate in the same minute. Send your storm multiplier at once, with NAS retransmission timers configured as they are in production. Queue depth and the retransmission rate tell you whether the platform is coping. Passing means the backlog drains within your recovery target and retries do not multiply the load into a spiral.

4. Accounting burst and buffering

The classic trigger is a BNG reboot, which produces a mass of session Starts or Stops while authentication carries on at peak. Watch accounting response time, database write latency and the billing feed. The pass criterion is strict: zero lost and zero duplicated records, proven by reconciling what was sent against what was stored. Eyeballing a dashboard isn’t enough.

5. Node failure under load

AAA failover testing starts here: kill an AAA node at peak, without a graceful shutdown. Watch session continuity, retransmissions at the NAS and time to full recovery. For Diameter peers, also measure how long watchdog detection takes (RFC 3539). Passing means no established session is lost and client retries succeed on the surviving nodes. The design side of this test is covered in how AAA failover should be tested.

6. Site failure and geo-failover

Cut off an entire site at the network level during peak. Then ask three questions. Does the surviving site carry the full load within your recovery time objective (RTO)? How do its latency and error rate hold up? Does split-brain protection do its job? Our high availability guide explains why both sites should carry live traffic every day.

7. Slow or unavailable backend (LDAP, HSS, SQL)

First add latency to one backend, then take it away completely. The thing to watch is whether the AAA server runs out of threads or connections while it waits, and whether that drags down traffic that never needed that backend in the first place. A good platform follows its designed fallback, whether that is serving from cache, rejecting or queueing, and keeps unrelated traffic moving. See AAA dependencies on LDAP and HSS for the integration side.

8. Malformed, replayed and rogue traffic

Mix malformed packets, replayed requests, wrong shared secrets and traffic from unknown NAS clients into a normal peak load. Drop counts, log entries and legitimate-traffic latency are the signals. Passing means bad traffic is dropped and logged without hurting real subscribers. For detection beyond the test lab, see detecting anomalies in authentication traffic.

9. Rolling upgrade under load

With peak traffic running, upgrade nodes one at a time. Watch error rate and latency at the NAS throughout. Passing means no visible dip. Five-nines availability allows only about five minutes of downtime a year, so a platform that needs a maintenance window for every upgrade will struggle to meet it.

Acceptance Criteria: Defining Carrier-Grade in Numbers

TPS testing only means something with written pass rules. Put thresholds into the proof-of-concept (PoC) scorecard before testing begins, so nobody negotiates them after the results arrive. The thresholds below are starting points. Adjust each one to your service-level commitments and your traffic profile.

Metric How to measure Starting-point threshold
Sustained throughput TPS held at the full traffic mix Design peak, held for the full soak (storm load is judged in scenario 3)
Authentication latency p50, p95 and p99 at the client p99 inside the agreed budget at design peak
Latency drift p99 trend across the soak Flat after warm-up
System errors Timeouts and errors not caused by bad credentials Below 0.1% at design peak
Accounting integrity Records sent vs. records stored Zero lost, zero duplicated
Session loss on failover Established sessions dropped per event Zero
Recovery time Failure to baseline latency Inside the agreed RTO
Saturation behavior Response beyond the ceiling Predictable slowdown or reject, no crash

Measure latency as percentiles: p99 is the response time that 99% of requests beat. An average can look healthy while one request in a hundred waits long enough for the NAS to retransmit. The Google SRE book explains why tail latency is the number users feel. Alepo AAA Server is engineered for 99.999% availability and rated at 36,000+ TPS (AAA Server datasheet), and we expect buyers to score it against criteria like these.

Questions to ask every vendor

Before a PoC begins, we’d put these to any vendor, ourselves included:

  • Will the test run against our traffic profile or yours?
  • Which authentication methods and accounting ratio produced your published TPS figure?
  • Who operates the load generators, and can we watch their CPU during the run?
  • Will you agree to these acceptance criteria in writing before the first test?

A vendor that answers all four without hesitation is ready to be tested properly.

RADIUS and Diameter Load Testing Tools Compared

Choose a tool only after the traffic profile is written. These are the realistic options, from lightweight utilities to full subscriber emulators:

Tool Type Protocol coverage Best fit Limits to plan for
radclient Open source (FreeRADIUS) RADIUS auth, accounting, CoA, Disconnect Functional checks, simple load Not built for full EAP-TLS or EAP-SIM conversations
radperf Network RADIUS utility RADIUS auth and accounting at set rates Throughput benchmarking Check method support for your mix
eapol_test Open source (hostap) Full EAP over RADIUS EAP-TLS, PEAP, EAP-SIM/AKA realism One supplicant per process
Seagull Open source (GPL) Diameter, RADIUS subset Basic Diameter load Last release 1.8.2; check dictionaries for your interfaces
Landslide Commercial Mobile core, Diameter, IMS, Wi-Fi Large-scale subscriber emulation Cost, lab setup time
Keysight IxLoad Commercial Application and subscriber traffic Converged lab environments Confirm RADIUS and Diameter coverage
Custom harness In-house Whatever you build Exact traffic profile Build and maintenance effort

Open-source: radclient, radperf, eapol_test, Seagull

For sending crafted packets and checking replies, radclient is the quickest RADIUS test tool. The radperf utility sends authentication and accounting traffic at set rates and reports offered load against accepted load. For EAP load testing, eapol_test runs complete EAP conversations, so scaling it means running many instances in parallel across several load-generator hosts. Seagull is a long-standing open-source traffic generator that covers Diameter and a subset of RADIUS. Check that its dictionaries cover SWm and S6b, the 3rd Generation Partnership Project (3GPP) interfaces an AAA server uses when subscribers attach over untrusted non-3GPP access, such as Wi-Fi calling through an ePDG (3GPP TS 29.273).

Commercial traffic generators: Spirent Landslide, Keysight IxLoad

Commercial generators emulate very large subscriber populations and full call flows across many interfaces at once. For example, emulates mobile core, Diameter, IMS and Wi-Fi nodes, including AAA and policy functions. They suit mobile-core and Wi-Fi offload testing where EAP-AKA and Diameter must run together. Confirm interface coverage against your profile before you commit lab time.

Executing the Test: A Step-by-Step Plan

  • Set up an isolated lab that mirrors production topology: NAS emulators, load balancers, AAA nodes, both sites and backends with realistic latency.
  • Instrument everything before you send traffic: client-side latency, server CPU and memory, backend latency and queue depth, exported to Prometheus and Grafana or your own stack.
  • Warm up at low load until caches and connection pools stabilize, then discard warm-up data.
  • Ramp in steps to design peak, then beyond it to saturation (scenario 1).
  • Run the remaining scenarios in order, resetting to a clean baseline between each.
  • Report against the scorecard, with the traffic profile, tool versions and configuration attached to every result.

Watch the load generators as closely as the AAA server. If a generator’s CPU saturates first, you are measuring the tool. For platforms on Kubernetes, the same metrics feed autoscaling, covered in autoscaling AAA on custom metrics.

Seven Mistakes That Produce Misleading AAA Benchmarks

  • Testing PAP only. Most production mixes include EAP methods that cost far more per authentication.
  • Ignoring accounting. The write path is often the first thing to fail.
  • Stubbing the backend. A zero-latency backend hides the real bottleneck.
  • Too few NAS clients. One client at high rate does not exercise load balancing or connection handling.
  • Saturating the client. A busy load generator caps the result below the server’s limit.
  • Miscounting retransmissions. Retries inflate TPS while the user still waits; RFC 5080 covers duplicate handling.
  • Running too short. A one-hour test will never find a leak that takes 48 hours to bite.

Stress Testing During Migration and Cutover

If you’re replacing an AAA platform, stress testing belongs inside the migration plan, not after it. Test the new platform against your recorded production profile before cutover. Then run it in parallel with the legacy system, both carrying live traffic, and compare the two before moving the remaining load across in measured steps. Each step needs its own rollback point.

Alepo AAA Server has more than ten carrier deployments worldwide. We think every vendor, ourselves included, should put its parallel-run and rollback plan in writing before the PoC begins.

Conclusion

AAA server stress testing is how you turn a vendor’s TPS claim into evidence about your own network. Build the traffic profile first, run all nine scenarios, judge them against written acceptance criteria and only then decide which tools and which platform pass.

Prefer to see the platform first? Walk through Alepo AAA Server with our team: authentication flows, accounting, failover behavior and the admin portal. Book a demo

Frequently Asked Questions

Q1. What is the difference between load testing and stress testing an AAA server?

Load testing confirms the platform meets its latency and error targets at planned peak volume. Stress testing pushes past that peak to find the saturation point and prove the platform fails gracefully. You need both. A soak test holds peak load for 24 to 72 hours to catch slow leaks.

Q2. How many TPS does a carrier AAA server need?

Start from your measured sustained peak, multiply by your storm multiplier, then add growth and headroom. Count accounting and CoA messages as well as authentications, because accounting often dominates. No market average will give you that number.

Q3. How do you load test EAP-TLS or EAP-SIM authentication?

Use eapol_test in parallel instances or a commercial generator that runs full EAP conversations. EAP-TLS needs test certificates for each synthetic subscriber. EAP-SIM and EAP-AKA need test SIM credentials and an HSS or simulator that returns authentication vectors.

Q4. What tools generate RADIUS test traffic?

For functional checks and simple load, use radclient. The radperf tool is built for RADIUS throughput benchmarking. eapol_test runs full EAP conversations. Commercial generators such as Landslide and Keysight IxLoad emulate large subscriber populations across several protocols.

Q5. What tools generate Diameter test traffic?

Seagull is an open-source option for basic Diameter load. Commercial generators such as Landslide cover larger mobile-core scenarios. Check each tool’s coverage of the 3GPP interfaces you use, such as S6b, SWm and SWx.

Q6. How do you test AAA failover under load?

Kill a node or isolate a site while peak traffic runs, without a graceful shutdown. Measure dropped sessions, NAS retransmissions and time back to baseline latency. Passing means zero session and accounting-record loss within your recovery target.

Q7. How long should an AAA soak test run?

At least 24 hours, and ideally 72. That window exposes memory growth, log rotation problems and database bloat that a short test cannot reveal.

Q8. Should stress testing be repeated after go-live?

Yes. Repeat it before launching a new access type, before major upgrades and at least once a year. Align each run with your capacity planning cycle so test results feed sizing decisions.

Want to see how this applies to your business? Let’s talk.

Share the Post:

Latest Posts

Receive the latest news

Subscribe To Our Newsletter

Subscribe to our Newsletter

Receive the latest news

Subscribe To Our Newsletter