Multi-Site AAA Server Deployment Best Practices

A multi-site AAA server deployment only survives a site loss if five things are executed, not assumed: subscriber and session state replicated across sites in real time and asynchronously to the auth path; every request routed to the nearest healthy site with NAS (network access server) timers set from measured latency; failover drills that isolate a whole site, not just a node; data classified for residency before it crosses a border; and an inter-site WAN (wide area network) sized, soak-tested, and physically diverse. The consideration table below turns those five into a checklist.

A single-site AAA (authentication, authorization, and accounting) cluster can survive a dead server. It cannot survive a dead building. Once an operator’s footprint spans regions, the question stops being “how do we make the AAA highly available?” and becomes “how do we make the second site behave like a real peer and not a box on a diagram?” That is a deployment problem more than a design problem.

Scope note: this article covers execution, meaning how to get a multi-site AAA server deployment synchronized, routed, tested, and compliant. The architectural choices behind it (site count, topology, active-active versus active-standby) are covered in the dedicated multi-region AAA server design guide. Read that one to decide what to build. Read this one to build it properly.

Multi-site AAA server deployment distributes authentication, authorization, and accounting infrastructure across two or more geographic locations, with synchronized subscriber data, latency-aware traffic routing, and regularly tested cross-site failover, so that losing an entire site reduces capacity instead of interrupting service.

Five practices separate a deployment that survives a site loss from one that only claims to.

1. Data synchronization across sites

The most common mistake in a multi-site AAA deployment is treating the second site as a backup of the first. Backups are copies taken at intervals. A peer site needs current state (subscriber profiles, active sessions, accounting counters, quota balances), because the moment traffic shifts, a RADIUS or Diameter request arriving at Site B must be answered from data written at Site A seconds earlier.

That means real-time, bidirectional replication of both the subscriber store and the session store, not a nightly database dump. The subscriber store is the easier half: it changes slowly and tolerates seconds of lag. Session and accounting state changes on every authentication and every interim update, and it has to cross the inter-site round-trip over the WAN without blocking the authentication path.

Two execution details decide whether this works. First, replication must be asynchronous with respect to the auth response. The NAS gets its Access-Accept from the local site immediately; the write reaches the remote site a moment later. Synchronous cross-site writes put WAN latency inside every authentication, and the NAS timeout budget will not forgive that at scale. Second, agree on the conflict-resolution rule before go-live: when the same session is touched at both sites during a partition, which write wins? Whatever the rule, it must be deterministic, documented, and tested. Alepo AAA Server is built on real-time database replication with stateless session storage backed by database persistence, and is engineered for 99.999% availability.

2. Latency-aware traffic routing

Picture a national fiber operator with AAA sites in the east and west of the country. Every BNG (broadband network gateway) in the west points to the western site as primary and the eastern site as secondary. One day the western site loses power, every western BNG fails over, and the eastern site absorbs the whole country. It holds, but authentication latency climbs for half the subscriber base and every interim accounting update from the west now crosses the WAN. Nobody planned for that traffic pattern, so nobody tested it.

Latency-aware routing means steering each authentication request to the nearest healthy site as a deliberate configuration decision, not as a side effect of NAS failover order. There are three layers to get right. At the NAS, configure primary and secondary AAA targets by proximity, with timeout and retry values matched to the measured latency of each; a secondary configured with a primary’s aggressive timeout will flap under load. In front of each site, place a load balancer or Diameter Routing Agent that health-checks the AAA nodes and stops advertising a site that is degraded, not only one that is down. Between the sites, define the overflow rule: when a site is lost, does its traffic go entirely to one neighbor or split across several?

Write the routing plan down as a table of source region → primary site → secondary site → expected latency, with the latency column taken from measurement, not from the WAN provider’s brochure. That table is also the input to capacity planning, because a site has to be sized for the traffic it will inherit, not the traffic it serves on a normal day. Our AAA server capacity planning guide covers the sizing arithmetic.

3. Testing site failover, not just assuming it

Most operators have run a node failover drill. Far fewer have run a site failover drill, and the two test different things. Killing a node proves the local cluster rebalances. Isolating a site proves that replication was current, that NAS failover timers actually fire, that the surviving site holds the inherited load, and that accounting written on the lost site can be reconciled afterward. None of that can be inferred from a node test.

A site drill follows the same staged discipline as any production test: shadow first, then a limited blast radius, then a full cut, with a written rollback trigger and one named engineer holding abort authority. The staging is covered in how to run an AAA failover drill without dropping live subscribers. What changes at site level is what you measure. Track replication lag at the moment of isolation (how many sessions were at risk), time-to-failover at the NAS, authentication success at the surviving site through the transition, accounting records lost or duplicated, and the behavior when the lost site comes back. Rejoining a site after an hour offline means resynchronizing an hour of state, and a badly ordered rejoin can overwrite newer data with older.

Rotate the drill so each site is isolated at least once a year, and retest after any change to the WAN, replication configuration, or NAS fleet. A drill that only ever kills Site B proves nothing about losing Site A.

4. Data residency and regulatory considerations

Replication does not respect borders unless you tell it to. When the sites in a multi-site AAA deployment sit in different jurisdictions, or one operator serves several countries from shared infrastructure, the subscriber data being replicated may fall under residency, retention, or lawful-intercept rules that differ per site. The AAA team is rarely the one who knows which.

The execution practice is to classify before you replicate. Separate what every site needs in order to authenticate (credentials, service profiles, active session state) from what can stay local (long-term accounting archives, audit logs, identifiers beyond what authentication requires). Map each class against each site’s obligations, with the operator’s legal or compliance function deciding what may cross which boundary. Where a class cannot leave a jurisdiction, the options are partitioned cross-region replication (the site pair inside the jurisdiction replicates fully; the cross-border link carries only the permitted subset) or a per-country deployment that shares operations tooling but not subscriber data. Encrypting replication traffic in transit is a baseline on every link. It is not a substitute for the classification. Residency rules vary and change, which is why this conversation belongs in the deployment plan and not in the post-incident review.

5. Network requirements between sites

The inter-site WAN is the one component of a multi-datacenter AAA deployment the AAA team does not own and cannot fix, so its requirements need to be specified early and verified before go-live. Four numbers matter. Bandwidth must carry peak replication volume (a function of authentication rate, interim accounting frequency, and record size) plus the burst that follows a site rejoin, when a backlog of state catches up. Latency sets the floor for replication lag, and therefore for the sessions at risk during a partition; it also decides whether a site can serve as secondary for a distant region within the NAS timeout budget. Jitter and loss matter more than raw latency for replication that retransmits; a link with low average latency but frequent loss shows erratic lag. Path diversity, meaning two physically separate routes, is what stops a single fiber cut from turning a multi-site deployment into two isolated single sites.

Before the first subscriber migrates, run replication over the production WAN for a full weekly cycle and graph lag against link utilization. If lag climbs at peak, the link is undersized or the replication is chattier than expected. Either is far cheaper to find in soak than in service.

Multi-site AAA deployment considerations at a glance

Consideration Recommendation
Data synchronization Real-time, asynchronous, bidirectional replication of subscriber and session state; documented conflict-resolution rule
Traffic routing Route to the nearest healthy site; NAS timers matched to measured latency; overflow plan written down
Site failover Staged drills that isolate a whole site, measuring replication lag, NAS failover time, and rejoin behavior
Data residency Classify data before replicating; compliance function decides what crosses borders; encrypt every link
Inter-site network Bandwidth for peak plus rejoin burst, measured latency, low loss, and physically diverse paths

Planning your multi-site AAA server deployment

If the multi-site AAA server deployment is still on paper, sequence it so each practice is proven before the next depends on it. Stand up the sites and soak-test replication over the real WAN. Build the routing table from measured latency and configure the NAS fleet against it. Run the first site-isolation drill against a small segment before any large migration wave lands. Confirm the residency classification with compliance before cross-border replication carries production data. Then migrate in waves, using the AAA server installation checklist as the per-wave gate.

The platform underneath decides how much of this is engineering and how much is fighting the product. Alepo AAA Server is built for multi-site operation. It runs N+1 and N+N redundancy with real-time database replication and stateless session storage backed by database persistence, so a surviving site answers from current subscriber and session state rather than a stale copy. It terminates RADIUS, Diameter, and TACACS+ on one platform. And it deploys as a containerized, clustered system with load balancing, so a site can be scaled horizontally to hold the load it inherits from a neighbor. It is engineered for 99.999% availability and runs in Tier-1 and Tier-2 carrier networks across the Middle East, Europe, Latin America, Africa, and Asia, in fixed, mobile, and converged deployments.

Watch a site fail over before you buy one. Book a demo and ask us to isolate a site while you watch the authentication graph at the survivor, or start with the Alepo AAA Server solution page to put the architecture in front of your network planning team.

Frequently asked questions

Q1. What are best practices for multi-site AAA deployment?

Five practices cover most of the risk: replicate subscriber and session state in real time across sites, route each request to the nearest healthy site with NAS timers matched to measured latency, run staged drills that isolate an entire site, classify data for residency before it crosses a border, and specify and soak-test the inter-site WAN before go-live.

Q2. How does data synchronize across AAA sites?

Through real-time, bidirectional replication of the subscriber and session stores, designed so the authentication response is served locally and the cross-site write happens asynchronously. The replication has to tolerate inter-site latency without putting it inside the auth path, and it needs a deterministic conflict-resolution rule for writes that collide during a partition.

Q3. How do I test multi-site failover?

With a scheduled drill that isolates a whole site and shifts real traffic to the survivor, staged from shadowing to a limited segment to a full cut, with a written rollback trigger. Measure replication lag at isolation, NAS failover time, authentication success at the surviving site, accounting records lost or duplicated, and rejoin behavior. Reviewing the runbook is not a test.

Q4. What is latency-aware routing in multi-site AAA deployment?

Steering each subscriber’s authentication request to the nearest site that is healthy, using proximity-based NAS primary/secondary configuration, health-checking load balancers or Diameter Routing Agents in front of each site, and a written overflow plan for where a lost site’s traffic goes.

Q5. How many sites should a geo-redundant AAA deployment use?

Commonly two or three. Two gives site-level redundancy with the simplest replication topology; three lets a surviving pair keep redundancy after one site is lost, at the cost of more replication paths to size and test. Beyond three, synchronization complexity usually outweighs the availability gain.

Q6. How does multi-site deployment differ from single-site HA?

Single-site high availability protects against component failure (a server, a process, a rack) with nodes that share a low-latency network and often a database. Multi-site deployment protects against losing the entire site, which introduces WAN latency into replication, NAS failover across distance, data-residency questions, and a failover drill that has to isolate a whole location.

Q7. What network requirements exist between AAA sites?

Enough bandwidth for peak replication volume plus the catch-up burst after a site rejoins, latency low enough that replication lag and secondary-site authentication stay within the NAS timeout budget, low packet loss and jitter, and physically diverse paths so a single cut cannot partition the deployment.

Want to see how this applies to your business? Let’s talk.

Share the Post:

Latest Posts

Receive the latest news

Subscribe To Our Newsletter

Subscribe to our Newsletter

Receive the latest news

Subscribe To Our Newsletter