Summary
- “24/7 support” only guarantees that a vendor will accept your ticket at any hour. It says nothing about response speed, escalation, expertise, or restoration.
- Demand response time tiers by severity, with two targets per tier: time to a qualified human (automated acknowledgments do not count) and time to workaround or restoration. You define the severity levels and declare the initial severity.
- Demand a written escalation path with a time bound on every hop, automatic escalation for an unresolved Severity 1, a named incident manager, and a documented shift-handover process.
- Demand specialist access: Severity 1 tickets bypass first-line triage and land with engineers who have run RADIUS or Diameter in a live operator network, ideally a named engineer or pod that knows your deployment.
- Before signing, ask for twelve months of response and restoration data by severity, two reference customers who have lived through a Severity 1, and an anonymized root-cause analysis. Records, not adjectives.
For instance, suppose it’s 2:40 a.m. Authentication success in one region has fallen off a cliff. The on-call engineer has already ruled out the network access server (NAS) and the subscriber store, and the AAA vendor’s portal confirms a ticket has been logged. The person who eventually calls back is a first-line agent working from a script who has never seen a Diameter Capabilities-Exchange failure. The clock is running on your uptime SLA, and nothing in the vendor’s “24/7 support” clause is being breached.
If you are negotiating a service-level agreement (SLA) for authentication, authorization, and accounting (AAA) infrastructure, here is the short version:
A protective AAA support SLA specifies response time tiers by severity, a defined escalation path, and access to engineers who understand carrier-grade AAA specifically. “24/7 support” on its own guarantees only that someone answers at any hour, not how fast they respond, how far a problem can escalate, or whether the person on the line has ever run RADIUS or Diameter in production.
Our guide to designing a high-availability AAA server covers the architecture side of reliability. This article covers the human side, where the uptime number gets defended or lost.
Response time tiers by severity
A single blanket promise (“we respond within four hours”) is the most common weakness in a vendor support SLA, and it fails in both directions. Four hours is far too slow for a full authentication outage and pointlessly aggressive for a cosmetic reporting bug. Insist instead on severity tiers, each with its own guaranteed response time and its own guaranteed time to workaround or restoration.
Define the tiers by subscriber impact rather than by how the vendor’s ticketing tool labels them. A workable structure for AAA: Severity 1 is subscribers unable to authenticate or sessions being torn down at scale; Severity 2 is a partial impairment such as one access type, one realm, or accounting records failing to write; Severity 3 is a degraded non-critical function with a workaround in place; Severity 4 is a question or change request. Write those definitions into the contract yourself. Vendor-written definitions tend to set a high bar for Severity 1.
Then attach two numbers to each tier. Response time is when a qualified human engages with the ticket, and an automated acknowledgment email does not count. Workaround or restoration time is when subscribers stop feeling the problem. Vendors commit readily to the first; the second is what protects your network. For a full outage on carrier-grade AAA, a defensible ask is a response measured in minutes and a workaround target measured in hours, with you declaring the initial severity. Our AAA server capacity planning guide explains why the partial-impairment tier matters more than it looks: an overloaded node rarely reads as an outage until it becomes one.
Escalation path guarantees
Response time tells you when the conversation starts. Escalation terms tell you what happens when the first person cannot fix it, and this is where most support contracts go quiet. Ask the vendor to show you, in writing, the path from first-line support to a senior engineer to the team that owns the code, with a time bound on each hop. If the answer is “our support manager decides case by case,” you do not have an escalation path. You have a hope.
A contractual escalation path for AAA names three things. The trigger: an unresolved Severity 1 escalates automatically after a fixed interval, without the customer having to ask. The destination: each level is a defined role with defined authority, up to the engineering team responsible for the product. The management track: a named incident manager who owns communication during a major incident, so the engineer fixing the problem is not also writing the status emails.
Ask specifically what happens to an open Severity 1 at shift change in a follow-the-sun model. Handover is where context gets dropped and escalations quietly reset; in our experience it is the single most common point where a two-hour incident becomes a six-hour one. A vendor with carrier experience will have a documented answer. A vendor selling enterprise IT support with a telecom logo on the slide often will not.
Access to engineers who understand carrier-grade AAA
Most “24/7” operations are staffed for coverage, not competence. Coverage means a person is awake. Competence means that person can read a RADIUS Access-Request carrying vendor-specific attributes, recognize a Diameter Device-Watchdog timeout, or spot an Extensible Authentication Protocol (EAP) negotiation failing on one handset firmware. Those are not first-line skills anywhere, and no hours-of-operation language changes that.
What you can demand is the composition of the team behind the phone number. Ask how many support staff have deployed AAA in a live operator network, how many have worked a Severity 1 on RADIUS or Diameter specifically, and whether Severity 1 tickets bypass first-line triage and land with a specialist. A named engineer or small pod that knows your deployment is worth more than any headcount figure. The alternative is explaining your architecture from scratch, at night, to whoever picked up.
Vendor structure matters here more than procurement usually credits. A vendor that builds, deploys, and supports its own AAA platform can put the people who wrote the code on your incident bridge; a reseller, or a vendor with outsourced support, structurally cannot. The Alepo AAA Server is engineered and supported by the same organization, with deployments across Tier-1 fiber and multi-million-subscriber mobile and broadband networks. In practice that means the engineer on a Severity 1 has often seen the failure mode on a live network before. Put the same question to every vendor you are evaluating.
What “24/7” actually guarantees (and doesn’t)
Read the clause literally. “24/7 support” commits the vendor to accepting your ticket at any hour. That is all. It does not commit them to responding within a defined time, to escalating, to staffing the night shift with anyone who knows your product, or to restoring service. Each of those has to be written in separately, and if it is not, the vendor is not in breach when it does not happen.
The gap between the phrase and the contract is rarely deception. It is marketing shorthand that buyers fill in with their own assumptions, and the assumptions are what surface during the outage. The table below is the checklist to carry into negotiation.
Support SLA terms to demand from an AAA vendor
| SLA term | Why it matters for AAA | What to demand |
| Response time tiers by severity | A full authentication outage and a reporting bug cannot share one response target; subscriber impact is immediate and total at Severity 1 | Four severity levels defined by subscriber impact; response and workaround targets for each; customer declares initial severity |
| Escalation path | First-line support cannot diagnose protocol-level or software faults; unbounded escalation lets a Severity 1 stall for hours | Automatic time-bound escalation to senior and engineering levels; named incident manager; documented shift-handover process |
| Specialist engineer access | AAA failures are protocol and integration failures; generalists must relearn your architecture on every call | Severity 1 bypasses first-line triage; named engineer or pod familiar with your deployment; vendor discloses support team AAA experience |
Questions that reveal real support quality
Contract terms tell you what a vendor has promised. Five questions during evaluation tell you what they deliver, and the useful ones all ask for evidence.
Ask for response and restoration data by severity for the last twelve months, aggregated across the vendor’s customer base; a vendor that meets its SLA will produce this readily. Ask to speak with two reference customers about a real Severity 1, and ask them how many people they had to talk to before it was fixed. Ask what happens to the ticket at shift change. Ask who, by name and role, would join your Severity 1 bridge in the first hour. And ask for an anonymized root-cause analysis (RCA) from a past major incident, because the quality of a vendor’s RCA is the most honest signal of its engineering culture you will get before signing.
Demanding a support SLA that actually protects you
Uptime is architecture plus people. Active-active redundancy decides whether a node failure is a non-event; support terms decide whether a software defect at 2:40 a.m. is a 20-minute incident or a 6-hour one. An SLA that specifies 99.999% availability and then says “24/7 support” underneath is only half written, and the unwritten half is the part you will negotiate hardest about after the outage.
So write it before. Severity tiers defined by subscriber impact, with response and workaround targets you set. Automatic, time-bound escalation to people with engineering authority. Specialists on a Severity 1 from the first minute. And records, from every vendor on the shortlist.
Alepo’s AAA server is designed for 99.999% availability with active-active session replication across live nodes, measured at 36,000+ sustained transactions per second in containerized deployment (Alepo production figures, not industry averages). The support model behind it follows the principle this article argues for: the engineers who deploy the platform are the engineers who answer when it matters.
Bring your draft support SLA to the conversation. Book a demo and ask us to walk through, clause by clause, how we would meet each term. Or start with the AAA Server datasheet to see the architecture the support model sits on.
Frequently asked questions
Q1. What should I demand in an AAA vendor’s support SLA beyond 24/7 availability?
Response time tiers tied to severity, with both response and workaround targets; an automatic, time-bound escalation path to engineering-level support; and access to engineers with hands-on carrier-grade AAA experience on every Severity 1.
Q2. What response time tiers should a support SLA include?
A separate guaranteed response time for each severity level, not one blanket promise. A full authentication outage should carry a response target in minutes and a workaround target in hours; a minor issue with a workaround can reasonably wait days.
Q3. What escalation path should an AAA vendor’s SLA guarantee?
A written path from first-line support through senior engineers to the team that owns the code, with a time bound on each hop, automatic escalation for unresolved Severity 1 tickets, and a named incident manager who owns communication.
Q4. What does “24/7 support” actually mean in practice?
Only that the support function will accept a ticket at any hour. It does not by itself guarantee response speed, escalation, specialist expertise, or restoration time; each must be written into the contract separately.
Q5. Should I ask for access to named support engineers?
For carrier-grade infrastructure, yes. AAA failures are protocol and integration failures that generalist first-line support rarely diagnoses, so a named engineer or small pod familiar with your deployment removes the costliest part of an incident: re-explaining your architecture under pressure.
Q6. How are support severity levels typically defined?
By subscriber impact. Severity 1 is subscribers unable to authenticate or sessions dropping at scale; Severity 2 is a partial impairment such as one access type or accounting failing; Severity 3 is a degraded function with a workaround; Severity 4 is a question or change request. Insist that the customer declares initial severity.
Q7. What questions reveal whether a vendor’s support is actually carrier-grade?
Ask for twelve months of response and restoration data by severity, two reference customers who have lived through a Severity 1, the shift-handover process for open incidents, who would join your bridge in the first hour, and an anonymized root-cause analysis from a past major incident.

