What an escalation policy actually is
An escalation policy is a sequence of “if not acknowledged, then…” rules. It governs what happens after an alert fires:
- Page the primary on-call. Wait N minutes for acknowledgment.
- If not acknowledged, page the secondary on-call. Wait N more minutes.
- If still not acknowledged, page someone else (a manager, the whole team, etc.).
The principle is that someone gets the incident, but you don’t blast everyone at once. The cost of waking the wrong person grows linearly with team size, and most incidents don’t need more than one human.
The structure looks simple, but every parameter is a tradeoff between detection speed, response speed, and on-call fatigue.
The numbers that matter
There are four numbers in any escalation policy. Tuning them is most of the work.
1. Time-to-first-alert (TTFA) — How long after the actual outage before the first page goes out. This is your detection latency. Affected by check frequency, re-check policy, and any deliberate alert suppression for flapping. For most teams 30 seconds to 2 minutes is the right zone.
2. Time-to-acknowledge (TTA) — How long the on-call has to acknowledge before the next person gets paged. Tradeoff between waking the secondary unnecessarily (too short) and letting an incident burn while the primary is in the shower (too long). Five minutes is the default most teams settle on.
3. Escalation depth — How many levels of escalation before everyone is paged. Two or three is normal. More than three usually means your policy is too defensive.
4. Re-escalation interval — If nobody acknowledges at level 1, how long before level 2 fires. Same tradeoffs as TTA. Usually identical.
The numbers that work for a 24/7 production site are different from the numbers that work for an internal tool. Calibrate them to the cost of being slow and the cost of waking the wrong person.
A starting template for small teams
For a team of 2–8 engineers running production:
Severity: Critical (site down, payments broken, data loss)
Level 1: Primary on-call. SMS + push + Slack channel.
Wait: 5 minutes.
Level 2: Primary on-call (re-page) + Secondary on-call. SMS + push.
Wait: 5 minutes.
Level 3: Page entire engineering channel. Slack @here.
Severity: High (degraded performance, partial outage)
Level 1: Primary on-call. Push + Slack channel.
Wait: 10 minutes.
Level 2: Secondary on-call. Push.
Wait: 15 minutes.
Level 3: Email manager. No further escalation.
Severity: Warning (single-region flake, deferred maintenance)
Level 1: Slack channel post only. No page.
No escalation. Review in the morning.
That’s the whole thing. Three severity levels, three escalation levels per severity, fewer for warnings. Most actual incidents are handled at Level 1.
What “severity” should mean
Severity isn’t about how much you personally care. It’s about how many resources should be mobilized.
- Critical = wake the on-call. Real human pain, real revenue impact, real data risk.
- High = ping the on-call during working hours, defer to morning otherwise unless it’s still firing then.
- Warning = put it in a channel for review. No one needs to drop what they’re doing.
The discipline is what gets classified as critical. The mistake teams make is being generous — calling things critical when they’re really high, calling things high when they’re really warnings. Everything gets escalated, everything wakes someone, everyone learns to ignore the alerts. Inflate severity at your own cost.
A useful test: for this alert, is it acceptable to wake someone at 3 AM? If the answer is yes, it’s critical. If you’d feel bad waking someone for it, it’s not.
What to monitor with what severity
Some rules of thumb:
Critical:
- The homepage / main entry point returning 5xx errors for multiple consecutive checks
- The checkout / payments endpoint returning 5xx for multiple consecutive checks
- The auth endpoint returning 5xx for multiple consecutive checks
- Database completely unreachable
- Multi-region monitor confirming a full outage
High:
- Response times >10x normal for sustained period
- Single-region flakes that persist
- API endpoints (non-revenue) returning errors
- Background job queue depth exploding
Warning:
- Response times degraded but functional
- SSL cert expiring in <30 days (escalates to high at <7)
- Disk space approaching 80%
- Error rate elevated but under threshold
If everything is critical, nothing is. If nothing is critical, you’ll find out about outages from customers.
The 5-minute acknowledgment problem
The default acknowledgment timeout (5 minutes) is too short for some incidents and too long for others. Two patterns help:
Stagger by time-of-day. During business hours, your acknowledgment timeout can be longer because someone will likely see the Slack message even if they didn’t page-acknowledge. Overnight, shorter — the primary is sleeping, and if they didn’t wake up in 3 minutes they’re probably not waking up at all.
Stagger by severity. Critical = 3 minute timeout. High = 10 minutes. Warning = no escalation. Make the system match the urgency.
On-call rotations
If you’re a team of 2–4, on-call is whoever’s not on vacation. If you’re 5–10, you should rotate by week. If you’re larger, formalize it.
Practical rotation rules:
- Weekly rotations are the standard. Daily is too disruptive; monthly is too long.
- Rotate Mondays, not Fridays. Starting your week on-call gives you time to ramp up. Ending your week on-call means you carry it through the weekend.
- Compensate. Either in cash, in time off, or in not-being-paged-during-other-things. On-call has a real personal cost.
- Have a clear handoff. The outgoing on-call should pass off open incidents, anything they’ve been watching, and any operational context. Five minutes of conversation prevents 5 hours of confusion.
For very small teams (1–2 engineers), formal on-call isn’t worth the overhead. You’re all on-call all the time. The escalation policy still matters, but it’s mostly about what wakes you up vs. what waits until morning.
Alert hygiene that actually matters
The single biggest improvement most teams can make is dedup at the source. One alert per incident, not one per failed check. If a monitor fails 10 times in 10 minutes, that’s still one alert.
Other high-ROI hygiene:
- Auto-resolve when the underlying condition clears. Don’t make humans manually acknowledge resolved alerts.
- Suppress during planned maintenance. Most monitoring tools have a “paused” or “maintenance window” mode. Use it.
- Group related alerts. Five different monitors firing for the same root cause should appear as one incident, not five.
- Don’t page on warning-level alerts. Slack them. The whole point of severity is to make this distinction.
- Review alerts that fired but didn’t need action. Once a week, look at the alert log. Anything that fired without being actionable is a candidate for tuning or deletion.
The goal isn’t to never get paged. The goal is for every page to be worth getting paged for.
What MyUptimeBot does, plainly
MyUptimeBot handles the detection side: multi-region rechecks before alerting, configurable severity per monitor, push/email/SMS channels with per-monitor routing.
For full escalation policy logic — multi-person rotations, ack-and-escalate flows, complex routing rules — most teams pair our detection with a dedicated incident tool like PagerDuty or Opsgenie via webhook. We’re the source of the alert; they’re the routing engine.
For small teams (1–3 people), pointing MyUptimeBot directly at SMS/push for critical monitors is often enough. You don’t need PagerDuty until you have a rotation to manage. When you do, hooking the two together is a 5-minute setup.
The principle, stated plainly
Every parameter in an escalation policy is a tradeoff between detecting incidents fast and not waking people unnecessarily. The right policy reflects what your business actually loses when each goes wrong.
Start conservatively (slightly longer waits, slightly fewer escalations), and tighten over time only if you find real incidents going unnoticed. The opposite path — starting aggressive and loosening after burnout — leaves a trail of muted phones and missed alerts.