The job of an alert
An outage alert has one job: get your attention fast enough that you can act before the outage compounds. Every notification channel handles that job differently, and the differences matter more than the technology behind them.
Three things determine whether an alert works:
- Latency — how long after the outage starts you actually see the notification
- Reliability — how often the alert successfully reaches you
- Disruption — how hard it is to ignore the alert when you do see it
A perfect alert is fast, reliable, and disruptive. Different channels nail different combinations.
Email is the universal default. Every monitoring service supports it, every operator has it, and it works on any device.
Latency: Usually 30 seconds to 2 minutes from send to inbox. Sometimes longer if your email provider is rate-limiting, in a slow-delivery situation, or if the alert gets routed through spam filtering.
Reliability: High in absolute terms (99%+ of legitimate emails are delivered) but the failure modes are nasty. A misconfigured spam filter or a renamed inbox folder can swallow alerts silently for weeks. You won’t know they’re missing until you check.
Disruption: Low. Most operators get hundreds of emails a day, and outage alerts compete with everything else. Outside of working hours, email goes unread.
Where it shines: Audit trail. Email gives you a permanent record of every alert with full context that you can forward, search, and archive. That’s why it stays in the mix even when other channels are louder.
Where it fails: As your only channel for critical alerts. If your site goes down on a Saturday morning and the alert is buried in your inbox between newsletters, you’re not finding out until Monday.
Push notifications (mobile app)
Push notifications go to a mobile app on your phone over Apple’s APNS or Google’s FCM infrastructure.
Latency: Typically 1–5 seconds from send to your phone, often faster. The fastest channel available to most operators.
Reliability: Generally high, but with two caveats. First, if your phone is off, in airplane mode, or out of signal, the push waits in queue until it comes back online. Second, mobile OS battery optimization can throttle or delay notifications from apps that haven’t been opened recently. The fix is to open the app once a week or whitelist it.
Disruption: Medium-to-high. Push notifications light up your screen, play a sound, and (on most phones) survive Do Not Disturb if you set the app as a priority. Hard to miss when your phone is nearby.
Where it shines: Real-time awareness for individual operators. If you’ve got the monitoring app on your phone and your phone with you, push is the fastest way to know.
Where it fails: Sleep. If your phone is on silent overnight, you’ll wake to a push waiting for you — but that’s not the same as being woken by it. For a sleeping operator, SMS is louder.
SMS / text message
SMS goes through carrier networks directly to your phone number.
Latency: Usually 5–30 seconds, depending on carrier and routing. Slower than push but more reliable through carrier-level disruptions.
Reliability: Very high. SMS predates the modern internet and works on basically any phone, including phones with no data connection. If your carrier has signal, you get the message.
Disruption: High, especially overnight. Most phones make a distinctive sound for SMS regardless of Do Not Disturb settings, and many people have their text alerts louder than their other notifications. SMS is what wakes you up at 3 AM.
Cost: This is the catch. SMS costs the monitoring provider real money — a few cents per message, more for international routing. Most providers either charge per SMS, include a small bundle in paid plans, or reserve SMS for higher-priced tiers.
Where it shines: Wake-up alerts for critical infrastructure. If your e-commerce site going down at 3 AM costs you serious money, SMS is what bridges from “asleep” to “fixing it.”
Where it fails: Bulk alerting. If your monitor flaps and you’d get 20 alerts an hour, SMS becomes both expensive and unbearable. Most setups rate-limit SMS or only send the first alert in a flapping cluster.
Webhook / chat integrations (bonus)
Webhooks let monitoring services post alerts directly into Slack, Discord, PagerDuty, Microsoft Teams, or any custom endpoint that accepts HTTP POST requests.
Latency: A few seconds. The receiving service then handles its own notification — Slack pings the channel, PagerDuty starts an escalation policy, etc.
Reliability: Depends on the receiving service. Slack has had multi-hour outages. PagerDuty is purpose-built for reliability. Build your own webhook and you own the reliability problem.
Disruption: Configurable per channel. Slack channels can be muted, PagerDuty can wake whoever’s on-call.
Where it shines: Team coordination. When more than one person needs to know about an alert, posting it into a shared channel beats blasting individuals.
The combination most operators land on
For most small teams running production sites, the pattern that works is two channels per alert level:
| Alert severity | Channels |
|---|---|
| Informational (“response time degraded”) | Email only — review later, no action needed |
| Warning (“monitor flaky from one region”) | Email + push — see it when you check your phone |
| Outage (“site is down, confirmed multi-region”) | Push + SMS — wake you up if needed |
| Critical / business hours | Push + SMS + Slack/PagerDuty for the team |
The principle is: the more important the alert, the more redundant the channels. A push notification is great until your phone is dead. An SMS is great until you don’t see your phone. Email is great as backup for both. Multiple channels for the same critical alert means failure of any one channel doesn’t lose the message.
A note on alert fatigue
The fastest way to make alerts ineffective is to send too many of them. Operators learn to mute their phones after a few false alarms, and once that happens the next real alert gets ignored too.
Two things help: aggressive de-duplication (one alert per incident, not one per failed check) and noise reduction (re-check from a second region before alerting at all). Both should be the monitoring service’s job, not yours.
If your current setup pages you for every flake, that’s a setup problem. Either the monitor is too sensitive, or the service doesn’t filter false alarms before sending — and either way, you’ll mute it within a week.
What we recommend, plainly
If you’re starting from zero: set up email + push first. That covers most situations and costs nothing extra. Add SMS only for the specific monitors where being asleep through an outage would be genuinely bad — usually your most important production site, not all of them.
If you’ve got a team: add a Slack or Discord webhook so everyone sees the same alert in one place. Reserve the personal-channel (SMS) for on-call rotations or critical incidents.
The mistake people make is either using only email (alerts get missed) or going straight to SMS for everything (alerts get muted). The mix is the answer.