The on-call killer isn’t outages — it’s noise
Anyone who has carried a pager has been here. At 2:43 AM your phone shrieks. Site is down. You stagger to your laptop, ssh in, run curl, and the site is fine. By the time you check the monitor, it has cleared itself.
You go back to sleep. The next night it happens again.
After two weeks of this you put the phone on silent. The next outage — the real one — happens during your silent hours, and you find out about it from a customer on Slack at 9:14 AM.
False alarms aren’t just annoying. They make real alerts less effective by eroding trust.
Why monitors flake
Even a perfectly reliable site will sometimes look down to a monitor that runs once per minute. The usual suspects:
- Network blip between monitor region and your host. A BGP route flapped. A peering session dropped. Both happen.
- Local DNS hiccup. Especially common right after DNS changes.
- TCP RST due to firewall state table eviction. Less common, but real on busy edges.
- TLS handshake timeout because the load balancer was rotating a node mid-request.
- CDN cache miss + cold origin — first request gets a slow response, the timeout fires before the response completes.
- Genuine 5xx for a single bad request. Race condition, dropped DB connection, retried successfully on the next try.
None of these mean your site is down. They mean something was wrong for one request, somewhere. If you’re checking from a single location every 60 seconds, any one of these will trip your monitor.
The second-check pattern
The simplest mitigation is brutally effective: when a check fails, don’t alert immediately. Re-check from a different region first.
T+0s Region A check fails
T+5s Trigger re-check from Region B
T+10s Region B succeeds → false alarm, suppress
Region B fails → real outage, page now
That’s it. That’s the whole technique. The downside is you’ve added 5–10 seconds to your detection time. The upside is you’ve filtered out roughly 95% of single-region flakes.
This works because flakes are usually local to a path: the monitor’s region, the network between the monitor and your host, or your host’s link to one transit provider. A second check from a different physical region exercises a different network path, different DNS resolver, and different last-mile peering. It’s an independent sample.
The probability that two independent regions both see your site as down at the same instant is much lower than either one alone — unless your site is actually down, in which case both will agree.
How long should the re-check wait?
Short. Long enough that the original check has fully timed out and you’re not just hitting a transient micro-blip, but short enough that you’re not delaying a real-outage alert.
A common pattern:
- Initial check: 10-second timeout
- Wait: 5 seconds
- Second-region check: 10-second timeout
- If second fails: alert
That’s an upper bound of 25 seconds added to detection. For most use cases that’s a worthwhile trade.
Don’t go further
You will be tempted to add more confirmation regions. “Let’s wait until 3 of 5 regions all fail.” Don’t. Each additional region you require adds detection latency but doesn’t proportionally reduce false alarms.
Two independent samples already filter the vast majority of single-region flakes. Adding a third does little except delay the real alert.
The exception is if you’re monitoring something with known regional variation — say, a CDN edge that genuinely has different reachability from different continents. In that case you might intentionally alert per-region rather than requiring global agreement.
A subtler trap: same monitor, different region, same provider
The second-region trick relies on independence. If your “Region A” and “Region B” monitors are both running on the same cloud provider in different availability zones, you’re not really sampling independent paths. Both monitors will see the same blip if the provider’s network has issues.
For a meaningful re-check, use a different cloud or a different ASN entirely. Even better, mix data centers — one bare-metal monitoring node, one cloud node — so a cloud-wide incident doesn’t blind you.
What about real flapping?
Sometimes your site genuinely is flapping — coming up and down rapidly. The re-check pattern handles this gracefully too. If Region A says down, Region B says up at the moment of check, you log the inconsistency but don’t page. The next minute’s check from Region A will fire again. If three or four consecutive minutes all show “A down, B up”, that’s a pattern worth investigating manually — but it’s not a midnight page.
For genuinely flapping incidents (the site is dying intermittently), you want the dashboard to show the pattern, not the pager to wake you for each oscillation.
The implementation in practice
If you’re rolling your own monitoring, the logic looks something like:
async function checkSite(url) {
const primary = await fetchWithTimeout(url, { timeout: 10000, region: 'us-east' });
if (primary.ok) return { status: 'up' };
// Primary failed — re-check from a different region before alerting
const secondary = await fetchWithTimeout(url, { timeout: 10000, region: 'eu-west' });
if (secondary.ok) {
log.info('primary-failed-secondary-ok', { url, primary, secondary });
return { status: 'up', note: 'single-region-blip' };
}
return { status: 'down', confirmed: true };
}
If you’re not rolling your own, this is the kind of thing a managed monitoring service should do for you. It’s worth asking your vendor explicitly — not every uptime tool implements it, and the ones that don’t will wake you up for blips.
The takeaway
Sleep is a finite resource. Every page that wasn’t a real outage costs you trust in the alert and rest you can’t replace.
A 5-second re-check from a different region is the highest-ROI piece of alerting logic you can add. It costs almost nothing in detection time and removes most of the noise that makes uptime monitoring miserable on-call.
If you’re building it yourself, build this. If you’re buying it, ask whether your provider does.
MyUptimeBot does the second-region re-check by default on every plan, including free. Try it →