monitoring

Multi-region health checks: implementing redundant uptime monitoring

A monitor that checks from one location can't tell the difference between a global outage and a local network glitch. Multi-region health checks fix that — but only if you implement them right. Here's the architecture, the pitfalls, and the math behind getting it to work.

MyUptimeBot Team · May 26, 2026 · 6 min read · For developers

The problem with single-region monitoring

A single probe checking your site from one location is enough to detect the most blatant outages: the server is off, DNS is broken, the SSL cert expired. It catches the failures that affect everyone.

What it can’t catch is the difference between:

  • Your site is down for everyone
  • Your site is down for the probe’s region
  • The path between the probe and your site has issues, but most users are fine

To a single-region monitor, all three look identical. You get an alert. You wake up. You check the site from your phone, which uses a different network path — and the site works. Now you’ve burned an hour of sleep on a transient blip that didn’t affect anyone real.

Multi-region monitoring solves this by sampling from independent locations. The question changes from “is the site down?” to “is the site down from where?” — and the answer drives smarter alerting decisions.

What “independent” actually means

The value of multi-region monitoring is independence between probes. If your “second region” is in the same data center as your first region, it’s not actually a second region — it’s two probes in the same failure domain.

Real independence means:

  • Different physical locations. Different cities, ideally different continents for global services.
  • Different network providers. Different last-mile transit, different ISPs, different peering arrangements.
  • Different cloud providers if possible. AWS us-east-1 and AWS us-west-2 are physically separate but share AWS-wide failure modes. Mixing AWS with GCP gives stronger independence.
  • Different DNS resolvers within each probe. If both regions ask the same recursive resolver, that resolver becomes a shared failure point.

The cheapest improvement most monitoring setups can make is running probes on at least two different cloud providers. Even two of the major three (AWS, GCP, Azure) gives you protection against any single-provider event taking out your monitoring.

The 2-of-3 (or 2-of-N) rule

Once you have multiple regions, you need a policy for what counts as a “real” failure. The simplest workable rule:

Alert only when ≥2 independent regions agree the site is down.

This is sometimes called the “2-of-3 quorum” or “majority report” approach. The logic:

  • One region fails → likely a local network blip. Investigate later if it persists. Don’t page.
  • Two regions fail → real outage signal. Page.
  • Three regions fail → definitely a real outage. Page immediately.

The trade-off is detection latency. If you require two regions to agree before alerting, you’ve added the time to run the second check before paging. Usually 5–10 seconds. For most cases, that’s a worthwhile trade for the false-alarm reduction.

Architecture: poll vs. push vs. heartbeat

There are three ways to set up multi-region checks:

Centralized polling. A single coordinator service tells each region “check this URL now.” Regions report results back. The coordinator combines them and decides whether to alert.

  • Pro: Simple to reason about, easy to add new regions.
  • Con: The coordinator is a single point of failure. If it goes down, all monitoring goes silent.

Independent regions with shared state. Each region checks independently on its own schedule. Results write to a shared store (a database, a queue). An alerting service reads the store and computes “did N regions fail in the last interval?”

  • Pro: No single coordinator. More resilient.
  • Con: More moving pieces. Slight clock-skew issues across regions.

Heartbeat-based. Each region acts as both a checker and a beat-sender. The absence of a beat from a region is itself a signal — either the region is down (don’t trust its checks) or it’s running fine.

  • Pro: Self-healing. Bad regions remove themselves from the quorum.
  • Con: Hardest to implement correctly, especially around clock-skew and split-brain scenarios.

For most teams, the “independent regions with shared state” pattern is the sweet spot. It’s resilient, well-understood, and most cloud platforms make the building blocks (queues, databases, lambda functions) easy to assemble.

A reference implementation sketch

If you’re rolling your own, the bones look like this:

[Region A probe] ─┐
[Region B probe] ─┼──► [Result store] ─► [Alerter] ─► [PagerDuty/Slack/etc]
[Region C probe] ─┘

Each probe is a small process (a cron job, a lambda, a kubernetes job) that:

  1. Runs every N seconds on its own schedule
  2. Performs the check (HTTP request with timeout, possibly TLS validation)
  3. Writes a result record: {region, target, timestamp, status, latency, error}

The alerter reads recent results from the store and applies a quorum rule:

def should_alert(target, window_seconds=120, min_failures=2):
    recent = result_store.query(target, since=now - window_seconds)
    regions_failed = {r.region for r in recent if r.status != 'up'}
    return len(regions_failed) >= min_failures

Real implementations have more nuance — debouncing, recovery detection, alert suppression for ongoing incidents — but the core is that simple.

Choosing regions

For most production sites, 3 regions hits the diminishing-returns wall. Pick them to maximize independence:

  • One in your primary user geography (e.g., us-east if most users are in North America)
  • One in a different continent (e.g., eu-west)
  • One on a different cloud provider (e.g., GCP us-central even if AWS is your main host)

Five regions adds geographic detail but rarely changes the alerting outcome — by the time three regions fail, you knew it was a real outage anyway.

Going beyond five regions is usually only valuable for services with truly global user bases where regional performance matters in itself, not just for confirming outages.

The cost question

Multi-region probing costs more than single-region. For most monitoring providers, it’s bundled in pricing tiers — basic plans give you a single region, mid-tier plans give you 2–3, top tier or enterprise gives you many.

If you’re building your own, the marginal cost of each region is small — a lambda invocation every minute is cents per month — but it adds up across many targets. A team monitoring 100 endpoints from 3 regions every 30 seconds is doing 17,280 checks per hour. Most of those are cheap, but the coordination overhead and result storage become real engineering work at that volume.

Practical pitfalls

A few mistakes I’ve seen teams make:

Probes on the same cloud as the monitored service. If your site runs on AWS and your monitoring probes also run on AWS, an AWS-wide event takes both down. The whole point of external monitoring is to be external. Use a different provider for at least your primary monitoring probe.

Probes inside your own firewall. “Internal monitoring” that lives inside your VPC can verify health-check endpoints but doesn’t tell you whether real users on the internet can reach you. The DNS, BGP, edge proxy, and TLS layers between users and your origin are exactly the layers internal monitoring skips.

Treating probe-side failures as application failures. A probe in Singapore failing to reach your site doesn’t necessarily mean the site is broken — it might mean the probe’s network has issues. Distinguish “the probe is failing” from “the target is failing” by health-checking the probe itself.

Single shared DNS resolver. Even with probes in three regions, if they all ask the same recursive resolver for DNS, that resolver is a shared failure point. Configure each probe to use a local resolver.

What MyUptimeBot does, plainly

We run probes from multiple regions on multiple cloud providers, and our alerting layer requires confirmation from at least two regions before paging. Single-region flakes get logged but don’t wake you up. That second-region re-check is the difference between “your phone screams for every cosmic ray that hits us-east” and “your phone only screams when something real happens.”

If you’re building this yourself, the architecture above is what to aim for. If you’d rather not, we already did it — the free plan covers one monitor from our multi-region checking infrastructure.

The principle, either way: never trust a single probe. Independence between checks is what turns noise into signal.

MyUptimeBot
Watching the internet

We build a friendly bot that watches your websites every 30 seconds and alerts you the moment something breaks. Notes here come from running that infrastructure, talking to the people who depend on it, and reading the postmortems no one publishes.

Stop finding out from customers.

One monitor, free forever. A friendly bot doing the worrying for you.