The numbers, translated
Uptime is usually quoted as a percentage of time a service is available. The convention is to talk in “nines.”
| Uptime | Downtime per year | Downtime per month | Downtime per day |
|---|---|---|---|
| 99% (two nines) | 3 days 15h | 7h 18m | 14m 24s |
| 99.5% | 1 day 19h | 3h 39m | 7m 12s |
| 99.9% (three nines) | 8h 45m | 43m 12s | 1m 26s |
| 99.95% | 4h 22m | 21m 36s | 43s |
| 99.99% (four nines) | 52m 35s | 4m 19s | 8.6s |
| 99.999% (five nines) | 5m 15s | 25.9s | 0.86s |
When someone says “we offer 99.9% uptime” they’re telling you they reserve the right to be down for nearly 9 hours a year, distributed however the failures fall. When someone says “99.99%” they’re committing to less than an hour. Five nines means less than 6 minutes a year, which is harder than it sounds — a single bad deploy followed by a slow rollback can blow your entire annual budget in one afternoon.
Why the gaps between tiers are massive
Each additional nine is roughly 10x harder to deliver than the one before it. The progression isn’t linear — it’s exponential. You can build 99% with one well-maintained server. 99.9% requires redundancy and monitoring. 99.99% requires multi-region failover, automated incident response, and chaos engineering. 99.999% requires a small team whose entire job is to drive that number, plus customers willing to pay the costs that produces.
In dollar terms, going from 99% to 99.9% might cost you a few hundred bucks a month for better hosting and basic monitoring. Going from 99.99% to 99.999% can cost millions for the staffing and infrastructure required to chase those last few minutes a year.
What’s in the SLA fine print
When a provider advertises an uptime SLA, the actual contract usually has more exclusions than the headline number suggests. Common carve-outs:
- Scheduled maintenance windows don’t count as downtime
- Force majeure (natural disasters, government action) doesn’t count
- Third-party failures outside the provider’s control don’t count
- Your own misconfiguration doesn’t count
- DDoS attacks above some threshold may not count
- Single-AZ services may have a different (lower) SLA than multi-AZ
So a “99.99% SLA” might really mean “99.99% during business hours not coinciding with scheduled maintenance, excluding the day Cloudflare went down.”
The credit for missing the SLA is also usually structured as a service credit — a partial refund of what you paid, not compensation for the business you lost. If your e-commerce site goes down for 2 hours during Black Friday and you’d otherwise have done $50K of revenue, your provider’s SLA gets you a few dollars off your next bill. Not a refund of your lost revenue.
What you should actually measure
The number that matters to you isn’t the provider’s headline SLA — it’s your end-to-end availability as a customer would experience it. That includes:
- Your hosting provider’s uptime
- Your CDN’s uptime
- Your DNS provider’s uptime
- Your SSL certificate not being expired
- Your application not throwing 500 errors
- Your third-party dependencies (Stripe, auth providers, analytics) still working
The product of all those is your actual uptime, and it’s typically lower than any single component’s SLA. If five different services each promise 99.9%, your combined availability is at best 99.5% — and that’s only if their failures never overlap.
This is why people who manage uptime professionally measure it from the outside. They put a monitor on a public URL that exercises the full stack — DNS resolution, TLS handshake, HTTP request, response time, application logic — and treat that measurement as the source of truth. The vendor SLAs become reference points, not promises.
How to set a realistic target
Pick a target by working backwards from what an outage actually costs you:
- Personal site / hobby project: 99% is fine. A handful of hours of downtime a year is invisible to your traffic.
- Small business marketing site: 99.5–99.9%. Aim for “rare but tolerable.”
- E-commerce, lead capture, anything that converts: 99.9%+. Every hour down is measurable revenue loss.
- Production SaaS with paying customers: 99.95%+. Your customers will notice and might churn.
- Infrastructure other businesses depend on: 99.99%+. You’re now writing SLAs of your own.
- Critical infrastructure (payments, healthcare, transportation): 99.999%. You need a team for this.
The cost-to-deliver climbs faster than the cost-of-downtime saved, eventually. Most businesses overshoot when they reach for five nines for non-critical systems and undershoot when they ignore monitoring entirely.
A practical framework
If you don’t know your current uptime, start measuring before you set a target. A free uptime monitor running for 30 days will tell you what your baseline is. You can’t improve a number you don’t track.
If your baseline is below your target, the highest-ROI fixes are usually:
- Add monitoring so you know when you’re down (most outages last longer because nobody noticed for 30 minutes)
- Renew SSL certificates automatically and monitor expiry — this single change kills a category of outages
- Set up DNS with multiple resolvers at a managed provider — your registrar’s bundled DNS is often the weakest link
- Cache static content on a CDN so your origin can be slow without taking everything down
- Practice your incident response so the recovery time after detection is short
Most of those are cheap or free. The “next nine” of uptime is usually in operational discipline, not infrastructure spend.
The marketing point
When you see “99.9% uptime” in someone’s sales pitch, read it as “we will be down for up to 9 hours a year, and the contract gives you a discount, not a refund, when we are.” That’s not necessarily bad — it’s appropriate for most use cases. But it’s not the magic number it sounds like.
The number you want to track is your own end-to-end availability, measured from outside your stack, against the threshold that matches what your business can actually tolerate.
If you don’t have that measurement set up yet, start a free monitor and let it run for a month. The data will tell you whether you have an uptime problem worth fixing — and if you do, where the easiest wins are.