Don’t immediately try to fix anything
The instinct when an alert fires is to dive into the server, the logs, the deploy history. That instinct is wrong for the first two minutes. Before you change anything, confirm what’s actually broken — because half of the time, the alert is misleading.
The first move is verification, not remediation. Take the URL from the alert. Open it in a fresh browser tab on your phone, not your laptop. Use cellular data, not your home Wi-Fi. If you can, ask someone in another location to do the same.
Three reasons to do this:
- Sometimes the site isn’t actually down. Local network blips, captive-portal Wi-Fi, an old DNS cache on your laptop, a browser extension blocking the page. All look like outages from your perspective and aren’t outages at all.
- Sometimes only part of the site is down. The homepage loads but the API is failing. The marketing site is fine but the dashboard is 500ing. Your fix depends on which.
- Sometimes the outage is regional. Cloudflare’s London edge is having a moment but US users are fine. The mitigation for “globally down” is different from “regional”.
Two minutes of verification beats two hours of fixing the wrong thing.
Minute 0–2: confirm and triage
- Open the affected URL on a non-work network. Is it actually broken?
- If you have monitoring from multiple regions, check: is it all regions, or just one?
- Open your hosting provider’s status page (or
downdetector.com, oristheinternetdown.com) — is this a known platform issue? - Open your CDN’s status page if you use one (Cloudflare, Fastly, CloudFront)
- Open your DNS provider’s status page
If a major provider is down, you’ve already done your most important diagnostic — and the answer is “wait it out, post a status update, stop debugging.” Most of the time it’s not them. But the 10% of the time it is, learning that in minute 2 saves you from spending the next hour chasing your own infrastructure.
Minute 2–4: tell people
This is the step everyone skips and everyone regrets later.
If you have a status page, post a brief note: “We’re investigating reports of [thing]. We’ll update in 15 minutes.” That’s it. No diagnosis. No timeline.
If you don’t have a status page, your customer-facing channels are it: Twitter / X, your support email auto-reply, a banner on whatever pages you can still reach. The point isn’t to deliver information — you don’t have any yet. The point is to head off the support tickets.
If you have a team, post in your incident channel: “I’m on [thing]. Status: investigating. I’ll update at HH:MM.” Even if there’s nobody else on, a paper trail of your own thinking will save you from forgetting what you tried.
Two minutes of communication buys you 30 minutes of focus, because the inbound messages slow down.
Minute 4–7: form a hypothesis before changing anything
Now you can actually look at the broken thing. But don’t reflexively start running commands. The fastest path to recovery is:
- Identify the symptom precisely. Which URL? What’s the HTTP status code? What’s the error message? Is it consistent or intermittent?
- Identify what changed recently. Deploys, config changes, third-party updates, scheduled cron jobs, SSL renewals. Pull up your deploy log. The most common cause of an outage is the most recent change.
- Form one hypothesis. Not three. One. “Probably the deploy that went out at 14:42.” Now you have something to test.
If your recent-change list is empty and nothing changed, your hypothesis space expands to: infrastructure (hosting outage, DNS, CDN), expiry (cert, domain, payment), or traffic (DDoS, viral spike, bot attack). Each has different fingerprints.
Minute 7–10: roll back, don’t debug
If you have a hypothesis and the hypothesis is “recent change”, the right move is almost always to roll back first and debug afterward.
This is the most counterintuitive piece of incident response. Engineers want to fix the bug. The business wants the site working. The fastest way to get the site working is to revert to the last known good state, then figure out what was wrong.
“Rolling back” means:
- For a code deploy: redeploy the previous version. Most platforms have a one-click rollback. Use it.
- For a config change: revert the config to what it was an hour ago. The change history in your config management is your friend.
- For an infrastructure change: undo whatever you just did. If you can’t, escalate.
- For a third-party update (library, plugin, theme): pin the previous version.
The “debug now, fix forward” instinct feels productive but lengthens the outage. The “revert first, then debug” instinct feels lazy but shortens it. Most postmortems include the line “if we had rolled back at minute 5 instead of debugging until minute 45, this would have been a 5-minute incident.” Almost nobody does this in the moment.
If the hypothesis isn’t a recent change, you don’t have a rollback to do — and you’re going to have to debug live. Set a 10-minute timer for yourself. If you haven’t found root cause in 10 more minutes, escalate or post a longer outage update. Don’t disappear into a rabbit hole for an hour.
What not to do
A few things that make outages worse:
- Don’t restart servers reflexively. It feels like progress but often loses the diagnostic state you needed.
- Don’t push fixes you haven’t tested. A bad fix during an outage extends it. If you wouldn’t push it normally, don’t push it now.
- Don’t disable monitoring. People do this to silence alerts during an incident. Then they forget to re-enable it. Then a new outage three days later goes undetected.
- Don’t argue with people in the incident channel. Discussion of root cause happens after the site is back up.
- Don’t promise specific recovery times. “We’ll have it back in 5 minutes” sets up failure if it actually takes 45.
The follow-up
When the site is back up, two things should happen within the next 24 hours:
- A short retrospective. What was the trigger? What was the root cause (often different from the trigger)? How long did each phase take? What slowed down recovery? Write it down. The format doesn’t matter — a Notion page, a Slack thread, a Google Doc. The act of writing is the value.
- One concrete improvement. Not five. One. The improvement that, if it had been in place, would have shortened this specific incident. Maybe it’s better monitoring on the thing that broke. Maybe it’s a rollback button. Maybe it’s a checklist for the kind of change that caused this.
Don’t try to fix all your reliability problems in one sprint. The incident gave you one specific thing to fix. Fix that. Next incident will give you the next thing.
A working playbook, condensed
If you want a single page to look at during the next outage:
T+0–2: Confirm the outage. Check from outside your network.
Check provider status pages.
T+2–4: Post a "we're investigating" notice.
Tell your team.
T+4–7: Form one hypothesis. What changed recently?
T+7–10: If the hypothesis is "recent change," roll back.
If it isn't, debug for 10 more minutes, then escalate.
After: Short retrospective. One concrete improvement.
That’s the whole thing. Stuck on a wall in your office, in your team’s wiki, or just in your head. The point isn’t to follow it religiously — every outage is different. The point is that having a playbook at all prevents the worst version of an incident, which is the one where you panic and start guessing.
And one more thing
Monitoring is what makes any of this possible. The 10-minute playbook only starts if you know the site is down — which you don’t, by default, until a customer tells you. The single highest-ROI thing you can do for incident response isn’t a fancy escalation policy or a status page. It’s setting up basic uptime monitoring so the clock starts when the outage starts, not when someone happens to notice.