What good outage communication is actually optimising for
Key Takeaways
For busy support leads: the first message should go out before you know the cause, should name the symptom in the customer's words rather than your architecture, and should state when the next update comes. Those three properties matter more than everything you write afterwards.
- 1Acknowledge before you diagnose. Silence reads as either not knowing or not caring, and both are worse than an incomplete update.
- 2Name the symptom, not the subsystem. Customers experience "cannot log in", not "elevated error rates on the auth service".
- 3Commit to a next update time, not a fix time. You control when you speak. You do not control when it is fixed.
- 4Tag outage conversations separately. Otherwise a bad afternoon contaminates your quality metrics and you draw the wrong conclusion next quarter.
- 5Turn off the troubleshooting. During a platform incident, both humans and AI should stop diagnosing individual reports, because every one of those investigations is wasted and mildly insulting.
Most templates optimise for sounding calm. That is the wrong target, because calm without information is indistinguishable from evasion.
The right target is decision support. Every customer in your product during an incident has a decision to make: retry now, wait, switch to a manual process, warn their own users, or escalate internally. Your update either enables that decision or it does not.
This reframes the two hardest judgement calls. Should you publish before you know the cause? Yes, because knowing that the thing is broken and not on their end is already a decision they can act on. Should you admit you have no estimate? Yes, because "we do not have an estimate yet" tells someone to stop waiting and start their workaround, which is genuinely useful, whereas an invented estimate you miss removes their ability to plan and costs you the trust you were trying to protect.
Should you write templates before you need one
Yes, and the reason is not speed. It is that judgement is worst under pressure. Nobody writes a well-calibrated first message while an engineer is on a call describing a database failing over.
Write three, one per severity tier, each covering the same four fields: what is affected in customer language, what is not affected, what to do in the meantime, and when the next update lands. Leave the cause blank. Templates give you structure, not sentences, and an update that reads like it came from a template is its own failure.
The severity tiers, and what each one commits you to
| Tier | What it looks like to a customer | First update | Cadence | Who is on it |
|---|---|---|---|---|
| Degraded | Slower than usual, one feature flaky, workarounds exist | Within 15 minutes | Hourly | On-call engineer, support lead informed |
| Major | A significant capability is unusable for many customers | Within 5 minutes | Every 30 minutes | Incident lead plus a named comms owner |
| Critical | The product is unusable, or data or money is at risk | Within 5 minutes | Every 15 minutes | Incident lead, comms owner, and a founder |
These are our numbers and I am stating them as our recommended defaults rather than as an industry standard, because I cannot point you at a study that establishes them and I would rather say so than dress a preference up as research.
The column that gets skipped is the last one. Naming a comms owner separately from the incident lead is the single change that most improves outage communication, because otherwise the person best placed to write the update is the person you least want to interrupt.
What goes in the first message
Four sentences, in this order.
The symptom, in the words a customer would use. "Checkout is failing for some customers", not "elevated 5xx from the payments gateway".
The scope, honestly bounded. If you do not know how many customers are affected, say you are still determining it. Do not guess low, because revising upwards costs more credibility than starting vague.
What to do now. Retry, wait, use the manual path, or nothing.
When you will speak again. A specific time. This is the sentence people actually rely on, and it is the one most often left out.
What does not go in it: the word "some" doing load-bearing work when you mean most, an apology longer than the information, and any speculation about cause. Cause speculation published early and revised later is how an outage becomes a credibility story.
What your support team should do while the queue floods
Stop investigating individual reports. This is the hardest instruction to follow because investigating is the job, and during a platform incident every individual investigation reaches the same conclusion at the cost of twenty minutes.
Instead: one acknowledgement, sent to everyone reporting the symptom, linking the status page and stating the next update time. Tag every one of those conversations with an incident identifier. When it resolves, send one update across the whole tagged set rather than replying individually.
Then triage for the exceptions, because they are the valuable part. Somebody in that flood is reporting a different symptom, or the same symptom on a path you thought was unaffected, or a data consequence you have not noticed. A flood of identical reports is noise. The one report that does not match is the signal, and it is only findable if you are not busy answering the identical ones by hand.
Tag the incident conversations separately from normal support metrics too. An outage is not a support quality failure, and letting it into your response-time averages means you will spend next quarter explaining a number rather than improving anything.
Configuring the AI for outage mode
An AI support system during an incident will do the worst possible thing by default: it will helpfully troubleshoot. It will ask the customer to clear their cache, check their connection and confirm their password, while your platform is down. Every one of those exchanges wastes the customer's time and reads as though nobody has noticed.
So the incident checklist needs an entry for it. Whatever tool you use, you need a way to inject the current incident state into what the AI says and to bias it toward handing over rather than diagnosing. The behaviour you want is: acknowledge the known issue, link the status page, offer a person immediately. How that transition should feel is covered in the guide to human handoff, and during an incident it should be looser than usual, because the cost of an unnecessary escalation is much lower than the cost of an unnecessary troubleshooting loop.
The other half is remembering to turn it off. An outage notice still showing three days later does more reputational damage than the outage did.
The post-incident report, and who it is really for
Publish within a day or two for anything that affected customers materially. Cover what happened in plain language, why, how you fixed it, who was affected and for how long, and what changes prevent a recurrence.
The audience is not the customer who was affected. They already know, they were there. The audience is the customer evaluating you next quarter, the customer's boss asking whether this vendor is safe, and your own team, who will handle the next incident better for having written this one down.
Which means specificity is the whole value. A report that says you have improved monitoring says nothing. A report that says a particular check had a threshold that could not fire on the failure mode that occurred, and that the threshold is now different, is a document that builds trust. Vagueness in a post-incident report reads as either not having found the cause or not wanting to say it.
The published standard for this is higher than most teams assume. Cloudflare's own write-up of its 18 November 2025 outage states that the failure began at 11:20 UTC, that "as of 17:06 all systems at Cloudflare were functioning as normal", and that it was triggered by "a change to one of our database systems' permissions". Nearly six hours, a named cause, and timestamps you can check against your own logs. It is a vendor's account of its own incident, which is exactly why the timestamps matter: a team that publishes a timeline stops having to argue about one.
Should you compensate, and how much
Work it as arithmetic rather than as a gesture, with every input an illustrative assumption.
Assume a customer pays twelve hundred dollars a year, so a hundred dollars a month (illustrative assumption). Assume a four hour outage in a month with roughly two hundred working hours (illustrative assumption). Strict pro rata is four divided by two hundred, times a hundred dollars, which is two dollars. Two dollars is an insult. Do not send it.
The useful frame is not pro rata, it is what the interruption cost them. If those four hours blocked their own customers, the disruption is worth far more than the fraction of the subscription it consumed. Customers price it that way too: in the 2025 National Customer Rage Survey, a study of 1,000 US respondents run by CCMC with Arizona State's Center for Services Leadership and published by its own authors, "fifty-nine percent of customers reported that their problem wasted time (an average of one full day), 45% cited a financial loss (an average of $1,008)". Those are self-reported estimates rather than audited losses, and they are still an order of magnitude away from a pro rata credit. So the practical options are a meaningful credit, one that is a visible fraction of a monthly bill rather than of an hourly one, or a contractual service credit if your agreement defines one, or nothing plus a genuinely good post-incident report.
Offering nothing is a legitimate choice for short incidents and it is better than a token. What is not legitimate is waiting to be asked. Compensation that arrives because a customer complained reads as a settlement. The same amount offered unprompted reads as a standard.
Where this playbook breaks
Three cases, and they cover most real incidents.
Partial outages are the hardest and the most common. When something fails for eight percent of customers on one code path, every choice is bad. Notify everyone and you alarm the ninety-two percent unnecessarily and train them to ignore future notices. Notify nobody and the affected eight percent conclude you did not notice. Our answer is to publish on the status page at a degraded tier, and to reach out directly to the customers you can identify as affected. That requires being able to identify them, which is an engineering capability you have to build before the incident, not during it.
Third-party outages break the ownership model. When your payment provider or your email delivery is down, you did not cause it, you cannot fix it, and you have no visibility beyond their status page. Say exactly that, name the provider, link their status page, and resist both blaming them and pretending to have control. Customers are much more understanding about a dependency failure than about being managed.
Silent data problems break the timeline entirely. If records were written incorrectly for six hours and nothing appeared broken, your incident started long before you noticed and the communication is about remediation rather than restoration. The status page format does not fit, and the honest version is a direct message to affected customers explaining what is wrong with their data, what you are doing about it, and what they should not rely on in the meantime. This is the incident type that damages trust most and it is the one nobody has templates for.
One more thing that breaks: hosting your status page on the infrastructure that just failed. Host it somewhere else. This is embarrassing every single time it happens and it happens constantly.
What to do after it is over
Hold a retrospective covering the communication as well as the technical response, and treat a delayed first update as a defect with the same seriousness as a missing alert.
Produce two or three specific changes to the playbook. Not "communicate faster". Something like "the comms owner is named in the incident channel within two minutes" or "the status page banner has an expiry".
Then re-read the templates. They will be slightly wrong, because they always are, and fixing them while the incident is fresh takes ten minutes and pays for itself the next time.
If you want to see how a support tool behaves when you tell it to stop troubleshooting and hand over, that is worth testing before you need it rather than during. The free trial is the cheapest way to run that rehearsal.