Customer-Facing Status Updates During an Outage
What to tell customers during an outage, when, and how — without over-promising or panicking them. A senior SRE's guide to status-page updates that build trust.
- #incident-response
- #sre
- #on-call
- #communication
- #reliability
When your service is down, your customers are having their own small incident — a checkout they can’t complete, a deploy they can’t ship, a dashboard that won’t load — and the thing that most shapes how they feel about it isn’t the outage itself. It’s whether you told them what was happening. A twenty-minute outage handled with clear, honest status updates does less damage to trust than a five-minute one met with total silence, because silence reads as “they don’t know it’s broken” or, worse, “they don’t care.”
Writing good customer-facing status updates under pressure is a real skill, and it’s separate from the internal comms and the technical work. This is what to tell customers, when, and how — the craft of the outward-facing update that keeps trust intact even when your service isn’t.
Why the update matters more than the outage
Customers are remarkably forgiving of outages if they feel informed and respected. What erodes trust is the sense of being left in the dark — refreshing a dead page, not knowing if it’s them or you, not knowing if you’re even aware. A prompt “we know, we’re on it, here’s when we’ll update you next” transforms that experience. It says: this is being handled by competent people who respect your time. That message, delivered early, is worth more than shaving a few minutes off the fix.
So treat the status update as a first-class part of the response, not an afterthought you get to once the fire’s out. Acknowledging the problem publicly is often the single most valuable thing you can do in the first few minutes for the customer relationship.
What a good update contains
A customer status update needs surprisingly little, and adding more usually makes it worse:
WHAT CUSTOMERS ACTUALLY NEED:
- Acknowledgement: we know there's a problem.
- Scope: what's affected, in THEIR terms ("completing purchases",
not "the payments-service").
- What we're doing: "investigating" / "we've found it, fixing" —
enough to show it's being worked.
- When we'll update next: a concrete time. This is the key one.
WHAT TO LEAVE OUT:
- Technical root cause / internal jargon.
- Blame (a vendor, an engineer, "a bad deploy").
- A specific fix ETA you can't guarantee.
- Speculation about what "might" be wrong.
The load-bearing element is the next-update time. Customers can tolerate not having a fix ETA; what they can’t tolerate is open-ended silence. “We’ll update by 14:30” gives them permission to stop refreshing and go do something else, which is a genuine kindness during an outage.
The update lifecycle
Match your message to the phase of the incident. The standard progression:
- Investigating — “We’re aware of an issue affecting [thing customers do] and are investigating. Next update by [time].” Post this fast, before you know the cause. Acknowledgement first.
- Identified — “We’ve identified the cause and are working on a fix.” Don’t say what the cause is; just that you found it. Reassuring without over-sharing.
- Monitoring — “We’ve applied a fix and things are recovering; we’re monitoring to confirm.” This manages the recover-then-relapse case — you’re not claiming victory yet.
- Resolved — “This is resolved as of [time]. We’re sorry for the disruption.” Clean, human, done.
Keep each one short. A status update is not an essay; it’s a signal.
Cadence: update even when there’s no news
The hardest discipline is posting on your promised cadence even when nothing has changed. “We’re still investigating and don’t yet have an estimate; next update by [time]” feels unsatisfying to write, but it’s exactly what maintains trust. A missed update time is a broken promise, and during an outage your customers are watching whether you keep your promises. Post the boring “no change yet” update on schedule. It reassures far more than it disappoints.
Tone: calm, honest, human
The register that works is calm and plainly honest, neither corporate-robotic nor panicked. Some guidance that’s saved me:
- Own it simply. “We’re sorry for the disruption” — once, sincerely. Over-apologizing reads as panic; none reads as cold.
- Don’t over-reassure. “Everything is totally fine!” while things are clearly broken destroys credibility. Match the message to reality.
- Never blame. Not the vendor, not an engineer, not “a bad deploy.” Customers don’t care whose fault it is, and public blame ages terribly.
- No premature all-clears. Declaring “resolved” before you’re sure, then having it break again, costs more trust than the original outage. Confirm it held first.
- Plain language always. If your grandparent couldn’t understand the update, rewrite it. No jargon, no internal service names.
A copy-paste update sequence
[14:07 · Investigating] We're aware of an issue affecting
customers' ability to complete purchases. We're investigating
and will post an update by 14:30.
[14:20 · Identified] We've identified the cause and are working
on a fix. You may continue to experience errors at checkout
until this is resolved. Next update by 14:45.
[14:35 · Monitoring] We've applied a fix and checkout is
recovering. We're monitoring closely to confirm it's fully
resolved. Next update by 15:00.
[14:50 · Resolved] This issue is resolved as of 14:48 and
checkout is working normally. We're sorry for the disruption
and appreciate your patience.
Common mistakes
- Silence. The default and the worst. Acknowledge early, even before you understand the problem.
- Missing your own cadence. Promising an update by 14:30 and going quiet. Keep the promise even with “no change.”
- Technical detail and jargon. Customers don’t want your service names or root cause; they want to know if it affects them and when it’s fixed.
- Blame. Naming a vendor or a “bad deploy.” It never helps and always ages badly.
- Premature “resolved.” Declaring victory before the fix held, then reopening. Confirm stability first.
- Over-promising an ETA. A missed fix time breaks trust twice. Promise the next update, not the fix.
Where AI helps — draft fast, human approves
Customer comms is a strong AI assist because it’s high-stakes writing under time pressure that a human can quickly verify. Hand the model the facts and the phase and have it draft in the right register: “Write a customer-facing status update for a checkout outage — plain language, no jargon, no root cause, no blame, acknowledge and promise the next update in 20 minutes.” It reliably produces the calm, correctly-scoped update faster than you’ll write it mid-incident. The rule is firm: a human reads and approves every word before it’s published. The model doesn’t know what you’ve actually confirmed, and a customer-facing channel is the worst place for a confident overstatement. The Incident Response tool generates audience-tuned comms drafts, including customer-facing ones, from a structured assessment.
Wrapping up
During an outage, the customer update is not an afterthought — it’s one of the highest-leverage things you do for the relationship. Acknowledge fast, describe the impact in the customer’s own terms, say what you’re doing, and always promise a next-update time you then keep. Stay calm, honest, jargon-free, and blame-free, and never call it resolved until it truly is. Handled this way, an outage can actually build trust, because customers see a team that communicates like it respects them.
AI-drafted customer updates are drafts. A human reviews and approves every message before it’s published.
Related
- The Communications Lead Role in Incident Response
- Building a Stakeholder Notification Matrix for Incidents
- Incident Communication Templates: Status Page, Stakeholder, and Exec
Get 500 Battle-Tested DevOps AI Prompts — Free
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.