Skip to content
DevOps AI ToolKit
Home

Incident Response: Triage, Comms and Postmortems

Faster RCAs, postmortems, runbooks, and on-call workflows powered by AI.

133 copy-paste prompts · 94 in-depth guides Jump to prompts Jump to guides

What actually breaks, and what to check first

The most expensive habit in incident response is diagnosing before mitigating. Understanding the cause is satisfying and frequently slow, while restoring service is often fast — a rollback, a feature flag, a failover. Every minute spent understanding a failure you could have stopped is a minute of impact you chose to accept, and the cause will still be there afterwards for a calmer investigation.

The second is that incidents fail on communication far more often than on technical skill. When nobody knows who is deciding, several people apply conflicting fixes at once; when nobody is writing anything down, the postmortem is reconstructed from memory and gets the timeline wrong. Separating the person deciding from the people investigating is the highest-leverage structural change available, and it costs nothing.

Postmortems are where most of the value is realised or lost. A postmortem that identifies a person has found a symptom; the useful question is why the system allowed a normal human action to cause an outage, and what would have caught it sooner.

Triage order

Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.

  1. What is the actual user impact?

    Name it concretely: which users, which operations, since when, and how you know.

    Impact drives severity, and severity drives everything else. "The dashboard is red" is not impact. Anchoring on a real user symptom stops the response from chasing an alert that does not matter.

  2. What changed?

    Deployments, feature flags, config changes, certificate expiries and scheduled jobs in the window before onset.

    Most incidents follow a change. Establishing the change window early either finds the cause quickly or rules out an entire category — and the negative result is genuinely useful.

  3. Can you mitigate now, without knowing the cause?

    Roll back, disable the flag, fail over, shed load, scale out.

    Do this before diagnosing. Mitigation and diagnosis are separate activities, and the order matters: stop the impact, then investigate with the pressure off.

  4. Who is deciding, and who is communicating?

    Name an incident lead and a communications owner explicitly, out loud.

    Unnamed roles mean either nobody or everybody acts. The lead does not have to be the most senior engineer — they have to be the one person deciding what gets tried next.

  5. Is anyone writing it down as it happens?

    A running timeline: time, observation, action, who did it.

    Reconstructed timelines are wrong in the details that matter most. Written contemporaneously, this is also the artefact that makes the postmortem take an hour instead of a day.

Diagnostic commands

Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.

kubectl rollout undo deployment/<name>
Changes state

Revert a workload to its previous version.

How to read it: Usually the fastest available mitigation. Confirm first that no database migration ran with the deploy — the workload rolls back, the schema does not.

git log --since="2 hours ago" --all --oneline
Read-only

List everything merged in the incident window across branches.

How to read it: Establishes the change set quickly. Combine with the deployment log — merged and deployed are different times, and the second is the one that matters.

kubectl get events -A --sort-by=.lastTimestamp | tail -50
Read-only

See what the cluster observed most recently.

How to read it: Strong first command when the blast radius is unclear. Events expire after about an hour, so capture them early — they are frequently gone by the time the postmortem is written.

curl -sS -o /dev/null -w "%{http_code} %{time_total}s\n" <endpoint>
Read-only

Measure what a user actually experiences from outside.

How to read it: Internal dashboards can be healthy while users fail. A synthetic check from outside the perimeter settles that disagreement immediately.

openssl s_client -connect <host>:443 -servername <host> </dev/null 2>/dev/null | openssl x509 -noout -dates
Read-only

Check certificate validity dates.

How to read it: Certificate expiry causes total, sudden, well-correlated outages and is trivially ruled in or out. Worth checking early precisely because it is so cheap.

Failure modes

These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.

The incident runs for hours with no clear owner.

Cause:
No incident lead was named, so decisions were made by whoever spoke most recently.
Fix:
Name a lead explicitly and out loud. Their job is deciding what is tried next and keeping only one change in flight at a time.

Several people apply fixes simultaneously and the system gets worse.

Cause:
Uncoordinated changes, with no single record of what has been tried.
Fix:
One change at a time, announced before it is made, recorded in the timeline. Parallel changes make it impossible to attribute the improvement or the regression.

Service is restored but the postmortem produces no change.

Cause:
The postmortem identified a cause but no action with an owner and a date.
Fix:
Every action item needs an owner, a date, and a place it is tracked. Items without those are a record of good intentions.

The same incident recurs a month later.

Cause:
The mitigation was applied and the underlying cause never addressed — an accepted risk that nobody decided to accept.
Fix:
Separate the mitigation from the fix in your tracking, and give the fix a date. Recurrence is otherwise the default outcome.

Nobody can reconstruct what happened.

Cause:
No contemporaneous timeline, and the ephemeral evidence — events, logs with short retention, dashboards with rolling windows — expired.
Fix:
Capture evidence during the incident, not after. Screenshot the dashboard, save the events, export the logs. Much of it is unavailable within the hour.

Common mistakes

  • Diagnosing before mitigating. Restore service first; the cause will still be available to investigate afterwards.
  • Letting the most senior person lead by default. The lead role is coordination, and it is often better filled by someone who is not deep in the debugging.
  • Making several changes at once, which destroys your ability to tell what helped.
  • Treating "the alert cleared" as resolution without confirming real user impact has ended.
  • Writing postmortems that name a person. The question is why the system permitted a normal action to cause an outage.
  • Skipping the postmortem for short incidents. A brief incident with a near-miss cause is exactly the cheap warning worth acting on.

Frequently asked questions

Should I find the root cause before fixing an incident?

No — mitigate first. Restoring service and understanding the failure are separate activities, and mitigation is usually much faster: a rollback, a feature flag, a failover. Every minute spent understanding a failure you could already have stopped is impact you have chosen to accept, and the evidence you need for diagnosis does not disappear once users are working again.

What roles does an incident actually need?

At minimum, someone deciding and someone communicating — and they should not be the same person, because updating stakeholders while debugging means doing both badly. Larger incidents add a scribe keeping the timeline and separate investigation leads per subsystem. The critical property is that the roles are named out loud; unnamed roles are filled by nobody or by everybody.

How long should a postmortem take to write?

About an hour, if a timeline was kept during the incident, and most of a day if it was not — which is the strongest practical argument for keeping one. Write it within a few days while memory is accurate, and keep it focused on what the system allowed rather than on what a person did.

What makes a postmortem action item useful?

An owner, a date, and a tracking location. Items without all three are intentions, and they are why the same incident recurs. It also helps to distinguish the mitigation you already applied from the fix that prevents recurrence — otherwise the mitigation gets ticked off and the real work quietly disappears.

Do small incidents need postmortems?

The short ones with a near-miss cause are often the most valuable, because you learned something at very low cost. A five-minute outage that revealed a missing alert is a cheap warning about a longer one. Keep the format proportionate — a paragraph and one action item is a legitimate postmortem.

Prompts

Guides

Recommended tools

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week