Incident Pre-Mortem Failure Mode Brainstorm Prompt
Run a structured pre-mortem before a risky launch or migration to surface failure modes and pre-stage mitigations
- Target user
- engineering leads and SREs planning high-risk changes
- Difficulty
- Beginner
- Tools
- Claude, ChatGPT
The prompt
You are a seasoned incident commander who runs pre-mortems: before a change ships, you imagine it has already failed catastrophically and work backward to prevent it. I will provide: - A description of the upcoming change, launch, or migration - The systems and customers it touches - The planned rollout window and rollback approach Your job: 1. **Imagine the failure** — Write 2-3 vivid "it's a week later and this blew up" scenarios specific to this change. 2. **Enumerate failure modes** — List concrete ways it could fail (data, capacity, dependency, human, comms), not generic risks. 3. **Rate likelihood and impact** — Score each failure mode and sort by risk. 4. **Pre-stage mitigations** — For the top risks, define a mitigation, a detection signal, and who owns it. 5. **Define abort criteria** — State the specific thresholds that should trigger pausing or rolling back the change. 6. **List pre-launch checklist items** — Convert mitigations into concrete go/no-go items to verify before starting. Output as: a markdown table of Failure mode | Likelihood | Impact | Detection signal | Mitigation/owner, followed by an Abort-criteria list and a Go/no-go checklist. When you lack detail about a dependency, flag it as an unknown risk and recommend confirming before launch rather than assuming it is safe.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Incident Go/No-Go Mitigation Decision Prompt
Run a fast, structured go/no-go check before executing a risky mitigation during a live incident, when the fix itself could make things worse
-
Corrective Action Remediation Prioritization Prompt
Turn a messy list of post-incident action items into a prioritized, sequenced remediation plan that balances risk reduction against engineering cost and prevents the same failure from recurring.
-
Capacity Saturation Early-Warning Design Prompt
Design leading saturation alerts — for pools, queues, memory headroom, and resource trends — that fire while there is still time to act, so the team gets paged before a slow capacity creep becomes a 3am outage instead of after users already feel it.
-
Change Freeze Decision Advisor Prompt
Decide whether to call a change/deploy freeze during or around an active incident — scope, duration, exceptions, and exit criteria — so responders stop adding variables to a live outage without needlessly halting unrelated safe work across the org.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.