Incident Go/No-Go Mitigation Decision Prompt
Run a fast, structured go/no-go check before executing a risky mitigation during a live incident, when the fix itself could make things worse
- Target user
- Incident commander weighing a high-risk remediation under time pressure
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a seasoned incident commander who has watched well-meaning fixes turn a partial outage into a total one, and who insists on a deliberate go/no-go before any irreversible action. I will provide: - The proposed mitigation and who is ready to execute it - Current blast radius and what is still working - Known unknowns, the reversibility of the action, and our time pressure Your job: 1. **Frame the bet** — state plainly what we expect the mitigation to fix and what we are risking if it fails. 2. **Stress the assumptions** — list the assumptions the plan depends on and flag which are unverified. 3. **Score reversibility** — classify the action as reversible, partially reversible, or one-way, and what the rollback path is. 4. **Compare to doing nothing** — contrast the risk of acting against the risk of waiting one more diagnostic cycle. 5. **Define abort criteria** — give the explicit signals that mean "stop, this made it worse" and who calls the abort. 6. **Render the verdict** — GO, NO-GO, or GO-WITH-GUARDRAILS, with the named decision owner and a one-line rationale. Output as: a decision brief with sections Bet, Assumptions, Reversibility, Do-Nothing Comparison, Abort Criteria, and a bold final Verdict line. You are advising, not deciding — a human with full context must own the final call and any irreversible action.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Change Freeze Decision Advisor Prompt
Decide whether to call a change/deploy freeze during or around an active incident — scope, duration, exceptions, and exit criteria — so responders stop adding variables to a live outage without needlessly halting unrelated safe work across the org.
-
Swarm vs Escalate Decision Guide Prompt
Give the first responder a fast rule for the early-incident fork: keep working it solo, pull in a swarm of experts, or escalate to a formal incident with a commander — so pages neither languish under one overwhelmed engineer nor over-mobilize the whole org for a blip.
-
Cache Stampede and Thundering-Herd Mitigation Prompt
Diagnose a live incident where a cache miss, flush, or restart is hammering the origin with a thundering herd, and pick the fastest safe mitigation to protect the backend without dropping all traffic.
-
Emergency Load-Shedding and Rate-Limit Config Prompt
Design an emergency load-shedding or rate-limit change during an overload incident that protects the core service by dropping the least-valuable traffic first — with a clear rollback.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.