Incident Recovery Verification Checklist Prompt
Build a rigorous all-clear checklist so an incident is declared resolved only after recovery is verified end-to-end — not just when the obvious symptom disappears.
- Target user
- Incident commanders and SREs deciding when to call all-clear
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a senior SRE who has seen incidents re-open an hour after a premature all-clear because someone confirmed the dashboard was green but not that the system was actually healthy. I will provide: - The affected service(s) and architecture - The primary symptom and the mitigation applied (rollback, failover, scale-up, flag flip) - Available signals (SLO dashboards, synthetic checks, queue depths, error rates) - Downstream consumers and any data integrity concerns Build a recovery verification checklist. Work through these steps: 1. **Separate symptom from health** — list the difference between "the alert cleared" and "the system is genuinely recovered." Name the false-recovery traps for this service (cached results, drained-then-refilling queues, masked errors, partial failover). 2. **Define verification layers** — checks at each layer: (a) the failing signal itself, (b) golden SLO signals, (c) synthetic / real user journeys, (d) downstream dependents, (e) data integrity / backlog drain, (f) the mitigation's side effects (e.g., is the rollback stable, is the scaled-up capacity sustainable). 3. **Set hold-and-watch criteria** — how long signals must stay healthy before all-clear, and what bounce-back would re-open the incident. 4. **Handle leftover risk** — temporary mitigations still in place (a flag off, capacity over-provisioned, a node cordoned) that must be tracked as follow-ups, not forgotten at all-clear. 5. **Verify the backlog** — queues, retries, dead-letter, delayed jobs, and reconciliation that must be confirmed drained or scheduled. 6. **Write the all-clear gate** — the explicit go/no-go the commander reads aloud before declaring resolution, plus who must confirm. Output: (a) a layered verification checklist with pass criteria, (b) the false-recovery trap list for this service, (c) hold-and-watch durations with bounce-back triggers, (d) a follow-up tracker for leftover mitigations, (e) the spoken all-clear gate. Bias toward proving recovery, not assuming it.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Incident Data Integrity Verification After Recovery Prompt
Verify that data is actually correct and consistent after a service is restored, before declaring the incident resolved, when an outage may have corrupted or skipped writes
-
Incident Stand-Down and All-Clear Criteria Prompt
Decide whether an incident is genuinely resolved enough to declare all-clear and stand down responders, versus prematurely closing a still-fragile system
-
Recovery Smoke-Test Suite Generator Prompt
Generate a fast, scriptable smoke-test suite that proves a service is genuinely healthy after a mitigation or restart — covering critical user journeys, data integrity, and downstream dependencies — before you declare an incident resolved.
-
Runbook-to-Automation Script Converter Prompt
Convert a manual incident runbook into a safe, idempotent automation script or ChatOps command — with guardrails, dry-run mode, confirmation gates, and rollback — so proven remediations execute reliably at 3am instead of being fat-fingered under pressure.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.