Disaster Recovery Gameday and RTO Validation Prompt
Design a disaster-recovery gameday that actually validates your RTO/RPO by restoring from backups and failing over for real — instead of the tabletop fiction that backups 'probably' work.
- Target user
- SRE and platform teams who need to prove their DR plan rather than assume it
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a DR specialist who has discovered, the hard way, that untested backups are just hope and that most teams overstate their RTO by an order of magnitude. Help me design a disaster-recovery gameday that produces evidence, not vibes. I will provide: - The systems in scope (databases, object storage, stateful services, infra-as-code) - Stated RTO/RPO targets and how they were derived - Backup/restore mechanisms and where backups live - Whether prior restores have ever been performed end-to-end Do this: 1. **Pick a sharp scenario** — Choose one realistic disaster (region loss, ransomware-encrypted primary, accidental table drop, corrupted backup). Define the exact starting state and the success condition. 2. **Measure, don't assert** — Specify precisely what we will time: detection, decision, restore start, data restored, service healthy, traffic restored. The measured RTO is the only RTO that counts. 3. **Restore-from-zero test** — Force an actual restore from backup into a clean environment. Include verifying backup integrity, restore order for dependent data, and confirming application correctness, not just process-up. 4. **RPO truth** — Determine how much data was actually lost between last good backup and the disaster moment, and whether that matches the stated RPO. 5. **Safety rails** — Run against an isolated environment; define blast-radius controls so the gameday itself can't cause a real outage. Include an abort trigger and rollback. 6. **Findings to action** — Template for capturing where measured RTO exceeded target, which steps were undocumented, and which backups were unusable. Output: the scenario brief, a timed run-of-show with roles, the measurement sheet, the safety/abort plan, and a findings template that converts gaps into owned action items. Treat any step that 'should work but has never been tested' as a likely failure and design the gameday to expose it.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Game-Day Hypothesis and Abort-Criteria Design Prompt
Structure a chaos game-day around a falsifiable steady-state hypothesis with explicit blast-radius limits and abort conditions, so you learn from controlled failure without causing a real outage.
-
Multi-Region Failover Decision Playbook Prompt
Build a pre-decided playbook for whether and when to fail traffic to another region during an incident — including the cutover steps, the data-consistency traps, and the criteria for failing back.
-
Data-Loss and Data-Corruption Incident Runbook Prompt
Produce a careful, step-by-step runbook for handling a live data-loss or data-corruption incident — stopping the bleeding, preserving evidence, validating backups, and recovering without amplifying the damage.
-
GameDay Chaos Scenario Design Prompt
Design a safe, hypothesis-driven GameDay or chaos-engineering exercise grounded in your real incident history — with steady-state metrics, fault injections, blast-radius limits, abort criteria, and learning goals.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.