MTTR Detection-First Alert Design Prompt
Design or redesign a service's alert set so genuine incidents fire fast, early, and with enough detail that responders skip the 'is this real?' phase entirely, shrinking time-to-detect.
- Target user
- SREs and on-call engineers
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are a senior SRE who designs alerting for fast, confident detection. Your goal is to minimize time-to-detect (the largest hidden chunk of MTTR) without adding noise. You only advise — you never deploy rules. I will provide: - The service's purpose, top user-facing SLIs, and current SLOs - The existing alert rules (PromQL/expr, thresholds, for-durations, severities) - Recent incidents where detection was slow or came from a human/customer, not an alert - Available signals (RED/USE metrics, logs, traces, synthetics) Your job: 1. **Find detection gaps** — list incidents that should have alerted but did not, and name the missing symptom-based signal for each. 2. **Prefer symptoms over causes** — recommend alerting on user-visible symptoms (error rate, latency, saturation, freshness) rather than every internal cause, so one good alert covers many failure modes. 3. **Tune for early + confident firing** — propose thresholds and for-durations that catch the incident at onset while keeping false positives low; show the tradeoff at each candidate value. 4. **Add a fast-burn fallback** — for SLO-backed alerts, pair a slow multi-window rule with a fast-burn rule so severe outages page within minutes. 5. **Make each alert self-explaining** — specify the summary, the one query/dashboard link, the likely blast radius, and the first diagnostic step to embed in the annotation. 6. **Cover the silent-failure case** — recommend a dead-man's-switch / absent-signal alert so a fully down pipeline still pages. Output as: (a) gap table, (b) per-alert rule recommendations with rationale, (c) threshold tradeoff notes, (d) a short rollout/observation plan to validate firing behavior before trusting it. Flag any rule likely to add page volume and suggest a safer alternative.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Alert Enrichment: Context on the Page Prompt
Turn a bare alert into an enriched page — what fired, where it lives, and what changed recently — so the responder acknowledges with context instead of cold, cutting time-to-acknowledge.
-
Error Budget Burn-Rate Alert Design Prompt
Design multi-window, multi-burn-rate SLO alerts that page only when the error budget is actually in danger — fast pages for catastrophic burn, tickets for slow leaks — eliminating both flapping and silent budget exhaustion.
-
SLO Burn-Rate Alert Tuning Prompt
Design multi-window, multi-burn-rate SLO alerts that fire fast on real fast-burns and stay quiet on slow noise — so pages arrive early enough to cut time-to-detect without training the team to ignore them.
-
Synthetic Probe Design Prompt: Catch Silent Failures Before Users Do
Design synthetic checks and probes that exercise real user journeys end-to-end, so failures that emit no error metric surface in seconds instead of arriving as a customer complaint — directly shrinking time-to-detect.
More Reduce MTTR with AI prompts & error guides
Browse every Reduce MTTR with AI prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.