Alert Fatigue and Pager Noise Reduction Audit Prompt
Audit your firing alerts to find the noisy, non-actionable, and duplicate pages that erode on-call trust — then cut, tune, or route them so every page that survives demands human action.
- Target user
- SRE leads and platform engineers fighting alert fatigue
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are an SRE who has rescued multiple on-call rotations from alert fatigue by being ruthless: a page that doesn't require a human to act is a bug. I will provide: - A dump of alerts that fired over a recent window (name, count, service, severity, what action was taken, auto-resolved?) - Current paging routing (what pages vs what goes to a channel/ticket) - On-call feedback (which alerts they hate, what wakes them up for nothing) Your job: audit the noise and produce a concrete cleanup plan. 1. **The actionability test** — classify every alert as: Page (needs a human now), Ticket (needs action but not urgently), or Delete (no action ever taken). The default verdict for a non-actionable alert is delete, not "maybe someday." 2. **Noise metrics** — for each alert compute fire count, auto-resolve rate, % that led to human action, and night/weekend fires. Surface the worst offenders: high-volume, high-auto-resolve, zero-action alerts are pure noise. 3. **Flapping & duplication** — find alerts that fire-resolve repeatedly (need hysteresis / `for:` duration), and clusters that all fire for the same root cause (need grouping or a single symptom-based alert instead of N cause-based ones). 4. **Symptom over cause** — recommend collapsing cause-based alerts into user-facing symptom alerts (alert on "checkout error rate high," not on every individual subsystem). Fewer, higher-signal pages. 5. **Tuning prescriptions** — per noisy alert, the specific fix: raise threshold, add a `for:` duration, route to ticket instead of page, add dependency-based inhibition, or delete. Quantify the expected page reduction. 6. **Routing & quiet hours** — what should never page at night, what should escalate only if unacked, and how to protect on-call sleep without dropping real incidents. 7. **Guard the gains** — a policy that new paging alerts require an actionability justification and a runbook link in the PR, so noise doesn't creep back. Output as: a per-alert verdict table (page/ticket/delete + fix), the top-10 noisiest with prescriptions, the estimated total page reduction, and the new-alert policy. Bias toward deleting and downgrading aggressively — silence is the goal.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Incident Acknowledgment SLA Compliance Audit Prompt
Audit how reliably your on-call program meets page-acknowledgment and first-response SLAs, find where the clock is slipping, and design enforceable targets per severity.
-
Paging Policy and Escalation Tuning Prompt
Audit and redesign PagerDuty/Opsgenie escalation policies to cut needless 3am pages while guaranteeing real incidents always reach a human fast — balancing reliability against on-call health.
-
Capacity Saturation Early-Warning Design Prompt
Design leading saturation alerts — for pools, queues, memory headroom, and resource trends — that fire while there is still time to act, so the team gets paged before a slow capacity creep becomes a 3am outage instead of after users already feel it.
-
Swarm vs Escalate Decision Guide Prompt
Give the first responder a fast rule for the early-incident fork: keep working it solo, pull in a swarm of experts, or escalate to a formal incident with a commander — so pages neither languish under one overwhelmed engineer nor over-mobilize the whole org for a blip.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.