Observability Gap Analysis From Incidents Prompt
Mine recent incidents to find where missing logs, metrics, or traces slowed detection and diagnosis, then prioritize the observability investments that would have shortened them most.
- Target user
- Observability and SRE teams prioritizing telemetry investments
- Difficulty
- Advanced
- Tools
- Claude, ChatGPT
The prompt
You are a staff observability engineer who treats every incident as evidence of where the system is blind. You know that "we eventually figured it out" usually hides expensive telemetry gaps. I will provide: - A set of recent postmortems or incident timelines - Current telemetry coverage (metrics, logs, traces, synthetic, RUM) per service - Detection sources (which signal caught each incident, or "customer reported") - Diagnosis notes (what engineers had to guess, query manually, or add mid-incident) Perform an observability gap analysis. Work through these steps: 1. **Score detection** — for each incident, classify how it was detected (proactive alert, dashboard, synthetic, support ticket, customer) and estimate the detection delay attributable to missing or noisy signals. 2. **Score diagnosis** — identify moments in each timeline where engineers stalled because a signal was missing, wrong-grained, unsampled, retention-expired, or uncorrelated across signals. Tag each stall with the missing telemetry. 3. **Aggregate the gaps** — cluster the per-incident gaps into recurring themes (e.g., no trace propagation across service X, no saturation metric on the queue, logs missing request IDs, dashboards lacking per-tenant breakdown). 4. **Estimate impact** — for each gap theme, estimate the detection or diagnosis time it would have saved across the incident set, and how many incidents it touches. 5. **Cost the fixes** — rough effort and ongoing cost (cardinality, storage, sampling) for each instrumentation change, and call out where more telemetry would add noise rather than signal. 6. **Prioritize** — rank gap fixes by saved-time-per-effort, and propose alerting / SLO changes that turn newly added signals into proactive detection. Output: (a) a per-incident detection + diagnosis scorecard, (b) clustered gap themes with affected incidents, (c) an impact-vs-cost prioritization table, (d) the top 5 instrumentation changes with concrete metric/log/trace specs, (e) the alerts or SLOs to add on top. Distinguish evidence-backed gaps from speculation, and flag where you would need more data.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Synthetic Monitoring for Faster Incident Detection Prompt
Design synthetic checks and journey probes that catch incidents before customers report them — closing the gap between failure and detection (the 'time-to-detect' phase of MTTR).
-
SLO Incident Dashboard Spec Generator Prompt
Specify a single incident-response dashboard for a service — the SLIs, burn-rate panels, saturation signals, and dependency health a responder actually needs at 3am — laid out so the first-on-call answers 'is it us, and how bad' in under a minute.
-
First-Alert Triage & Hypothesis Ranking Prompt
Take a freshly fired alert plus a snapshot of metrics, logs, and recent changes, and produce a ranked list of failure hypotheses with the cheapest next diagnostic step for each — without taking any action on the system.
-
Live Incident Hypothesis Tracker Prompt
Keep a live incident's debugging organized — track every hypothesis, the evidence for and against it, what's been ruled out, and the next highest-value experiment — so the team converges on the cause instead of chasing in circles.
More Incident Response prompts & error guides
Browse every Incident Response prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.