Guided troubleshooting paths
Stack Missions
Missions turn this site’s deepest troubleshooting content into structured journeys. Each one walks a real production failure path — from mental model to a simulated incident you run in the Troubleshooting Workspace. Progress is tracked as you go; sign in to keep it across devices.
- Flagship
OpenStack Production Troubleshooting
Trace a request from the load balancer to compute, storage, and networking — and diagnose the failures that page you at 3am. Built on this site’s deepest OpenStack coverage.
8 modules
-
Linux Production Troubleshooting
The host-level skills every other stack sits on: disk, filesystem, memory, processes, and the commands that expose what a box is actually doing.
4 modules
-
Docker Production Readiness
Take a container from "works on my machine" to production-safe: build errors, runtime failures, Compose, and a hardening pass with the auditor.
4 modules
-
Kubernetes Production Troubleshooting
The failures that actually page you: pods that won’t start, networking and ingress, storage, and a security pass on the cluster.
4 modules
-
Prometheus & Monitoring Operations
Turn raw metrics into alerts that mean something: PromQL, alert rules tied to symptoms, Grafana, and the observability stack.
4 modules
-
PostgreSQL Production Troubleshooting
The database failures that take an app down: connection exhaustion, locks, sequence/constraint errors, and recovery.
3 modules
More missions (Terraform, GitLab CI/CD) are in progress. Want one prioritized? Tell me →