Kubernetes KEDA Event-Driven Autoscaling Prompt
Scale Kubernetes workloads on real event sources — queue depth, Kafka lag, cron, Prometheus queries — with KEDA, including scale-to-zero, ScaledObject/ScaledJob design, and avoiding flapping or stuck consumers.
- Target user
- Engineers scaling queue/stream workers beyond CPU-based HPA
- Difficulty
- Intermediate
- Tools
- Claude, ChatGPT
The prompt
You are an SRE who scales async workers on the metric that actually matters — backlog — instead of CPU, and has tuned KEDA to scale to zero without losing work. Provide: - The workload (queue consumer, Kafka consumer, webhook processor, batch jobs) - The event source (SQS/RabbitMQ/Kafka/Redis/Prometheus/cron) and its auth - Latency/throughput goals and cost pressure (is scale-to-zero worth it?) - Whether work is idempotent and how a mid-process pod kill is handled Design the autoscaling: 1. **HPA vs KEDA** — explain that KEDA *creates and drives an HPA* from external triggers; pick KEDA when the right signal is a queue/lag/external metric, and keep plain HPA when CPU/memory genuinely tracks load. 2. **ScaledObject design** — choose the trigger(s) and tune `pollingInterval`, `cooldownPeriod`, `minReplicaCount`, `maxReplicaCount`, and the per-trigger `threshold` (e.g. messages-per-replica). For Kafka, scale on consumer-group lag and respect partition count as a real max. Show multi-trigger composition. 3. **Scale-to-zero, safely** — when `minReplicaCount: 0` is appropriate, the cold-start latency tradeoff, the activation threshold, and how to avoid killing a pod mid-message (graceful shutdown + visibility timeout / commit-after-process). 4. **ScaledJob for batch** — when discrete jobs beat long-running consumers (each message → a Job), with `maxReplicaCount`, parallelism, and completion semantics. 5. **Auth (TriggerAuthentication)** — wire trigger auth via workload identity / a referenced Secret rather than inline creds, scoped to the one queue. 6. **Anti-flap & observability** — cooldown vs HPA stabilization window, the metrics to watch, and how to debug "why isn't it scaling?" (KEDA operator logs, the generated HPA, the metric value KEDA reports). Output: (a) the ScaledObject (or ScaledJob) + TriggerAuthentication for my source, (b) a threshold/tuning table with rationale, (c) a scale-to-zero safety checklist, (d) a flapping-diagnosis runbook, (e) the HPA-vs-KEDA verdict for my workload.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Vertical Pod Autoscaler (VPA) Tuning Prompt
Roll out the Vertical Pod Autoscaler safely — recommendation-only mode, update policies, container resource bounds, and the HPA coexistence trap — to right-size requests without restart storms.
-
Kubernetes Karpenter NodePool & Disruption Budget Tuning Prompt
Design and tune Karpenter NodePool, EC2NodeClass, and disruption/consolidation policies so the cluster bin-packs aggressively without churning workloads or violating PDBs.
-
Kubernetes HPA Debugging Prompt
Diagnose HorizontalPodAutoscaler issues — flapping replicas, `unable to fetch metrics`, custom metrics adapter, behavior tuning, scale-from-zero patterns.
-
Helm Secrets + SOPS Encrypted Values Workflow Prompt
Design a GitOps-safe workflow for encrypting Helm values with the helm-secrets plugin and SOPS (age/KMS) — encrypted values in git, decryption at deploy time, key rotation, and CI wiring.
More Kubernetes & Helm prompts & error guides
Browse every Kubernetes & Helm prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.