Auto-Scaling Cost vs Latency Tuning Prompt
Tune auto-scaling parameters to balance cost against latency and reliability, choosing the right metrics, thresholds, and cooldowns to avoid flapping and over-provisioning.
- Target user
- SRE and platform engineers optimizing scaling behavior and cloud spend
- Difficulty
- Advanced
- Tools
- Claude, Gemini
The prompt
You are a senior reliability and cost engineer who tunes auto-scaling for the right balance of latency, reliability, and spend. I will provide: - The workload profile (traffic shape, spikiness, warm-up time per instance) - The current scaling config (HPA/KEDA, ASG, or cloud autoscaler) and metrics used - Latency/SLO targets and the cost budget - Observed problems (flapping, slow scale-up, idle over-provisioning) Your job: 1. **Pick scaling signals** — choose between CPU, RPS, queue depth, p95 latency, or custom KEDA metrics, and explain why the current signal may be wrong. 2. **Set thresholds and targets** — recommend target utilization, scale-out/in thresholds, and min/max bounds tied to the SLO. 3. **Stabilize** — tune cooldowns, stabilization windows, and step/percent policies to stop flapping. 4. **Handle warm-up** — account for instance/pod warm-up and connection draining to avoid cold-start latency during scale-up. 5. **Cut cost** — propose scheduled scaling for predictable cycles, spot/preemptible usage, and scale-to-zero where safe. 6. **Predictive option** — assess whether predictive/scheduled scaling beats reactive for this traffic shape. 7. **Validate** — define a load test and the dashboards/alerts to confirm the new config holds the SLO. Output as: (a) the recommended scaling config, (b) the signal/threshold rationale, (c) a cost-vs-latency trade-off table, (d) a load-test and rollback plan. Roll out changes to min/max bounds gradually and keep the prior config ready to restore; never let cost-driven minimums drop below what the SLO requires during peak.
Run this prompt with AI
Test it, get an AI-improved version, or compare models — live in the Prompt Workspace. No copy-paste.
Related prompts
-
Self-Hosted Runner Autoscaling Automation Prompt
Design autoscaling for self-hosted GitHub Actions runners — webhook-driven scale-up, idle scale-down, and ephemeral runner lifecycle — so CI capacity tracks demand without leaving zombie runners or leaking credentials between jobs.
-
Auto-Scaling Policy Automation Prompt
Design data-driven auto-scaling policies for HPA, KEDA, or cloud ASGs — picking the right metrics, thresholds, stabilization windows, and guardrails to avoid flapping and runaway scale-up.
-
Automation Client-Side Rate Limiter Token Bucket Design Prompt
Design a client-side rate limiter for automation that calls external APIs, using a token-bucket to stay under provider quotas, absorb bursts, and coordinate limits across concurrent workers without tripping 429s.
-
Cross-Region Automation Failover Orchestration Design Prompt
Design the orchestration that fails automation control planes and scheduled jobs over to a secondary region, avoiding split-brain double-execution while guaranteeing critical jobs still run during a regional outage.
More Automation prompts & error guides
Browse every Automation prompt and troubleshooting guide in one place.
Reading prompts? Get all 500 in one free PDF
500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.
- 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
- Instant PDF download — yours free, forever
- Plus one practical AI-workflow email a week (no spam)
Single opt-in · unsubscribe anytime · no spam.