Skip to content
DevOps AI ToolKit
All guides
AI for Automation By James Joyner IV · · 13 min read

Migrate GitLab Runner Autoscaling Before May 2027: Practical Playbook

Migration playbook for GitLab Runner autoscaling: executor selection, key tuning rules, provider notes, and a cutover checklist by May 2027.

Migrate GitLab Runner Autoscaling Before May 2027: Practical Playbook

Use GitLab Runner Autoscaler, either the Docker Autoscaler or the Instance executor, and run it from persistent Runner Manager hosts rather than spot instances. If you’re still on Docker Machine, start planning your migration now. Tune IdleCount and concurrent first since they control most of your cost and latency tradeoff, and remember that the autoscaler plugins already cover AWS, GCP, and Azure. Kubernetes workloads get their own pre-warming tricks.


TL;DR:

  • Docker Machine support is ending in May 2027, so teams should prioritize migrating to Docker Autoscaler or Instance executor to avoid last-minute issues.
  • Use the Docker Autoscaler for Linux workloads needing containerization, and switch to the Instance executor when jobs require full OS access or targeting Windows or macOS agents.
  • Keep the Runner Manager on a persistent host with proper credentials, as it coordinates scaling and must not be on spot instances to prevent loss of visibility.
  • Tuning parameters like IdleCount, IdleTime, and MaxGrowthRate are crucial to balancing cost and queue management, with metrics guiding incremental adjustments.
  • Avoid conflicting cloud autoscalers by setting provider autoscaling to manual or disabled, and monitor key metrics such as queue depth and failed creation counts for reliable scaling.

Table of Contents

Configure the runner manager: host requirements and credentials

I’ve migrated enough runner fleets to say this plainly: your Runner Manager is infrastructure, not cattle. It holds the state that coordinates every scaling decision, so putting it on a spot or preemptible instance is asking for trouble. When the manager disappears mid-cycle, you lose visibility into which instances are idle, which are mid-job, and which need cleanup. GitLab’s Docker Machine configuration guidance is explicit that the manager host needs to be persistent, and I’d add that it should also get normal patching and monitoring like any other production box.

A few practical notes from setting these up:

  • Run the manager on a small, always-on instance. Ubuntu LTS or Amazon Linux both work fine for the GitLab Runner package.
  • Install GitLab Runner from the official repository, not a generic package manager, so you get timely security updates.
  • Prefer cloud-native credentials: IAM instance profiles on AWS, Workload Identity on GCP, managed identities on Azure. Credential files work but create a rotation headache you don’t need.
  • Register runners with clear tags that map to workload types (build, test, deploy), since tags are how you’ll later split capacity and debug queuing.
  • If you’re using a custom driver that needs SSH, lock down security groups to the manager’s IP range only.

One manager config can register multiple child runners, which is handy when you want separate IdleCount and limit settings per workload without standing up new hosts.

Autoscaling executors and the Docker Machine migration

Three executors cover almost every autoscaling scenario: Docker Autoscaler, Instance executor, and Kubernetes executor. Docker Autoscaler provisions cloud instances and runs jobs in Docker containers on them, which is the closest conceptual replacement for Docker Machine. Instance executor skips the container layer and gives jobs the whole host, which matters if you need macOS or Windows runners or if a job needs direct device access. The Instance executor docs confirm it supports Linux, macOS, and Windows, which Docker Autoscaler does not.

Here’s the part that should move migration up your backlog: Docker Machine is deprecated and scheduled for removal in GitLab 20.0, expected May 2027. That sounds far off until you remember how long infrastructure changes take to land through change advisory boards and staging environments.

A rough decision order:

  1. Default to Docker Autoscaler for standard Linux CI workloads where containers are sufficient.
  2. Use Instance executor when jobs need the full OS, GPU access, or you’re targeting Windows or macOS build agents.
  3. Use Kubernetes executor when your workloads already run inside a cluster and you want pod-level isolation instead of managing cloud instances directly.
  4. Migrate incrementally: inventory existing Docker Machine configs, install the relevant fleeting plugin, test in staging, then cut over tag by tag.

Pro Tip: Keep your old Docker Machine config file saved somewhere boring, like a Git tag, so rollback is a five-minute job instead of an afternoon of memory reconstruction.

For a deeper comparison of executor tradeoffs, see our guide to choosing between shell, Docker, and Kubernetes executors.

How the autoscaling algorithm and tuning knobs actually interact

This is where most teams either overspend or create a queue of pending jobs, and the fix is understanding how a handful of parameters talk to each other.

  • IdleCount sets how many idle machines the autoscaler tries to keep warm, ready to absorb a sudden burst of jobs.
  • IdleTime is how long an idle machine sits before it gets terminated, trading a few extra minutes of cost for faster pickup on the next job.
  • concurrent caps total parallel jobs across the whole runner, while limit caps the number of machines a specific runner config can create.
  • MaxGrowthRate throttles how many new machines can be created per cycle, which protects your cloud account from a sudden spike in API calls.
  • IdleScaleFactor lets the idle pool grow proportionally with current load instead of staying fixed, which is useful for spiky workloads.

capacity_per_instance and use_count matter too: if one instance can run multiple jobs, your effective concurrency is higher than your machine count suggests, and the autoscaler accounts for that when deciding whether to spin up more capacity. According to GitLab’s autoscale configuration documentation, IdleScaleFactor and a minimum idle count setting exist specifically to give you adaptive behavior instead of a single static number.

The real lesson here is that most queuing problems trace back to one misconfigured parameter, not a cloud outage. Start with a conservative IdleCount, watch your machine creation duration and queue depth, then raise IdleCount or MaxGrowthRate incrementally based on what the metrics tell you, not what feels safe.

Illustration of queue signals tuning runner capacity

Preemptive mode, available on the Instance executor, requests a replacement machine before the current job finishes, which shaves startup latency at the cost of occasionally over-provisioning by one instance. Non-preemptive mode is more conservative and usually the better default until you’ve got a few weeks of metrics to justify the tradeoff.

Provider notes: AWS, GCP, Azure, and Fargate

The single most common mistake across every cloud is letting the provider’s own autoscaler fight with GitLab Runner’s. Set your cloud-side autoscaling resource to manual or “do not autoscale” everywhere, because the Docker Autoscaler documentation is blunt about what happens otherwise: conflicting scaling commands and failed jobs.

  • AWS: Set your Auto Scaling Group’s scaling mode to none and let the fleeting plugin manage instance count directly. Spot instances cut cost but add interruption risk, and GitLab’s AWS autoscaling guide notes that spot price volatility can push you into API request limits during machine creation. Use scale-in protection on instances running active jobs, and disable AZRebalance if it’s fighting your instance counts.
  • GCP: Use a managed instance group configured to “do not autoscale,” and prefer Workload Identity over downloaded key files for the fleeting plugin’s permissions.
  • Azure: In your VM Scale Set, set overprovision to false. Azure’s default overprovisioning behavior creates temporary extra VMs that can conflict with how GitLab Runner tracks instance state, which shows up as jobs assigned to VMs that are about to disappear.
  • Fargate: Workable for Docker Autoscaler but with real constraints: you need a compatible task image, a working SSH entrypoint for the driver, and a decision on private versus public IP assignment. Fargate suits bursty, stateless workloads better than long-running or highly custom build environments.

For AWS-specific spot tuning with working examples, our Fleeting plugin and spot instance walkthrough covers the configuration in more depth than fits here.

Kubernetes executor tuning: pause pods and pre-warming

If your workloads already live in a cluster, Kubernetes executor is usually the better fit over instance-based autoscaling, mainly because pod startup is faster than booting a cloud instance from scratch. The trick to making it feel instant is pause pods: low-priority placeholder pods that sit idle and get preempted the moment a real job pod needs the space. GitLab Runner creates a PriorityClass called gitlab-runner-idle-capacity with priority -1 by default specifically to make this preemption automatic, according to the Kubernetes executor docs.

  • Keep a small dedicated node pool for CI workloads so pause pods don’t compete with application pods for scheduling.
  • Pre-pull common build images onto that node pool to cut the biggest single source of job startup delay.
  • Set realistic resource requests and limits per job type; over-requesting wastes pause pod capacity, under-requesting causes throttling mid-build.

Pro Tip: If your builds are consistently slow to start despite pause pods, check image pull time first. It’s the most common hidden bottleneck.

Our Kubernetes executor tuning guide walks through resource sizing and caching strategies in more detail.

Fault tolerance and the metrics that actually matter

Run at least two Runner Managers sharing identical tags from day one. A single manager is a single point of failure for your entire CI pipeline, and the fleet scaling documentation recommends this pairing specifically for fault tolerance, not as an optional extra.

Once you have redundancy, watch these metrics before you touch any tuning parameter:

  • Machine creation duration tells you whether your cloud provider is becoming a bottleneck.
  • Failed creation counts flag credential problems, quota limits, or capacity shortages before they become outages.
  • Queue depth shows whether jobs are waiting on available capacity.
  • Idle instance count tells you if you’re overpaying for unused capacity.
  • Job saturation shows how close you are to your concurrent ceiling.

Dashboards built around the autoscaling algorithm, a queuing overview, and creation duration histograms give you the full picture in one glance, and the same fleet scaling guide stresses using these dashboards to set realistic IdleTime and MaxGrowthRate values rather than guessing. These same metrics double as your incident runbook: a spike in failed creations points at IAM or quota issues, while rising queue depth with low failed creations usually means your IdleCount is set too conservatively.

Troubleshooting pending jobs and common autoscaling errors

Most pending-job incidents fall into a handful of patterns.

  1. “Unable to acquire instance” errors usually trace back to IAM permissions or hitting a cloud quota. Check the fleeting plugin logs first.
  2. SSH tunnel EOF errors often mean a security group or network ACL is blocking the manager from reaching the new instance.
  3. VMSS overprovisioning conflicts on Azure show up as jobs assigned to instances that vanish mid-job. Confirm overprovision is set to false.
  4. Jobs pending while the runner shows online is frequently a misconfigured idle capacity setting rather than a real outage. According to GitLab’s support documentation on excessive queuing, over-configuring idle capacity can actually prevent runners from requesting new jobs at all.

Enable debug logs on the manager before escalating. Nine times out of ten, the answer is sitting in there.

Migration checklist and playbook

  1. Inventory every Docker Machine config and map each runner tag to the workload it serves.
  2. Provision persistent Runner Manager hosts and install the fleeting plugin for each target cloud.
  3. Switch credentials to IAM roles or Workload Identity, and tighten permissions to only what the plugin needs.
  4. Test in staging with capacity_per_instance=1 and a low max_instances ceiling before touching production traffic.
  5. Keep your original config saved and a staging-runner toggle ready, then monitor job success rates closely during cutover.

Pro Tip: Migrate one runner tag at a time instead of flipping the whole fleet at once. A bad config on one workload shouldn’t take down every pipeline in the company.

Author perspective and practical trade-offs

Latency and cost pull in opposite directions, and the only way to find your actual balance point is watching real metrics, not assumptions. Adopt incrementally, one tag at a time, and bring in outside help once the IAM and permission work exceeds what your team can safely audit alone.

— James

How DevOps AI ToolKit helps with autoscaling migrations

Migrating off Docker Machine touches IAM policy, instance group configuration, and Kubernetes resource sizing all at once, which is a lot to get right on a deadline. We offer a Terraform / IaC Audit to catch misconfigured autoscaling resources before they cause conflicts, plus a Kubernetes Health Check and an Observability Review for teams that need dashboards built around the metrics covered above.

Devopsaitoolkit

  • Audits are offered with defined pricing.
  • Remediation checklists are provided following audits.
  • Consulting is available for teams seeking hands-on assistance during migration.

Full service details and current pricing are on our work with me page and our pricing page.

Authoritative documentation and high-value references

FAQ

What is GitLab Runner autoscaling?

GitLab Runner autoscaling automatically provisions and removes compute instances or pods to match CI/CD job demand, controlled by a persistent Runner Manager host. The GitLab Runner Autoscaler documentation describes it as the modern replacement for the deprecated Docker Machine executor.

Why are my GitLab runner jobs stuck pending?

Jobs commonly stay pending because of IAM or credential failures, a cloud quota limit, or an idle capacity setting that’s misconfigured. According to GitLab’s support guidance, over-tuned idle capacity settings are a frequent, overlooked cause.

When does Docker Machine stop working in GitLab?

Docker Machine executor support is scheduled for removal in GitLab 20.0, expected around May 2027. Teams should migrate to Docker Autoscaler or Instance executor well before that date to avoid a forced last-minute cutover.

Should I use Docker Autoscaler or the Instance executor?

Use Docker Autoscaler for standard containerized Linux workloads, since it’s the direct successor to Docker Machine. Choose the Instance executor when jobs need full host access or when you’re targeting macOS or Windows runners, which Docker Autoscaler does not support.

How do pause pods reduce Kubernetes job startup time?

Pause pods are low-priority placeholder pods that sit idle on pre-warmed nodes and get preempted instantly when a real job pod needs the capacity. GitLab Runner creates a PriorityClass named gitlab-runner-idle-capacity with priority -1 by default to make this preemption automatic, according to the Kubernetes executor docs.

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week
Free download · 368-page PDF

Get 500 Battle-Tested DevOps AI Prompts — Free

500 battle-tested, copy-paste AI prompts engineered by a senior systems engineer — every one with fill-in placeholders and safety/back-out notes. Drop your email and it's yours.

  • 500 prompts: Linux · Kubernetes · Terraform · OpenStack · GitLab · Docker · Monitoring · Incident Response
  • Instant PDF download — yours free, forever
  • Plus one practical AI-workflow email a week (no spam)

Single opt-in · unsubscribe anytime · no spam.