Skip to content
DevOps AI ToolKit
Home

GitLab CI/CD Troubleshooting & Pipeline Design

Debug pipelines, generate jobs, and review .gitlab-ci.yml with AI.

153 copy-paste prompts · 125 in-depth guides Jump to prompts Jump to guides

What actually breaks, and what to check first

GitLab CI failures fall into three groups that look identical in the job log and have nothing else in common: the pipeline did not run the job you expected, the runner could not start the job, or your commands failed inside it. Sorting a failure into the right group takes about thirty seconds and saves most of the debugging.

`rules` and `only/except` are the usual cause of the first group. A job that "did not run" almost never failed — it was never created, because no rule matched. The pipeline view shows what was created, never what was skipped, so the absence is easy to misread as a runner problem.

The second group is environmental: no runner with matching tags, a registry rejecting authentication, a Docker-in-Docker service that has not finished starting. These fail before your script executes, and the giveaway is that the log contains no output from your own commands.

Triage order

Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.

  1. Did the job get created at all?

    Compare the pipeline’s job list against .gitlab-ci.yml, then use CI Lint (Pipelines → Editor → Lint) with a merged result.

    A missing job is a `rules` problem, not a failure. Rules evaluate top to bottom and stop at the first match, so an early broad rule can shadow every later one.

  2. Did a runner pick it up?

    Read the top of the job log: "Running with gitlab-runner …" and the runner name.

    A job stuck in pending with no runner line means no runner matches its tags, or none is available. Tag mismatches are the most common cause and produce no error — just a job that waits.

  3. Did it fail before your script started?

    Look for the "Getting source from Git repository" and "$ <your first command>" markers.

    Failure before your first command is environmental — image pull, registry auth, cache restore or service startup. Failure after it is your build. This single distinction routes the whole investigation.

  4. Is it authentication or availability?

    grep -iE "401|403|denied|unauthorized|no space left|timed out" job.log

    401/403 against the registry usually means an expired `CI_JOB_TOKEN` scope or a wrong credential; "no space left" means the runner host needs cleanup; timeouts point at network or an overloaded runner.

  5. Is it reproducible outside CI?

    docker run --rm -it <same image> bash, then run the script steps by hand.

    Reproducing in the job’s own image separates "my command is wrong" from "the CI environment differs". Most "works on my machine" cases are a missing package that the runner image does not ship.

Diagnostic commands

Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.

gitlab-runner verify --delete
Changes state

Check registered runners against the GitLab instance and remove stale registrations.

How to read it: Run on the runner host. Stale registrations cause jobs to be assigned to runners that no longer exist, which presents as jobs hanging in pending.

gitlab-runner --debug run
Read-only

Run the runner in the foreground with verbose logging.

How to read it: Shows the job acquisition loop and executor errors that the job log never displays. The right tool when jobs are not being picked up at all.

curl --header "PRIVATE-TOKEN: $TOKEN" "https://gitlab.example.com/api/v4/projects/:id/pipelines/:pid/jobs"
Read-only

List every job in a pipeline with its status, via the API.

How to read it: Shows jobs that the UI collapses or hides, and confirms whether a job was created at all — the fastest way to prove a rules problem.

docker system df && docker system prune -af --filter "until=72h"
Destructive

Inspect and reclaim disk on a runner host.

How to read it: "No space left on device" mid-pipeline is nearly always accumulated build layers. Prune with a time filter; an unfiltered prune destroys the layer cache and slows every subsequent build.

git log --oneline -1 && git status --porcelain
Read-only

Confirm which commit the job actually checked out.

How to read it: Catches shallow-clone and detached-HEAD surprises, particularly in merge-request pipelines where the checked-out ref is a merge result rather than your branch.

Failure modes

These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.

The job never appears in the pipeline.

Cause:
No `rules` clause matched, so the job was never created.
Fix:
Rules stop at the first match. Reorder from most specific to most general, and use CI Lint with merged results to see what the final configuration evaluates to after includes.

Job sits in pending forever.

Cause:
No runner matches the job’s tags, or every matching runner is busy or offline.
Fix:
Compare the job’s tags against the runner’s tags exactly — they are case-sensitive and must all be present. A job with a tag no runner has will wait until it times out.

401 Unauthorized pulling from the container registry.

Cause:
The job token lacks scope for the target project, or a stored credential expired.
Fix:
`CI_JOB_TOKEN` is scoped to its own project by default; cross-project pulls need the allowlist configured. For external registries use a deploy token rather than a personal one, which breaks when the person leaves.

Full walkthrough →

Job fails with "no space left on device".

Cause:
The runner host has filled with accumulated images, layers and caches.
Fix:
Prune with a time filter and add a scheduled cleanup. Shared runners without periodic pruning fill predictably, so this recurs until it is automated.

Full walkthrough →

Docker-in-Docker commands fail with connection refused.

Cause:
The `docker:dind` service has not finished starting when the script begins.
Fix:
Wait for the daemon before the first Docker command, and set `DOCKER_HOST`/TLS variables consistently. The failure is a race, which is why it is intermittent and reruns sometimes "fix" it.

Cache never restores; every build starts cold.

Cause:
The cache key varies per job, or the cache path is outside the build directory.
Fix:
Key on the lockfile so the cache invalidates only when dependencies change, and confirm the path is inside the project — runners cannot cache paths outside it.

Common mistakes

  • Ordering `rules` from general to specific. The first match wins, so a broad early rule makes every later one unreachable.
  • Using a personal access token for CI. It stops working when that person’s access changes, usually at the worst moment.
  • Running an unfiltered `docker system prune -af` on a shared runner, which destroys the layer cache for every project on that host.
  • Assuming merge-request pipelines check out your branch. They check out a merge result, which is why some failures cannot be reproduced locally.
  • Setting `retry` on jobs to hide flakiness. It converts a visible intermittent failure into an invisible one that costs pipeline minutes.
  • Echoing variables to debug them. Masked variables leak once concatenated or base64-encoded, and CI logs are widely readable.

Frequently asked questions

Why did my GitLab job not run?

Because it was almost certainly never created. A job whose `rules` did not match is not skipped or failed — it simply does not exist in the pipeline, and the pipeline view only shows jobs that were created, so the absence looks like a bug. Use CI Lint with merged results to see what your configuration evaluates to after all `include` directives are resolved.

My job is stuck in pending. What is wrong?

No runner matches it. Tags are exact and case-sensitive, and a job requiring a tag that no runner carries will wait until it times out with no error message. Compare the job’s tag list against the runner’s and check the runner is online; those two cover nearly every case.

How do I tell whether a failure is my build or the CI environment?

Find the line where your first command runs. Anything failing before it — image pull, registry authentication, cache restore, service startup — is environmental. Anything after it is your build. It is a fast check and it routes the entire investigation correctly.

Why does my pipeline pass locally but fail in CI?

Usually because the runner image differs from your machine, or because a merge-request pipeline checks out a merge result rather than your branch. Reproduce by running the job’s own image with `docker run` and executing the script steps by hand; if it passes there, the difference is the ref that was checked out.

What is the safest way to handle credentials in GitLab CI?

Use masked, protected CI/CD variables, and prefer project or group deploy tokens over personal access tokens so access does not depend on one person’s account. Never echo a variable to debug it: masking is defeated by any transformation such as base64 or string concatenation, and job logs are visible to everyone with reporter access.

Prompts

Guides

Recommended tools

Newsletter

Free: the DevOps AI Incident-Triage Cheat Sheet

Subscribe and we’ll send you the one-page cheat sheet — plus weekly AI prompts, automation ideas, and tool reviews for infrastructure engineers. One email a week. No spam, unsubscribe anytime.

  • AI Incident-Triage Cheat Sheet (PDF)
  • Access to 2,778 DevOps AI prompts
  • One practical workflow email per week