What actually breaks, and what to check first
GitLab CI failures fall into three groups that look identical in the job log and have nothing else in common: the pipeline did not run the job you expected, the runner could not start the job, or your commands failed inside it. Sorting a failure into the right group takes about thirty seconds and saves most of the debugging.
`rules` and `only/except` are the usual cause of the first group. A job that "did not run" almost never failed — it was never created, because no rule matched. The pipeline view shows what was created, never what was skipped, so the absence is easy to misread as a runner problem.
The second group is environmental: no runner with matching tags, a registry rejecting authentication, a Docker-in-Docker service that has not finished starting. These fail before your script executes, and the giveaway is that the log contains no output from your own commands.
Triage order
Work down this list in order. Each step either finds the fault or rules out a whole class of cause — a negative result is progress, not a wasted step.
-
Did the job get created at all?
Compare the pipeline’s job list against .gitlab-ci.yml, then use CI Lint (Pipelines → Editor → Lint) with a merged result.A missing job is a `rules` problem, not a failure. Rules evaluate top to bottom and stop at the first match, so an early broad rule can shadow every later one.
-
Did a runner pick it up?
Read the top of the job log: "Running with gitlab-runner …" and the runner name.A job stuck in pending with no runner line means no runner matches its tags, or none is available. Tag mismatches are the most common cause and produce no error — just a job that waits.
-
Did it fail before your script started?
Look for the "Getting source from Git repository" and "$ <your first command>" markers.Failure before your first command is environmental — image pull, registry auth, cache restore or service startup. Failure after it is your build. This single distinction routes the whole investigation.
-
Is it authentication or availability?
grep -iE "401|403|denied|unauthorized|no space left|timed out" job.log401/403 against the registry usually means an expired `CI_JOB_TOKEN` scope or a wrong credential; "no space left" means the runner host needs cleanup; timeouts point at network or an overloaded runner.
-
Is it reproducible outside CI?
docker run --rm -it <same image> bash, then run the script steps by hand.Reproducing in the job’s own image separates "my command is wrong" from "the CI environment differs". Most "works on my machine" cases are a missing package that the runner image does not ship.
Diagnostic commands
Every command is labelled by what it can do to the system. Read-only commands are safe to run during an incident; the others are not, and are marked accordingly.
gitlab-runner verify --delete Check registered runners against the GitLab instance and remove stale registrations.
How to read it: Run on the runner host. Stale registrations cause jobs to be assigned to runners that no longer exist, which presents as jobs hanging in pending.
gitlab-runner --debug run Run the runner in the foreground with verbose logging.
How to read it: Shows the job acquisition loop and executor errors that the job log never displays. The right tool when jobs are not being picked up at all.
curl --header "PRIVATE-TOKEN: $TOKEN" "https://gitlab.example.com/api/v4/projects/:id/pipelines/:pid/jobs" List every job in a pipeline with its status, via the API.
How to read it: Shows jobs that the UI collapses or hides, and confirms whether a job was created at all — the fastest way to prove a rules problem.
docker system df && docker system prune -af --filter "until=72h" Inspect and reclaim disk on a runner host.
How to read it: "No space left on device" mid-pipeline is nearly always accumulated build layers. Prune with a time filter; an unfiltered prune destroys the layer cache and slows every subsequent build.
git log --oneline -1 && git status --porcelain Confirm which commit the job actually checked out.
How to read it: Catches shallow-clone and detached-HEAD surprises, particularly in merge-request pipelines where the checked-out ref is a merge result rather than your branch.
Failure modes
These are distinct problems, not variations of one. Matching the symptom to the right cause is most of the work.
The job never appears in the pipeline.
- Cause:
- No `rules` clause matched, so the job was never created.
- Fix:
- Rules stop at the first match. Reorder from most specific to most general, and use CI Lint with merged results to see what the final configuration evaluates to after includes.
Job sits in pending forever.
- Cause:
- No runner matches the job’s tags, or every matching runner is busy or offline.
- Fix:
- Compare the job’s tags against the runner’s tags exactly — they are case-sensitive and must all be present. A job with a tag no runner has will wait until it times out.
401 Unauthorized pulling from the container registry.
- Cause:
- The job token lacks scope for the target project, or a stored credential expired.
- Fix:
- `CI_JOB_TOKEN` is scoped to its own project by default; cross-project pulls need the allowlist configured. For external registries use a deploy token rather than a personal one, which breaks when the person leaves.
Job fails with "no space left on device".
- Cause:
- The runner host has filled with accumulated images, layers and caches.
- Fix:
- Prune with a time filter and add a scheduled cleanup. Shared runners without periodic pruning fill predictably, so this recurs until it is automated.
Docker-in-Docker commands fail with connection refused.
- Cause:
- The `docker:dind` service has not finished starting when the script begins.
- Fix:
- Wait for the daemon before the first Docker command, and set `DOCKER_HOST`/TLS variables consistently. The failure is a race, which is why it is intermittent and reruns sometimes "fix" it.
Cache never restores; every build starts cold.
- Cause:
- The cache key varies per job, or the cache path is outside the build directory.
- Fix:
- Key on the lockfile so the cache invalidates only when dependencies change, and confirm the path is inside the project — runners cannot cache paths outside it.
Common mistakes
- Ordering `rules` from general to specific. The first match wins, so a broad early rule makes every later one unreachable.
- Using a personal access token for CI. It stops working when that person’s access changes, usually at the worst moment.
- Running an unfiltered `docker system prune -af` on a shared runner, which destroys the layer cache for every project on that host.
- Assuming merge-request pipelines check out your branch. They check out a merge result, which is why some failures cannot be reproduced locally.
- Setting `retry` on jobs to hide flakiness. It converts a visible intermittent failure into an invisible one that costs pipeline minutes.
- Echoing variables to debug them. Masked variables leak once concatenated or base64-encoded, and CI logs are widely readable.
Frequently asked questions
Why did my GitLab job not run?
Because it was almost certainly never created. A job whose `rules` did not match is not skipped or failed — it simply does not exist in the pipeline, and the pipeline view only shows jobs that were created, so the absence looks like a bug. Use CI Lint with merged results to see what your configuration evaluates to after all `include` directives are resolved.
My job is stuck in pending. What is wrong?
No runner matches it. Tags are exact and case-sensitive, and a job requiring a tag that no runner carries will wait until it times out with no error message. Compare the job’s tag list against the runner’s and check the runner is online; those two cover nearly every case.
How do I tell whether a failure is my build or the CI environment?
Find the line where your first command runs. Anything failing before it — image pull, registry authentication, cache restore, service startup — is environmental. Anything after it is your build. It is a fast check and it routes the entire investigation correctly.
Why does my pipeline pass locally but fail in CI?
Usually because the runner image differs from your machine, or because a merge-request pipeline checks out a merge result rather than your branch. Reproduce by running the job’s own image with `docker run` and executing the script steps by hand; if it passes there, the difference is the ref that was checked out.
What is the safest way to handle credentials in GitLab CI?
Use masked, protected CI/CD variables, and prefer project or group deploy tokens over personal access tokens so access does not depend on one person’s account. Never echo a variable to debug it: masking is defeated by any transformation such as base64 or string concatenation, and job logs are visible to everyone with reporter access.