Diagnosing a failing GitLab CI/CD pipeline¶
A red pipeline is rarely a mystery once you know where to look. This article covers how a GitLab CI/CD pipeline is structured, how to read a job log to the first error instead of the last, and how to classify almost any failure into one of a few root causes — each with a matching fix.
In GitLab, a pipeline is defined by a file named .gitlab-ci.yml
at the repository root, and its jobs are executed by GitLab Runners.
Every push runs the pipeline, so when linting, tests, or a deployment break, the evidence
is already sitting in the job log waiting to be read.
How a pipeline is structured¶
A pipeline is a set of jobs grouped into ordered stages. Jobs in the same stage run
in parallel; a stage starts only after the previous stage fully succeeds. Each job declares
a container image, an optional before_script for setup, and a script of shell
commands.
stages: [build, test, deploy]
lint:
stage: build
image: python:3.12
script:
- pip install ansible-lint
- ansible-lint playbooks/
Two rules explain most of what you'll see. A job fails the moment any command in its
script returns a non-zero exit code, and a failed stage halts the pipeline — later
stages never run — unless the job is marked
allow_failure: true. That is why the
first red job matters: everything downstream is often just fallout.
Read the job log from the top¶
The instinct is to read the last line of a failed log. Resist it. GitLab prints the job's
environment first, then each command it runs prefixed with $, then that command's output,
and finally an exit code. The useful signal is the first $ command whose output looks
wrong — the last error is frequently a downstream symptom.
Read the log in this order:
- Top of the log — which runner and image picked up the job? A wrong or missing
imageexplains "command not found" instantly. - First
$command that errors — this is usually the real cause; everything after it is a knock-on effect. - The exit code — non-zero fails the job. A script that must tolerate an error needs
explicit handling (
|| true, anifguard), not silence.
The three most common failures¶
Once you've found the first error, classify it. The overwhelming majority of pipeline failures fall into three buckets.
| Symptom in the log | Category | Fix |
|---|---|---|
command not found, exit code 127, or a Python ModuleNotFoundError |
Missing dependency | Install and pin it in before_script, or bake it into a custom image |
| An explicit version-constraint error | Incompatible versions | Align the versions — upgrade one side or pin the other |
A lint rule, pytest assertion, or post-check that failed |
Failed test | Fix the code the test caught |
Missing dependency¶
The job calls a tool or library that isn't in the runner image:
$ ansible-playbook deploy.yml
/bin/sh: ansible-playbook: not found
ERROR: Job failed: exit code 127
Exit code 127 is the shell's universal "command not found". Install the dependency in
before_script and pin it so the pipeline stays reproducible:
deploy:
image: python:3.12
before_script:
- pip install "ansible-core==2.16.*" # install + pin
script:
- ansible-playbook deploy.yml
Incompatible versions¶
The components all exist, but their versions clash — a collection that needs a newer
ansible-core, a library that dropped support for your Python, a Terraform provider that
requires a newer Terraform:
ERROR: cisco.ios 5.3.0 requires ansible-core>=2.15.0,
but you have ansible-core 2.13.4
Nearly every "it worked yesterday" failure is an unpinned dependency that silently
upgraded. Pin the tool, library, collection, and base image (python:3.12,
ansible-core==2.16.*, hashicorp/terraform:1.7) so a build is bit-for-bit repeatable.
The fix for an incompatibility is to align the versions: upgrade the constrained component
or pin the other to a compatible release.
Failed test¶
Here the environment is fine but a test asserted something untrue — a yamllint rule, a
pytest assertion, or a pyATS post-check that found
the network state didn't match intent:
FAILED tests/test_configs.py::test_vlan_present
assert 10 in configured_vlans
E assert 10 in [20, 30, 99]
ERROR: Job failed: exit code 1
This is the pipeline working, not obstructing you. The fix is to correct the flagged code — never to delete the test or loosen the linter to force the job green.
When the job never really runs¶
Some failures aren't in your code at all; the job never got a fair chance to execute. Run through this short checklist:
- Stuck in
pending— no runner matched. Either no active runner is assigned to the project, or the job'stagsdon't match any runner's tags. GitLab even tells you: "This job is stuck because you don't have any active runners online with any of these tags." Fix the tags on the job or the runner, or bring a runner online. - Wrong image —
command not foundfor a tool you expected meansimage:doesn't include it. - Blank CI/CD variable — a credential resolves to empty. The classic trap: the variable is marked protected, but the pipeline is running on an unprotected feature branch, so it never gets injected. Protected variables reach only protected branches and tags.
- YAML error — the pipeline doesn't start at all because
.gitlab-ci.ymlis malformed. Catch this before pushing with the CI Lint tool (orglab ci lintlocally), so you skip a slow round-trip through a runner.
A repeatable method¶
Turn "the pipeline is red" into a specific fix every time:
- Open the failed job from the pipeline view — start at the job that went red, not in the code.
- Read the log to the first error; capture the command and its exit code.
- Classify it: tool/module not found → missing dependency; a constraint message →
incompatible versions; an assertion or lint failure → failed test; a
pendingjob or blank variable → an environment problem. - If unsure, reproduce locally by running the same command inside the same image
(
docker run --rm -it python:3.12 sh). - Apply the matching fix, then let the pipeline re-run to confirm it's green.
Key takeaways¶
- A job fails on the first non-zero exit code, and a failed stage halts everything after it — so read the log top-down and fix the earliest error.
- Exit code
127/command not foundis a missing dependency; an explicit version-constraint message is a version clash; a failed assertion or lint rule is a real problem the test caught. - Most "it worked yesterday" breakage is an unpinned dependency — pin the tool, library, collection, and base image to keep builds reproducible.
- A job stuck in
pendingor a blank secret is an environment issue: check runner tags, and remember that protected variables only reach protected branches. - Validate
.gitlab-ci.ymlwith CI Lint before pushing to avoid failures that never even reach a runner.