Skip to content

Diagnosing a failing GitLab CI/CD pipeline

A red pipeline is rarely a mystery once you know where to look. This article covers how a GitLab CI/CD pipeline is structured, how to read a job log to the first error instead of the last, and how to classify almost any failure into one of a few root causes — each with a matching fix.

In GitLab, a pipeline is defined by a file named .gitlab-ci.yml at the repository root, and its jobs are executed by GitLab Runners. Every push runs the pipeline, so when linting, tests, or a deployment break, the evidence is already sitting in the job log waiting to be read.

How a pipeline is structured

A pipeline is a set of jobs grouped into ordered stages. Jobs in the same stage run in parallel; a stage starts only after the previous stage fully succeeds. Each job declares a container image, an optional before_script for setup, and a script of shell commands.

stages: [build, test, deploy]

lint:
  stage: build
  image: python:3.12
  script:
    - pip install ansible-lint
    - ansible-lint playbooks/

Two rules explain most of what you'll see. A job fails the moment any command in its script returns a non-zero exit code, and a failed stage halts the pipeline — later stages never run — unless the job is marked allow_failure: true. That is why the first red job matters: everything downstream is often just fallout.

Read the job log from the top

The instinct is to read the last line of a failed log. Resist it. GitLab prints the job's environment first, then each command it runs prefixed with $, then that command's output, and finally an exit code. The useful signal is the first $ command whose output looks wrong — the last error is frequently a downstream symptom.

Read the log in this order:

  • Top of the log — which runner and image picked up the job? A wrong or missing image explains "command not found" instantly.
  • First $ command that errors — this is usually the real cause; everything after it is a knock-on effect.
  • The exit code — non-zero fails the job. A script that must tolerate an error needs explicit handling (|| true, an if guard), not silence.

The three most common failures

Once you've found the first error, classify it. The overwhelming majority of pipeline failures fall into three buckets.

Symptom in the log Category Fix
command not found, exit code 127, or a Python ModuleNotFoundError Missing dependency Install and pin it in before_script, or bake it into a custom image
An explicit version-constraint error Incompatible versions Align the versions — upgrade one side or pin the other
A lint rule, pytest assertion, or post-check that failed Failed test Fix the code the test caught

Missing dependency

The job calls a tool or library that isn't in the runner image:

$ ansible-playbook deploy.yml
/bin/sh: ansible-playbook: not found
ERROR: Job failed: exit code 127

Exit code 127 is the shell's universal "command not found". Install the dependency in before_script and pin it so the pipeline stays reproducible:

deploy:
  image: python:3.12
  before_script:
    - pip install "ansible-core==2.16.*"   # install + pin
  script:
    - ansible-playbook deploy.yml

Incompatible versions

The components all exist, but their versions clash — a collection that needs a newer ansible-core, a library that dropped support for your Python, a Terraform provider that requires a newer Terraform:

ERROR: cisco.ios 5.3.0 requires ansible-core>=2.15.0,
but you have ansible-core 2.13.4

Nearly every "it worked yesterday" failure is an unpinned dependency that silently upgraded. Pin the tool, library, collection, and base image (python:3.12, ansible-core==2.16.*, hashicorp/terraform:1.7) so a build is bit-for-bit repeatable. The fix for an incompatibility is to align the versions: upgrade the constrained component or pin the other to a compatible release.

Failed test

Here the environment is fine but a test asserted something untrue — a yamllint rule, a pytest assertion, or a pyATS post-check that found the network state didn't match intent:

FAILED tests/test_configs.py::test_vlan_present
    assert 10 in configured_vlans
E   assert 10 in [20, 30, 99]
ERROR: Job failed: exit code 1

This is the pipeline working, not obstructing you. The fix is to correct the flagged code — never to delete the test or loosen the linter to force the job green.

When the job never really runs

Some failures aren't in your code at all; the job never got a fair chance to execute. Run through this short checklist:

  • Stuck in pending — no runner matched. Either no active runner is assigned to the project, or the job's tags don't match any runner's tags. GitLab even tells you: "This job is stuck because you don't have any active runners online with any of these tags." Fix the tags on the job or the runner, or bring a runner online.
  • Wrong image — command not found for a tool you expected means image: doesn't include it.
  • Blank CI/CD variable — a credential resolves to empty. The classic trap: the variable is marked protected, but the pipeline is running on an unprotected feature branch, so it never gets injected. Protected variables reach only protected branches and tags.
  • YAML error — the pipeline doesn't start at all because .gitlab-ci.yml is malformed. Catch this before pushing with the CI Lint tool (or glab ci lint locally), so you skip a slow round-trip through a runner.

A repeatable method

Turn "the pipeline is red" into a specific fix every time:

  1. Open the failed job from the pipeline view — start at the job that went red, not in the code.
  2. Read the log to the first error; capture the command and its exit code.
  3. Classify it: tool/module not found → missing dependency; a constraint message → incompatible versions; an assertion or lint failure → failed test; a pending job or blank variable → an environment problem.
  4. If unsure, reproduce locally by running the same command inside the same image (docker run --rm -it python:3.12 sh).
  5. Apply the matching fix, then let the pipeline re-run to confirm it's green.

Key takeaways

  • A job fails on the first non-zero exit code, and a failed stage halts everything after it — so read the log top-down and fix the earliest error.
  • Exit code 127 / command not found is a missing dependency; an explicit version-constraint message is a version clash; a failed assertion or lint rule is a real problem the test caught.
  • Most "it worked yesterday" breakage is an unpinned dependency — pin the tool, library, collection, and base image to keep builds reproducible.
  • A job stuck in pending or a blank secret is an environment issue: check runner tags, and remember that protected variables only reach protected branches.
  • Validate .gitlab-ci.yml with CI Lint before pushing to avoid failures that never even reach a runner.