Building a GitLab CI/CD pipeline for network automation¶
A good network pipeline mirrors how a careful engineer works by hand — check the change is
well-formed, dry-run it, apply it, then confirm it worked — but does it automatically on
every commit. This article shows how to build that pipeline in GitLab CE: the four stages,
the .gitlab-ci.yml keywords that express them, and how to gate production safely.
In GitLab the pipeline is defined by a file named
.gitlab-ci.yml at the repository root, and its jobs
run on GitLab Runners. If you only need to read a
red pipeline, see Diagnosing a failing GitLab CI/CD pipeline;
here we construct one from scratch.
Four stages that mirror careful manual work¶
Structure a network-automation pipeline around four ordered stages. Each stage only starts once the previous one fully succeeds, so a bad change is stopped as early — and as cheaply — as possible.
| Stage | Purpose | Typical tools |
|---|---|---|
| build | Prove the change is syntactically valid before touching anything | yamllint, ansible-lint, ansible-playbook --syntax-check, terraform validate/fmt -check, render Jinja2 templates |
| prevalidate | Dry-run the change without touching production | ansible-playbook --check --diff, terraform plan, unit tests, apply to a virtual lab topology |
| deploy | Apply the change to the real network | ansible-playbook, terraform apply — gated and restricted to main |
| postvalidate | Confirm the live network matches intent | pyATS/Genie tests, operational-state checks over RESTCONF |
The order is the whole point: deploying before you validate defeats the pipeline. Build and prevalidate are your safety net; post-validation is the proof that reality now matches the desired state.
The building blocks of .gitlab-ci.yml¶
A handful of keywords express everything above:
stages— the ordered list of phases.- A job — a named block with a
stage, a containerimage, and ascriptof shell commands. before_script— setup commands run before every job'sscript(e.g. installing tools).rules— decide whether and when a job runs.rulesis the modern replacement for the legacyonly/exceptkeywords.artifacts— files a job produces that later stages consume (a rendered config, a test report).variablesand CI/CD variables — configuration and secrets. Store credentials as masked, protected project variables — never in the file.environment— names a deployment target such asstagingorproduction, which enables protected-environment controls.
A complete example pipeline¶
This four-stage .gitlab-ci.yml lints and syntax-checks a change, dry-runs it, gates the
production apply behind a manual click, and verifies the result:
stages:
- build
- prevalidate
- deploy
- postvalidate
variables:
ANSIBLE_FORCE_COLOR: "true"
default:
image: python:3.12
before_script:
- pip install "ansible-core==2.16.*" ansible-lint yamllint "pyats[full]"
- ansible-galaxy collection install cisco.ios
lint_and_syntax:
stage: build
script:
- yamllint .
- ansible-lint playbooks/
- ansible-playbook playbooks/deploy.yml --syntax-check
artifacts:
paths: [playbooks/]
expire_in: 1 hour
dry_run:
stage: prevalidate
script:
- ansible-playbook playbooks/deploy.yml --check --diff -i inventory/lab.ini
deploy_prod:
stage: deploy
environment:
name: production
rules:
- if: '$CI_COMMIT_BRANCH == "main"'
when: manual # a human must click to deploy
script:
- ansible-playbook playbooks/deploy.yml -i inventory/prod.ini
verify:
stage: postvalidate
script:
- pyats run job tests/postcheck_job.py --testbed-file testbed.yml
Every job pins its versions (python:3.12, ansible-core==2.16.*), so the pipeline is
reproducible instead of silently drifting when an upstream release lands. --check --diff
in the prevalidate stage is Ansible's dry run: it reports what would change without pushing
anything to a device.
Gating production safely¶
The deploy job combines three controls, and they matter more than any other part of the file:
- Restricted to
main— therulescondition$CI_COMMIT_BRANCH == "main"means the production job only appears on the main branch, never on feature branches. when: manual— GitLab waits for a person to press play. Insiderules, a manual job defaults toallow_failure: false, making it a blocking gate: nothing downstream runs until someone triggers it.- A protected
environment— restrict who is allowed to run deployments toproductionto a specific list of users, so the manual click is also an authorization check.
Secrets never live in the YAML. A committed .gitlab-ci.yml is visible to everyone with repo
access, and job logs can leak, so keep device credentials in
masked, protected CI/CD variables
and reference them at runtime as environment variables — never hard-code them in the file, a
commit message, or an artifact. Note that a protected variable is only injected on protected
branches and tags, which is usually exactly what you want for a production credential.
Runners, tags, and executors¶
The pipeline defines what to run; a runner is the agent that actually runs it. Each
runner uses an executor — most commonly Docker, where every job gets a fresh container from
your image:. Tags are how a job requests a specific
runner:
deploy_prod:
stage: deploy
tags: [network-lab] # only a runner with this tag can reach the devices
script:
- ansible-playbook playbooks/deploy.yml -i inventory/prod.ini
This is how a cloud-hosted GitLab drives devices on a private network: the runner lives inside the lab, and device-touching jobs are tagged for it. If a job hangs in pending, the usual cause is that no available runner carries the job's tags.
Passing data between jobs: artifacts vs cache¶
Jobs run in isolated environments, so two mechanisms move data between them, and confusing them causes flaky pipelines:
- Artifacts are outputs a job hands
forward — a rendered config, a JUnit test report, a validation log. They are guaranteed to
be available to later stages in the same pipeline and expire on a schedule you set. Use
artifacts(optionally narrowed withneedsordependencies) to pass a file the deploy stage depends on. - Cache speeds up repeated runs by reusing things like a Python virtualenv or downloaded packages across pipelines. It is best-effort and can be missing, so it must never hold anything a job depends on for correctness.
The rule of thumb: cache is an optimization that is safe to lose; artifacts are deliverables that must arrive.
Key takeaways¶
- Structure the pipeline as build → prevalidate → deploy → postvalidate — the same order a careful engineer follows, enforced on every commit.
- Build lints and syntax-checks; prevalidate dry-runs with
--check --diffandterraform plan; deploy applies the change; postvalidate proves the live network matches intent with pyATS and operational-state checks. - Gate production with three controls together: restrict to
main, requirewhen: manual(a blocking gate), and use a protectedenvironmentto control who may deploy. - Keep credentials in masked, protected CI/CD variables — never in the YAML, a commit, or an artifact.
- Pin every image, tool, and collection so the pipeline is reproducible, and route
device-touching jobs to a capable runner with
tags. - Pass required files forward as artifacts; use cache only for speed, never for correctness.