Skip to content

Building a GitLab CI/CD pipeline for network automation

A good network pipeline mirrors how a careful engineer works by hand — check the change is well-formed, dry-run it, apply it, then confirm it worked — but does it automatically on every commit. This article shows how to build that pipeline in GitLab CE: the four stages, the .gitlab-ci.yml keywords that express them, and how to gate production safely.

In GitLab the pipeline is defined by a file named .gitlab-ci.yml at the repository root, and its jobs run on GitLab Runners. If you only need to read a red pipeline, see Diagnosing a failing GitLab CI/CD pipeline; here we construct one from scratch.

Four stages that mirror careful manual work

Structure a network-automation pipeline around four ordered stages. Each stage only starts once the previous one fully succeeds, so a bad change is stopped as early — and as cheaply — as possible.

Stage Purpose Typical tools
build Prove the change is syntactically valid before touching anything yamllint, ansible-lint, ansible-playbook --syntax-check, terraform validate/fmt -check, render Jinja2 templates
prevalidate Dry-run the change without touching production ansible-playbook --check --diff, terraform plan, unit tests, apply to a virtual lab topology
deploy Apply the change to the real network ansible-playbook, terraform apply — gated and restricted to main
postvalidate Confirm the live network matches intent pyATS/Genie tests, operational-state checks over RESTCONF

The order is the whole point: deploying before you validate defeats the pipeline. Build and prevalidate are your safety net; post-validation is the proof that reality now matches the desired state.

The building blocks of .gitlab-ci.yml

A handful of keywords express everything above:

  • stages — the ordered list of phases.
  • A job — a named block with a stage, a container image, and a script of shell commands.
  • before_script — setup commands run before every job's script (e.g. installing tools).
  • rules — decide whether and when a job runs. rules is the modern replacement for the legacy only/except keywords.
  • artifacts — files a job produces that later stages consume (a rendered config, a test report).
  • variables and CI/CD variables — configuration and secrets. Store credentials as masked, protected project variables — never in the file.
  • environment — names a deployment target such as staging or production, which enables protected-environment controls.

A complete example pipeline

This four-stage .gitlab-ci.yml lints and syntax-checks a change, dry-runs it, gates the production apply behind a manual click, and verifies the result:

stages:
  - build
  - prevalidate
  - deploy
  - postvalidate

variables:
  ANSIBLE_FORCE_COLOR: "true"

default:
  image: python:3.12
  before_script:
    - pip install "ansible-core==2.16.*" ansible-lint yamllint "pyats[full]"
    - ansible-galaxy collection install cisco.ios

lint_and_syntax:
  stage: build
  script:
    - yamllint .
    - ansible-lint playbooks/
    - ansible-playbook playbooks/deploy.yml --syntax-check
  artifacts:
    paths: [playbooks/]
    expire_in: 1 hour

dry_run:
  stage: prevalidate
  script:
    - ansible-playbook playbooks/deploy.yml --check --diff -i inventory/lab.ini

deploy_prod:
  stage: deploy
  environment:
    name: production
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
      when: manual            # a human must click to deploy
  script:
    - ansible-playbook playbooks/deploy.yml -i inventory/prod.ini

verify:
  stage: postvalidate
  script:
    - pyats run job tests/postcheck_job.py --testbed-file testbed.yml

Every job pins its versions (python:3.12, ansible-core==2.16.*), so the pipeline is reproducible instead of silently drifting when an upstream release lands. --check --diff in the prevalidate stage is Ansible's dry run: it reports what would change without pushing anything to a device.

Gating production safely

The deploy job combines three controls, and they matter more than any other part of the file:

  • Restricted to main — the rules condition $CI_COMMIT_BRANCH == "main" means the production job only appears on the main branch, never on feature branches.
  • when: manual — GitLab waits for a person to press play. Inside rules, a manual job defaults to allow_failure: false, making it a blocking gate: nothing downstream runs until someone triggers it.
  • A protected environment — restrict who is allowed to run deployments to production to a specific list of users, so the manual click is also an authorization check.

Secrets never live in the YAML. A committed .gitlab-ci.yml is visible to everyone with repo access, and job logs can leak, so keep device credentials in masked, protected CI/CD variables and reference them at runtime as environment variables — never hard-code them in the file, a commit message, or an artifact. Note that a protected variable is only injected on protected branches and tags, which is usually exactly what you want for a production credential.

Runners, tags, and executors

The pipeline defines what to run; a runner is the agent that actually runs it. Each runner uses an executor — most commonly Docker, where every job gets a fresh container from your image:. Tags are how a job requests a specific runner:

deploy_prod:
  stage: deploy
  tags: [network-lab]        # only a runner with this tag can reach the devices
  script:
    - ansible-playbook playbooks/deploy.yml -i inventory/prod.ini

This is how a cloud-hosted GitLab drives devices on a private network: the runner lives inside the lab, and device-touching jobs are tagged for it. If a job hangs in pending, the usual cause is that no available runner carries the job's tags.

Passing data between jobs: artifacts vs cache

Jobs run in isolated environments, so two mechanisms move data between them, and confusing them causes flaky pipelines:

  • Artifacts are outputs a job hands forward — a rendered config, a JUnit test report, a validation log. They are guaranteed to be available to later stages in the same pipeline and expire on a schedule you set. Use artifacts (optionally narrowed with needs or dependencies) to pass a file the deploy stage depends on.
  • Cache speeds up repeated runs by reusing things like a Python virtualenv or downloaded packages across pipelines. It is best-effort and can be missing, so it must never hold anything a job depends on for correctness.

The rule of thumb: cache is an optimization that is safe to lose; artifacts are deliverables that must arrive.

Key takeaways

  • Structure the pipeline as build → prevalidate → deploy → postvalidate — the same order a careful engineer follows, enforced on every commit.
  • Build lints and syntax-checks; prevalidate dry-runs with --check --diff and terraform plan; deploy applies the change; postvalidate proves the live network matches intent with pyATS and operational-state checks.
  • Gate production with three controls together: restrict to main, require when: manual (a blocking gate), and use a protected environment to control who may deploy.
  • Keep credentials in masked, protected CI/CD variables — never in the YAML, a commit, or an artifact.
  • Pin every image, tool, and collection so the pipeline is reproducible, and route device-touching jobs to a capable runner with tags.
  • Pass required files forward as artifacts; use cache only for speed, never for correctness.