Selecting the right network automation approach¶
The previous articles each covered a tool. This one covers judgement: given a real requirement — technical and business — which approach do you reach for? You'll learn the three broad categories, the criteria that separate them, and a heuristic that resolves most real-world tickets quickly.
Three categories¶
Almost every network automation solution falls into one of three buckets:
- Infrastructure-as-code (IaC) frameworks — Ansible and Terraform. You get version-controlled, repeatable definitions with a moderate learning curve. Ansible is procedural and excels at pushing configuration to many CLI devices; Terraform is declarative and stateful and excels at lifecycle management on controller-based platforms. Both are open, widely skilled, and integrate cleanly with Git and CI/CD.
- Low-code / no-code platforms — Cisco Catalyst Center templates and workflows, the Meraki Dashboard, SD-WAN (vManage) policies, or Ansible Automation Platform with self-service surveys. These lower the barrier so less-specialized staff can run standardized changes through a UI. The trade-off is flexibility: you are limited to what the platform exposes.
- Custom applications — usually Python. These offer unlimited flexibility: encode any business logic, integrate any system, present any interface. The cost is that you build and maintain everything yourself — error handling, testing, security, and lifecycle — an ongoing investment someone has to fund.
| Category | Strength | Trade-off |
|---|---|---|
| IaC framework | Repeatable, version-controlled, Git/CI-CD native | Moderate learning curve |
| Low-code / no-code | Non-programmers run standard changes | Limited to what the platform exposes |
| Custom application | Unlimited flexibility, any integration | You build and maintain everything |
There is no universally correct answer. The skill is matching the approach to the constraints as written, not defending a personal favourite. This is also the consensus in current practice: the mature teams aren't debating Terraform versus Ansible — they master both and use each where it fits, often in the same Git repo and pipeline.
Decision criteria¶
Work through the requirement against these axes:
- Declarative vs. imperative. Do you want to declare an end state and detect drift (Terraform), or run an ordered set of operations (Ansible playbooks, a Python script)?
- State and lifecycle. Do you need to track, change, and tear down resources over time? That favours Terraform, whose state file is what makes drift detection and clean teardown possible.
- Scale and repetition. Many near-identical changes across many devices favour an IaC framework over hand-built scripts.
- Team skillset. YAML and HCL have a far lower barrier than maintaining a Python codebase; no-code suits non-programmers.
- Platform fit. Controller platforms (ACI, Meraki, Catalyst Center, SD-WAN) have first-class Terraform providers and REST APIs; brownfield CLI estates often favour Ansible or Netmiko.
- Maintainability and ecosystem. Prefer the option with the strongest community modules/providers for your platform — someone else already carries the maintenance.
- Business factors. Cost, time-to-value, compliance and audit needs, vendor support, and who operates it day to day.
A selection heuristic¶
Most decisions collapse to three questions asked in order:
- Is there a maintained module or provider that already does it? → Use the IaC framework.
- Are the operators non-programmers running standard changes? → Use low-code / no-code.
- Is the logic genuinely custom and unsupported by anything above? → Build a custom Python application — and still wrap it in Git and CI/CD.
On a real backlog, most tickets resolve at step 1: "there's already a module/provider — use the IaC framework." Custom Python is reserved for the handful of integrations nothing else covers, precisely because every custom app becomes something the team maintains forever.
Worked scenarios¶
| Requirement | Best fit | Why |
|---|---|---|
| Provision 200 branch sites on Meraki with identical policy | Terraform | Declarative lifecycle + first-class Meraki provider + drift detection |
| Push a one-off ACL update to 500 brownfield IOS XE switches tonight | Ansible | Agentless push, idempotent ios_acls resource module, fast to write |
| Let the NOC run a self-service "add a VLAN" job through a form | Low-code | Ansible Automation Platform survey or a Catalyst Center template |
| Reconcile IPAM, ticketing, and device config with bespoke business rules | Custom Python | No provider covers cross-system logic; consume multiple REST APIs |
Notice how the one-off CLI push lands on Ansible, not Terraform: introducing per-switch Terraform state for a single change adds bookkeeping you'd immediately throw away. And the 200-site Meraki job lands on Terraform rather than a Python loop over the API — the provider gives you idempotency and drift detection for free, whereas the script would reinvent both, badly.
Idempotency: the property that often decides¶
Beyond features, the property that most often settles the choice is idempotency — whether running the same automation twice leaves the system unchanged the second time.
Declarative tools (Terraform, Ansible resource modules) are idempotent by design: they describe a
desired end state and converge toward it, so a re-run is safe and reports changed=0. A raw script
that sends CLI commands is procedural — it does exactly what you coded, every time, whether or not
the change is already in place, and can double-apply or error.
That distinction matters operationally: idempotent convergence means you can run automation on a schedule to enforce state and correct drift without fear that re-running causes damage. When a requirement says "continuously enforce" or "self-heal," that points to a declarative, idempotent approach — not a one-shot script.
A quick way to map language to category:
- "desired state / enforce / converge / drift" → declarative (Terraform, resource modules)
- "run these exact steps, in this order" → procedural (a script, or classic command modules)
- "let non-coders trigger it" → low-code / UI
Anchor everything to a source of truth¶
Whatever you choose, the intended configuration should live in a source of truth — a Git repository, a NetBox/IPAM system, or a data file — rather than only on the devices. Automation then renders and enforces that intent.
Combined with Git and CI/CD, this is the GitOps model: the repository is authoritative, every change is reviewed and versioned, and the pipeline deploys it. The device running-config and an engineer's laptop are downstream, not the source of truth. This applies equally to Ansible, Terraform, and custom apps, and is exactly how teams keep audit records and hold back configuration drift at scale.
Key takeaways¶
- Choose by requirements, not preference. Identify the constraints — declarative vs. imperative, state/lifecycle, scale, skillset, platform — and pick the option that matches them.
- Terraform for declarative, stateful lifecycle (especially controllers); Ansible for agentless config push (especially brownfield CLI). They're complementary, not rivals.
- Low-code / no-code lowers the skill barrier for standard changes, at the cost of flexibility.
- Custom Python only when the logic is genuinely unsupported — it carries an ongoing maintenance cost you own forever.
- Prefer a maintained module/provider before writing code, and anchor intent in a source of truth deployed via GitOps, regardless of tool.
Sources: Scalr — Terraform + Ansible: when to use each and how to integrate, Red Hat — Ansible vs. Terraform, CodiLime — Ansible vs. Terraform in networks.