A logging strategy: syslog and webhooks¶
A network automation solution has to record what it does and what it sees — otherwise the first hard question in an incident ("what changed, and when?") has no answer. A logging strategy answers three questions: what to log, at what severity, and to which destination. This article walks through both destinations that matter in practice — syslog for the durable, searchable record and webhooks for real-time notification — plus the structured-logging habits that make either one useful at scale.
What to log, and at what severity¶
Log the three things you will actually go looking for later: events (a playbook started or finished), changes (what was modified, on which device, by whom), and errors (with enough context to diagnose). Attach a timestamp and a device identifier to every record, and never write a secret.
Severity is what lets you separate signal from noise. Syslog and most logging frameworks share the same eight levels defined in RFC 5424, numbered 0–7:
| Code | Level | Use it for |
|---|---|---|
| 0 | Emergency | System is unusable |
| 1 | Alert | Action must be taken immediately |
| 2 | Critical | A critical condition |
| 3 | Error | A failure — a device unreachable, a change rejected |
| 4 | Warning | Something recoverable that an operator should see |
| 5 | Notice | Normal but significant |
| 6 | Informational | Routine progress — "deploy started on 42 devices" |
| 7 | Debug | Developer detail |
The trap to avoid: lower number means higher severity. 0 (emergency) is the most severe and 7 (debug) is the most verbose — not the other way around. This ordering is the same on Cisco devices and in Python's logging module, which is what makes routing by severity work consistently end to end.
Syslog: the durable record¶
Syslog ships each message — tagged with a facility and severity — to a centralized server or SIEM. It is the right destination for durable, searchable, audit-grade records of everything the automation does.
On a Cisco IOS XE device, point logging at a collector and cap it by severity:
logging host 10.0.0.50
logging trap informational
logging origin-id hostname
service timestamps log datetime msec show-timezone
logging trap informational sends severity 6 and everything more severe (lower-numbered) — so informational through emergency, but not debug. Specifying any level always includes that level plus all lower numbers (IOS XE System Message Logs).
Your automation host logs to syslog too. Python's standard library ships straight there with SysLogHandler:
import logging
from logging.handlers import SysLogHandler
log = logging.getLogger("netauto")
log.setLevel(logging.INFO)
handler = SysLogHandler(address=("syslog.example.com", 514)) # UDP/514 by default
handler.setFormatter(logging.Formatter("netauto: %(levelname)s %(message)s"))
log.addHandler(handler)
log.info("playbook deploy.yml started on 42 devices")
log.error("device r3 unreachable — skipping")
Webhooks: real-time notification¶
A webhook is an HTTP POST your automation sends to a URL the moment an event occurs — the event-driven complement to syslog. Where syslog is for the complete durable record, a webhook is for immediate human notification (Slack, Teams, an incident tool) or for triggering downstream automation (ChatOps). You build a small payload and POST it to an endpoint you keep as a secret configuration value:
import os
import requests
def notify(text):
url = os.environ["WEBHOOK_URL"] # secret kept out of code
resp = requests.post(url, json={"text": text}, timeout=5)
resp.raise_for_status()
notify("Deployment complete: 42/42 devices succeeded")
Route by severity: log everything, page selectively¶
Syslog and webhooks are not an either/or. Most solutions do both: log everything to syslog for the record, and fire a webhook only for high-severity events so humans are paged for what matters, not for routine noise. Because logging levels are numeric, you can wire this directly into the logging tree — one handler per destination, each with its own threshold:
import logging
import os
import requests
from logging.handlers import SysLogHandler
class WebhookHandler(logging.Handler):
"""POST high-severity records to a chat or incident webhook."""
def emit(self, record):
try:
requests.post(
os.environ["WEBHOOK_URL"], # secret, not in code
json={"text": self.format(record)},
timeout=5,
)
except Exception:
self.handleError(record) # logging must never crash the app
log = logging.getLogger("netauto")
log.setLevel(logging.INFO)
syslog = SysLogHandler(address=("syslog.example.com", 514))
syslog.setLevel(logging.INFO) # everything, for the durable record
webhook = WebhookHandler()
webhook.setLevel(logging.ERROR) # only what should page a human
log.addHandler(syslog)
log.addHandler(webhook)
log.info("deploy.yml started on 42 devices") # → syslog only
log.error("device r3 unreachable — skipping") # → syslog *and* webhook
The setLevel on each handler does the routing: the INFO line lands in syslog only, while the ERROR reaches both. Wrapping emit in try/except and calling handleError matters — a flaky webhook endpoint should never take down the automation that is trying to report a problem.
Structured logs and correlation IDs¶
Free-form text is hard to search once you have more than one runner. Structured logging emits each record as JSON with consistent fields, so a collector can filter on any of them. Log a small set of audit fields on every change — who triggered it, what changed, on which device, when, and the result — plus a correlation ID that ties together every record for one run:
import json
import logging
class JsonFormatter(logging.Formatter):
def format(self, record):
return json.dumps({
"ts": self.formatTime(record),
"level": record.levelname,
"run_id": getattr(record, "run_id", None),
"device": getattr(record, "device", None),
"msg": record.getMessage(),
})
handler = logging.StreamHandler()
handler.setFormatter(JsonFormatter())
log = logging.getLogger("netauto")
log.addHandler(handler)
log.setLevel(logging.INFO)
log.info("vlan 10 added", extra={"run_id": "a1b2", "device": "r1"})
When one pipeline run touches 200 devices, a shared run_id lets you pull every log line for that run — across the automation host, the API gateway, and each device — into a single searchable timeline instead of a scattered mess. Inside your scripts, reach for logging over print: it gives you levels, timestamps, and central routing for free, and log.exception("config push failed") records the full traceback at ERROR automatically.
Two rules that keep logging trustworthy¶
- Logging is not alerting. Logging records what happened; alerting decides what a human must act on now. Paging on every INFO line trains people to ignore alerts (alert fatigue). Log everything, but route only actionable, high-severity events to a webhook that pages someone.
- Never log secrets. Logs are widely readable and long-retained, so a password or token written to a log is effectively leaked. Scrub credentials before logging, and keep webhook URLs and tokens in secret storage — not in the log body, not in code. Set retention deliberately, too: keep operational logs long enough to investigate incidents and meet any audit requirement, then rotate or archive so storage does not grow without bound.
Key takeaways¶
- A logging strategy decides what to log (events, changes, errors — with context), at what severity (0 emergency → 7 debug), and where.
- Lower severity number = more severe.
0is emergency,7is debug — on the device and in Python alike. - Syslog is the durable, centralized, audit-grade record (
logging trapon IOS XE,SysLogHandlerin Python). Webhooks deliver real-time event notification via HTTP POST. - Use both: log everything to syslog, POST only high-severity events to a webhook. Per-handler
setLeveldoes the routing. - Prefer structured JSON logs with a correlation ID so one run's records reassemble into a single timeline.
- Logging ≠ alerting, and never log secrets — keep webhook URLs and tokens in secret storage.
Sources: RFC 5424 — The Syslog Protocol,
Cisco IOS XE — Configuring System Message Logs,
Python logging.handlers.SysLogHandler,
Python Logging HOWTO.