Skip to content

Consuming REST APIs at scale

Calling one endpoint once is easy. Building automation that stays reliable against a real controller — over thousands of objects, a token that expires mid-run, and a server that pushes back when you go too fast — is a different skill. This article covers the pieces that make REST consumption robust at scale: authentication patterns, persistent tokens, pagination, rate limiting with backoff, error handling, and the efficiency levers (webhooks and conditional requests) that keep you inside your quota.

Modern Cisco platforms — Catalyst Center, Meraki, ISE, SD-WAN vManage, and Webex — are all driven by REST APIs, so the same techniques carry across every controller you touch.

The foundation: read the status code

A REST API exposes resources at URLs, manipulates them with HTTP verbs (GET/POST/PUT/PATCH/DELETE), and exchanges JSON. Every response carries a status code your code must check: 2xx succeeded, 4xx means your request was wrong, 5xx means the server failed. Everything below builds on reading those codes correctly rather than assuming success.

Recognize the authentication pattern

Different platforms authenticate differently, and each maps to a recognizable pattern:

Pattern Where you see it How it works
API key in header Meraki (X-Cisco-Meraki-API-Key) Send the key on every call; protect it
HTTP Basic Cisco ISE ERS Base64 user:pass on every call, over HTTPS
Token / Bearer Catalyst Center POST credentials → token → resend as X-Auth-Token
OAuth 2.0 Webex, cloud services Short-lived access token plus a refresh flow

Persistent authentication

Re-authenticating on every request is wasteful and can trip rate limits. Persistent authentication means obtaining a token once, caching it, reusing it until it nears expiry, and refreshing it proactively — or reactively when a 401 Unauthorized tells you it is no longer valid.

A reusable requests.Session is the right container: it persists your token header, reuses TCP connections (connection pooling), and applies default settings across calls. The Catalyst Center token lives for 60 minutes, so refreshing at ~50 minutes leaves comfortable headroom.

# dnac_client.py — token auth with caching, proactive refresh, and 401 recovery
import time
import requests

class DnacClient:
    def __init__(self, base, user, pwd):
        self.base = base
        self.user, self.pwd = user, pwd
        self.session = requests.Session()
        self.token = None
        self.token_time = 0.0

    def _login(self):
        url = self.base + "/dna/system/api/v1/auth/token"
        r = self.session.post(url, auth=(self.user, self.pwd), timeout=10)
        r.raise_for_status()
        self.token = r.json()["Token"]
        self.token_time = time.time()
        self.session.headers.update({"X-Auth-Token": self.token})

    def _ensure_token(self):
        # refresh if missing or older than ~50 minutes (token lifetime is 60)
        if not self.token or (time.time() - self.token_time) > 3000:
            self._login()

    def get(self, path):
        self._ensure_token()
        r = self.session.get(self.base + path, timeout=10)
        if r.status_code == 401:          # token expired early → refresh once and retry
            self._login()
            r = self.session.get(self.base + path, timeout=10)
        r.raise_for_status()
        return r.json()

Pagination: walk every page

APIs rarely return thousands of objects in one response. Pagination splits results into pages, and your code must walk every page or it will silently miss data — one of the most common bugs in API automation. Three styles dominate:

  • Offset / limit (Catalyst Center): repeat with offset=1&limit=500, then offset=501, until a short or empty page comes back.
  • Cursor based (Meraki): pass perPage and follow the startingAfter / endingBefore cursors.
  • Link header (RFC 8288): the response's Link header contains a rel="next" URL; follow it until no next exists.
# pagination_link_header.py — Meraki-style pagination via the Link header
import re

def get_all(session, url):
    items = []
    while url:
        r = session.get(url, params={"perPage": 1000}, timeout=30)
        r.raise_for_status()
        items.extend(r.json())
        # follow rel="next" if the server offered one, else stop
        link = r.headers.get("Link", "")
        m = re.search(r'<([^>]+)>;\s*rel="next"', link)
        url = m.group(1) if m else None
    return items

Do not try to "fix" pagination by requesting a giant limit — servers cap page size, so you would still miss records. Always loop until the page runs out.

Rate limiting and backoff

To protect themselves, APIs cap how many requests you may send. Meraki, for example, allows 10 requests per second per organization (and 100 per source IP). Exceed the limit and you receive 429 Too Many Requests, usually with a Retry-After header stating how long to wait. Correct automation honours Retry-After first and falls back to exponential backoff with jitter for transient server errors, instead of hammering a server that is already struggling.

# rate_limit_backoff.py — retry 429 and 5xx, respecting Retry-After
import time
import random

def request_with_retry(session, method, url, max_tries=5, **kwargs):
    for attempt in range(max_tries):
        r = session.request(method, url, timeout=30, **kwargs)
        if r.status_code == 429:
            wait = float(r.headers.get("Retry-After", 1))  # server tells you how long
            time.sleep(wait)
            continue
        if 500 <= r.status_code < 600:                      # transient server error
            time.sleep((2 ** attempt) + random.random())    # backoff + jitter
            continue
        r.raise_for_status()
        return r
    raise RuntimeError(f"Exhausted retries for {url}")

The jitter matters. If every client retries after exactly 2, 4, and 8 seconds, they all collide again in synchronized waves — the "thundering herd." A small random offset spreads the retries so the server can recover.

Error handling and idempotency

Robust API code treats failure as normal:

  • Set a timeout on every call so a hung server cannot freeze your program.
  • Check status codes explicitly. Retry 429 and 5xx; do not blindly retry other 4xx codes — the request itself is wrong, so fix it, don't repeat it.
  • Convert bad codes into exceptions with raise_for_status(), then log the failure with context — and never leak tokens or credentials into logs.

Retry only idempotent operations. GET, PUT, and DELETE yield the same end state when repeated, so they are safe. A naive retry of a non-idempotent POST can create duplicates; when you must retry one, use an idempotency key or first check whether the resource already exists.

Pull less: webhooks and conditional requests

Scaling is also about when you pull. Polling — asking "has anything changed?" on a timer — is simple and works everywhere, but most calls return "no change," wasting quota and adding latency equal to the poll interval. Webhooks invert the flow: you register an HTTPS endpoint and the platform calls you the instant something happens, so you react immediately with zero wasted requests. The trade-off is operational — you must run a reachable, signature-validating endpoint that returns quickly and does heavy work asynchronously. Choose webhooks when the requirement is "react immediately"; choose polling when there is no inbound connectivity or a simple periodic sync is enough.

When you do poll, avoid re-fetching data that hasn't changed. HTTP conditional requests use validators: the server returns an ETag (a version fingerprint) or Last-Modified header, and your next request sends it back as If-None-Match or If-Modified-Since. If nothing changed, the server replies 304 Not Modified with no body.

# revalidate a cached resource with its ETag
resp = session.get(url, headers={"If-None-Match": '"a1b2c3"'}, timeout=10)
if resp.status_code == 304:
    data = cache[url]                       # unchanged → reuse what we already have
else:
    data = resp.json()
    cache[url] = data
    etags[url] = resp.headers.get("ETag")   # remember the new validator

Many APIs do not count 304 responses against your rate-limit budget, so conditional requests both speed up your automation and stretch your quota.

Key takeaways

  • Read the status code first. 2xx succeeded, 4xx is your fault, 5xx is theirs.
  • Recognize the auth pattern: API-key header, HTTP Basic, token/Bearer (Catalyst Center /auth/token → X-Auth-Token), or OAuth 2.0.
  • Persist authentication: obtain a token once, cache it in a Session, reuse it, and refresh near expiry or on a 401.
  • Always paginate — follow offset/limit, cursors, or the Link header until there is no next page.
  • 429 means slow down: honour Retry-After, then use exponential backoff with jitter for 5xx.
  • Set a timeout on every request; retry only 429/5xx, and only idempotent methods (GET/PUT/DELETE).
  • Pull less: prefer webhooks over constant polling, and use ETag/If-None-Match conditional requests to skip unchanged data for free.

Sources: Cisco Catalyst Center — Authentication, Meraki Dashboard API — Rate Limit, RFC 8288 — Web Linking, MDN — HTTP conditional requests.