Network automation
Collect device state, parse CLI output into data and push configuration changes safely with Netmiko, NAPALM, Nornir and Ansible.
On this page
Cheatsheet#
| Task | Tool or snippet |
|---|---|
| SSH to one device, run a command | ConnectHandler(**dev).send_command("show ip interface brief") (Netmiko) |
| Same, parsed into a list of dicts | send_command(cmd, use_textfsm=True) |
| Command with an interactive prompt | send_command(cmd, expect_string=r"confirm") |
| Slow command | send_command(cmd, read_timeout=120) |
| Sync or asyncio SSH with a smaller footprint | scrapli (Scrapli, AsyncScrapli) |
| Vendor-neutral getters and config replace/merge | napalm.get_network_driver("ios") |
| Many devices from an inventory, threaded | Nornir with nornir_netmiko or nornir_napalm |
| Declarative, idempotent changes with check mode | Ansible cisco.ios.* modules over network_cli |
| Parse CLI text offline | ntc_templates.parse.parse_output (TextFSM) |
| Structured data from the device | NETCONF, RESTCONF, gNMI; | json on NX-OS |
| Diff before apply | NAPALM compare_config(); Ansible --check --diff |
| Confirmed commit | NAPALM commit_config(revert_in=300) then confirm_commit() |
| Bulk reachability | fping -a -g 192.0.2.0/24 |
Choosing the interface#
| Interface | When |
|---|---|
| SSH and CLI scraping (Netmiko, scrapli) | Always available. The fallback when nothing structured exists |
| NETCONF / RESTCONF (YANG models) | Structured config and state; candidate datastore and validate-then-commit where the platform supports it |
| gNMI | Streaming telemetry and config on current platforms |
| Vendor or controller REST API | Controller-managed fabrics (ACI, Meraki, Catalyst Center) where the controller owns the config |
Prefer a structured interface when the platform has one. CLI output is text for humans, its format changes between software releases, and every parser is a maintenance item. Screen scraping works by sending a command, reading until the prompt pattern reappears and returning the text in between; most failures come from that prompt detection.
Connecting and collecting with Netmiko#
Netmiko (4.x) wraps Paramiko SSH with per-platform prompt handling. On connect it runs the platform’s session preparation, which for IOS includes terminal length 0, so you do not disable paging yourself.
import os
from netmiko import ConnectHandler
device = {
"device_type": "cisco_ios", # cisco_xe is an alias; cisco_nxos, arista_eos, juniper_junos...
"host": "sw-01.example.com",
"username": os.environ["NET_USER"],
"password": os.environ["NET_PASS"],
"secret": os.environ.get("NET_ENABLE", ""),
"conn_timeout": 10, # TCP connect timeout, seconds
}
with ConnectHandler(**device) as conn:
conn.enable() # needs "secret" if not already privileged
raw = conn.send_command("show ip interface brief") # str
rows = conn.send_command("show ip interface brief", use_textfsm=True) # list[dict], or str if no template matchedsend_command waits until it sees the prompt or expect_string, up to read_timeout seconds (default 10), then raises ReadTimeout. delay_factor and max_loops from Netmiko 3 are deprecated. send_command_timing instead returns when output stops arriving, which suits commands whose prompt is unpredictable but can return early on slow devices.
out = conn.send_command("copy running-config startup-config", expect_string=r"\[startup-config\]\?")
out += conn.send_command("\n", expect_string=r"#")More of the Netmiko surface that comes up in real jobs:
conn.send_config_from_file("acl.cfg") # config mode, one line at a time, exits
conn.send_config_set(cmds, cmd_verify=False) # faster on slow devices; skips echo checks, so errors are not caught
conn.send_multiline(["copy tftp: flash:", "192.0.2.40", "c9300.bin", "\n"]) # answer a chain of prompts
conn.send_command("show run", use_genie=True) # pyATS/Genie parser instead of TextFSM (pip install pyats genie)
conn.send_command("show ip route", use_ttp=True, ttp_template="route.ttp") # TTP templates
conn.write_channel("show version\n"); time.sleep(1); conn.read_channel() # raw, for pathological prompts
conn.find_prompt() # current prompt; changes when hostname changes
conn.disconnect()
from netmiko import file_transfer
file_transfer(conn, source_file="c9300-universalk9.17.12.04.SPA.bin", dest_file="c9300.bin",
file_system="flash:", direction="put", overwrite_file=False) # SCP; needs ip scp server enable
from netmiko import SSHDetect
guesser = SSHDetect(**{**device, "device_type": "autodetect"})
device["device_type"] = guesser.autodetect() # cisco_ios, cisco_nxos, arista_eos...
# Connect through a jump host with an SSH config file
device["ssh_config_file"] = "~/.ssh/config" # ProxyJump and per-host keys honouredsend_config_set returns the echoed output, so scan it for % Invalid input or % Incomplete command and fail the job: IOS does not raise on a bad line, it prints and moves on.
Many devices with Nornir#
Nornir 3 is a Python framework that provides inventory, filtering and a threaded runner. Tasks are plain Python functions, so ordinary debugging works.
# config.yaml
inventory:
plugin: SimpleInventory
options:
host_file: inventory/hosts.yaml
group_file: inventory/groups.yaml
runner:
plugin: threaded
options:
num_workers: 10from nornir import InitNornir
from nornir_netmiko.tasks import netmiko_send_command
from nornir_utils.plugins.functions import print_result
nr = InitNornir(config_file="config.yaml")
result = nr.filter(site="syd").run(
task=netmiko_send_command, command_string="show version", use_textfsm=True
)
print_result(result)
print(result.failed_hosts) # dict of host -> MultiResult for hosts that raisedA failure on one host does not stop the others. Check result.failed or failed_hosts rather than assuming success.
# inventory/hosts.yaml
sw-01.example.com:
hostname: sw-01.example.com
platform: ios
groups: [syd-access]
data: { site: syd, role: access }
# inventory/groups.yaml
syd-access:
groups: [ios]
ios:
platform: ios
connection_options:
netmiko: { extras: { secret: "" } } # per-plugin options; credentials come from defaults.yaml or envfrom nornir.core.filter import F
from nornir_napalm.plugins.tasks import napalm_get, napalm_configure
from nornir_utils.plugins.tasks.files import write_file
core = nr.filter(F(site="syd") & F(role="core") & ~F(platform="nxos")) # boolean filters on host data
facts = core.run(task=napalm_get, getters=["facts", "interfaces_ip"])
for host, multi in facts.items():
print(host, multi[0].result["facts"]["os_version"])
def backup(task): # a task that calls other tasks
cfg = task.run(task=napalm_get, getters=["config"]).result["config"]["running"]
task.run(task=write_file, filename=f"backups/{task.host}.cfg", content=cfg)
core.run(task=backup)
r = core.run(task=napalm_configure, configuration="ntp server 192.0.2.123", dry_run=True) # diff only
print({h: res[0].diff for h, res in r.items()})Every task result carries .result, .diff, .changed, .failed and .exception. nr.data.reset_failed_hosts() clears the failed set so a retry run includes them; without it Nornir skips hosts that failed earlier in the same process.
Vendor-neutral state with NAPALM#
NAPALM gives the same getters and config methods across drivers (ios, eos, junos, nxos, nxos_ssh, iosxr_netconf). The ios driver uses Netmiko underneath and returns data parsed by NAPALM.
import os
from napalm import get_network_driver
driver = get_network_driver("ios")
with driver("sw-01.example.com", os.environ["NET_USER"], os.environ["NET_PASS"]) as dev:
facts = dev.get_facts()
interfaces = dev.get_interfaces()
neighbours = dev.get_lldp_neighbors()
arp = dev.get_arp_table()
bgp = dev.get_bgp_neighbors() # {"global": {"peers": {ip: {"is_up": ..., "address_family": {...}}}}}
env = dev.get_environment() # fans, power, temperature, cpu, memory
counters = dev.get_interfaces_counters() # tx/rx errors and discards per interface
running = dev.get_config(retrieve="running", sanitized=True)["running"] # secrets masked
out = dev.cli(["show ip route summary", "show clock"]) # {command: output} for anything without a getter
ok = dev.ping("192.0.2.1", source="192.0.2.2", count=5)["success"]["packet_loss"] == 0NAPALM also validates state against a YAML file, which turns “did the change work” into a pass/fail report:
# validate.yml
- get_facts:
os_version: "17.12" # substring match
- get_bgp_neighbors:
global:
peers:
203.0.113.1:
is_up: true
address_family:
ipv4:
received_prefixes: { _mode: ">=", value: 10 } # comparison operators on numbers
- get_interfaces:
GigabitEthernet1/0/1:
is_up: true
is_enabled: truereport = dev.compliance_report("validate.yml")
print(report["complies"], {k: v for k, v in report.items() if isinstance(v, dict) and not v.get("complies", True)})Parsing CLI output#
from ntc_templates.parse import parse_output
rows = parse_output(platform="cisco_ios", command="show ip interface brief", data=raw)
# [{'interface': 'GigabitEthernet1/0/1', 'ip_address': '198.51.100.2', 'status': 'up', 'proto': 'up'}, ...]ntc-templates selects a TextFSM template by platform and command. Field names were standardised in ntc-templates 4.0 (for example ipaddr became ip_address), so code written against older releases may look up keys that no longer exist.
When no template exists, write one rather than an ad hoc regex. A TextFSM template is a state machine: Value lines declare fields, rules match lines and Record emits a row. A line that matches no rule is ignored, so an unexpected format produces missing rows rather than wrong values.
Value INTERFACE (\S+)
Value IP_ADDRESS (\S+)
Value STATUS (up|down|administratively down|deleted)
Value PROTO (up|down)
Start
^${INTERFACE}\s+${IP_ADDRESS}\s+\w+\s+\w+\s+${STATUS}\s+${PROTO} -> RecordWhere the device can produce structured output, ask it instead. NX-OS supports | json; IOS XE does not, so use RESTCONF or NETCONF there.
import json
data = json.loads(conn.send_command("show ip route | json")) # NX-OSConfiguration changes#
Diff every change before it is applied, and know how you will back it out.
NAPALM: load, diff, commit#
with driver("sw-01.example.com", user, password) as dev:
dev.load_merge_candidate(filename="vlan.cfg")
diff = dev.compare_config()
if diff and approved(diff):
dev.commit_config(revert_in=300) # confirmed commit: reverts in 5 minutes unless confirmed
verify(dev) # your own post-change checks
dev.confirm_commit()
else:
dev.discard_config()| Capability | EOS | Junos | IOS | NX-OS | IOS XR (NETCONF) |
|---|---|---|---|---|---|
| Replace and merge | Yes | Yes | Yes | Yes | Yes |
Commit confirm (revert_in) | Yes | Yes | Yes | No | No |
| Atomic merge (all or nothing) | Yes | Yes | No | No | Yes |
The IOS driver copies the candidate file to the device with SCP and applies it with configure replace or a merge, so it needs ip scp server enable and an archive path on local flash for rollback. See the NAPALM IOS notes and support matrix.
Netmiko: send commands#
with ConnectHandler(**device) as conn:
conn.enable()
out = conn.send_config_set([
"interface GigabitEthernet1/0/2",
"description uplink to core",
"switchport mode trunk",
])
conn.save_config() # write memory; without it the change is lost on reloadsend_config_set enters config mode, sends each line and exits. It does not diff, check idempotence or roll back, so pair it with a before/after capture.
Ansible: declarative with check mode#
Ansible network modules run on the control node and talk to the device over the network_cli connection. See Ansible for playbook basics.
# group_vars/ios.yml
ansible_connection: ansible.netcommon.network_cli
ansible_network_os: cisco.ios.ios
ansible_become: true
ansible_become_method: enable- name: Configure access VLANs
hosts: ios
gather_facts: false
tasks:
- name: Ensure VLANs exist
cisco.ios.ios_vlans:
config:
- vlan_id: 20
name: servers
state: merged
- name: Set uplink description
cisco.ios.ios_config:
parents: interface GigabitEthernet1/0/24
lines:
- description uplink to core
backup: true
save_when: modifiedansible-playbook vlans.yml --check --diff --limit sw-01 # show what would change, change nothingResource modules such as ios_vlans and ios_interfaces take structured data and a state: merged adds, replaced rewrites the listed items, overridden rewrites and removes anything not listed, deleted removes, and gathered reads current state back as data. Two more states run without a device: rendered turns the data into the commands it would send, and parsed turns saved show running-config output into the module’s data model, which is how you bootstrap a source of truth from an existing network.
- name: Facts, commands and resource modules
hosts: ios
gather_facts: false
tasks:
- name: Collect structured facts and the running config
cisco.ios.ios_facts:
gather_subset: [min, config]
gather_network_resources: [interfaces, l2_interfaces, vlans] # resource-module data models
register: facts
- name: Show commands, waiting for a condition
cisco.ios.ios_command:
commands:
- show ip ospf neighbor
- show ip bgp summary
wait_for:
- result[0] contains FULL # retry until true or fail
retries: 6
interval: 10
register: show
- name: Interfaces as data, overriding what is there
cisco.ios.ios_l2_interfaces:
config:
- name: GigabitEthernet1/0/2
mode: access
access: { vlan: 20 }
- name: GigabitEthernet1/0/24
mode: trunk
trunk: { allowed_vlans: [10, 20, 30], native_vlan: 999 }
state: replaced # only the listed interfaces are rewritten
- name: What commands would the data become (no device contact)
cisco.ios.ios_vlans:
config: [{ vlan_id: 20, name: servers }]
state: rendered
register: rendered
- name: Reach into a saved config and produce data
cisco.ios.ios_interfaces:
running_config: "{{ lookup('file', 'backups/sw-01.cfg') }}"
state: parsed
register: parsed
- name: Free-form lines with a diff against the intended config
cisco.ios.ios_config:
src: templates/baseline.j2 # Jinja template rendered on the control node
diff_against: intended
intended_config: "{{ lookup('template', 'templates/baseline.j2') }}"
diff_ignore_lines: ['^ntp clock-period', '^! Last configuration']
backup: true
backup_options: { dir_path: backups/, filename: "{{ inventory_hostname }}.cfg" }
save_when: modifiedios_command never enters config mode and is the right module for read-only checks. ios_config match: line (default) sends only lines that differ; match: exact sends when the block differs in order too, and match: none sends everything regardless, which is what you want with before: [no ip access-list extended MGMT-IN] to rebuild an ACL atomically. Connection settings that matter: ansible_command_timeout (default 30 s, raise for show tech or long commits) and ansible_persistent_connect_timeout; the persistent connection is reused across tasks in a play, so a hostname change mid-play breaks prompt detection for every later task.
ansible.netcommon.cli_backup (cli_backup in the netcommon collection) is the vendor-neutral backup task, and ansible.utils.cli_parse runs TextFSM, TTP or native parsers on any command output inside a playbook.
A config push can remove your own access
Use a confirmed commit (revert_in, configure replace ... time) or reload in 10 before the change and reload cancel after verifying. On a device with neither, arrange console access first. See Cisco change safety.
Order matters for anything on the path you are connected over: add the new configuration, verify, then remove the old. An ACL applied before your own permit ends the session.
Configuration diff and rollback#
A change is safe when three artefacts exist before it runs: the running config as it was, the intended config, and the diff between them. Which tool makes the diff depends on the platform, but the shape is the same.
| Approach | Diff | Rollback | Notes |
|---|---|---|---|
NAPALM load_replace_candidate + compare_config | Exact, device-generated on EOS/Junos; configure replace diff on IOS | commit_config(revert_in=) then confirm_commit(), or rollback() after commit | Replace is the only way to remove config you did not list |
NAPALM load_merge_candidate | Lines to add | rollback() restores the pre-change archive on IOS | Cannot remove lines |
Ansible ios_config --check --diff | Line-based, what Ansible would send | backup: true gives a file; restoring it is a configure replace you run yourself | diff_against: intended for compliance runs |
Netmiko send_config_set | None; capture show run before and after and diff -u | Manual | Pair with reload in on the device |
Git-backed intended configs plus configure replace | git diff | configure replace flash:previous.cfg | The device applies exactly the file; the repository is the source of truth |
# Replace the whole config from a rendered template, with a timed revert and post-checks
import difflib
with driver("sw-01.example.com", user, password) as dev:
before = dev.get_config(retrieve="running")["running"]
dev.load_replace_candidate(filename="rendered/sw-01.cfg")
diff = dev.compare_config()
if not diff:
dev.discard_config(); print("no change"); raise SystemExit
print(diff)
dev.commit_config(revert_in=300) # device reverts in 5 minutes unless confirmed
checks = dev.compliance_report("validate.yml")
if checks["complies"] and dev.ping("192.0.2.1")["success"]["packet_loss"] == 0:
dev.confirm_commit()
else:
dev.rollback() # explicit, rather than waiting for the timer
raise SystemExit("post-checks failed; rolled back")
after = dev.get_config(retrieve="running")["running"]
open("diffs/sw-01.diff", "w").writelines(difflib.unified_diff(before.splitlines(True), after.splitlines(True), "before", "after"))has_pending_commit() tells you whether a previous run left a timer running; call it first and confirm_commit() or rollback() before loading a new candidate. IOS keeps the pre-change snapshot under the archive path (rollback-<n>), so rollback() works after confirm_commit() too, until the archive rotates. On NX-OS, which has no confirmed commit, use checkpoint and rollback running-config checkpoint <name> through cli() around the change.
# Ansible: back up, change with a diff, verify, and restore the backup on failure
- name: Change with rollback
hosts: ios
gather_facts: false
tasks:
- name: Backup
cisco.ios.ios_config:
backup: true
backup_options: { dir_path: "backups/{{ inventory_hostname }}", filename: pre-change.cfg }
check_mode: false
- name: Copy the backup to flash for configure replace
ansible.netcommon.net_put:
src: "backups/{{ inventory_hostname }}/pre-change.cfg"
dest: flash:pre-change.cfg
check_mode: false
- name: Apply
block:
- name: Push change
cisco.ios.ios_config:
src: templates/change.j2
diff_against: running
- name: Verify
cisco.ios.ios_command:
commands: [show ip ospf neighbor]
wait_for: [result[0] contains FULL]
retries: 6
interval: 10
rescue:
- name: Roll back to the pre-change config
cisco.ios.ios_command:
commands:
- command: configure replace flash:pre-change.cfg force
- name: Fail the play after restoring
ansible.builtin.fail:
msg: "verification failed on {{ inventory_hostname }}; configuration restored"Normalise before diffing: strip timestamps, ntp clock-period, ! Last configuration change and the Building configuration header, or every backup differs. Type 7 passwords re-encode with a different salt on some platforms, which is another false diff; diff_ignore_lines in Ansible and a sed filter in scripts handle both.
Idempotence and templates#
Render the intended configuration from data, compare it with the running configuration, and push only the difference. Reruns are then safe and the repository becomes the source of truth.
from jinja2 import Environment, FileSystemLoader, StrictUndefined
env = Environment(loader=FileSystemLoader("templates"), trim_blocks=True, lstrip_blocks=True,
undefined=StrictUndefined) # fail on a missing variable instead of rendering blank
cfg = env.get_template("switch.j2").render(hostname="sw-01", vlans=[10, 20, 30])hostname {{ hostname }}
{% for vlan in vlans %}
vlan {{ vlan }}
name VLAN{{ vlan }}
{% endfor %}Batfish analyses configuration files offline and answers reachability and ACL questions, which catches a broken ACL before deployment. pyATS/Genie parses and compares device state before and after a change.
Bulk operations safely#
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
def collect(host: str) -> tuple[str, str]:
with ConnectHandler(host=host, **base) as c:
return host, c.send_command("show version")
with ThreadPoolExecutor(max_workers=10) as pool: # devices and TACACS/RADIUS servers cap concurrent sessions
for host, out in pool.map(collect, hosts):
Path(f"out/{host}.txt").write_text(out)Stage changes: one device, then one site, then the fleet, verifying between each. Capture state before and after on every device so a rollback has something to compare against.
Discovery and verification#
fping -a -g 192.0.2.0/24 2>/dev/null # addresses that answer ICMP
nmap -sn 192.0.2.0/24 -oG - # host discovery, greppable output
snmpwalk -v2c -c "$SNMP_COMMUNITY" sw-01.example.com IF-MIB::ifDescr # interface list via SNMPdev.get_lldp_neighbors() from NAPALM builds a topology from the devices themselves.
Verify after every change against state collected before it: interface counters, neighbour tables, route counts and a targeted reachability test. A command that applied without error has not been verified.
Troubleshooting#
| Symptom | Cause | Check |
|---|---|---|
NetmikoAuthenticationException on some devices | AAA policy differs by device group, or local fallback account differs | Log in by hand with the same account; check TACACS/RADIUS logs |
NetmikoTimeoutException | TCP connect failed: routing, ACL on VTY lines, wrong port | nc -vz host 22; conn_timeout |
ReadTimeout: Pattern not detected | Prompt changed (hostname change, config mode, confirmation prompt) or command slower than read_timeout | Set expect_string or raise read_timeout; enable session_log |
use_textfsm=True returns a string | No template for that platform and command, or output did not match | Check the ntc-templates index; parse manually or add a template |
| Parsed keys missing after an upgrade | ntc-templates 4.0 renamed fields | Update key names (ip_address not ipaddr) |
| Works interactively, fails in a script | enable not called, or the command needs config mode | conn.check_enable_mode(), send_config_set |
| Random failures at scale | Too many concurrent sessions; VTY lines or AAA rate limits exhausted | Lower workers; show users on the device |
| NAPALM IOS commit fails | SCP server disabled or no archive configured | ip scp server enable, archive path flash:... |
Ansible network_cli hangs or times out | Wrong ansible_network_os, or enable not configured | ansible-playbook -vvvv; set ANSIBLE_PERSISTENT_COMMAND_TIMEOUT |
| Config applied but gone after reload | Never saved | save_config(), save_when: modified, copy run start |
send_config_set “succeeded” but the config is wrong | IOS printed % Invalid input and carried on | Check the returned output for %; use cmd_verify=True (default) |
NAPALM compare_config shows the whole config as changed | Candidate missing lines the device adds itself, or line-ending/encoding differences | Start from get_config() output, edit, and re-replace; check for \r |
NAPALM commit_config(revert_in=) raises on IOS | Archive not configured, or an earlier pending commit | archive + path flash:...; has_pending_commit() |
Ansible resource module reports changed every run | Device normalises values (VLAN lists, case, ranges) differently from the data | Compare gathered output with your data and match its form |
ios_command wait_for never passes | Wrong conditional syntax or output has ANSI/paging | Use result[0] contains X; check terminal length 0 ran |
| Ansible persistent connection hangs after a hostname change | Prompt no longer matches | Finish the play, or meta: reset_connection after the change |
net_put / SCP transfer fails | SCP server disabled or no space on flash | ip scp server enable; dir flash: |
| Nornir skips hosts on a second run in the same process | Hosts marked failed earlier | nr.data.reset_failed_hosts() |
| TextFSM template parses but returns fewer rows than expected | Output format changed with the software release | Compare raw output with the template’s regex; open an ntc-templates issue or add a template |
import logging
logging.basicConfig(filename="netmiko_debug.log", level=logging.DEBUG) # library debug log
device["session_log"] = "session.log" # full transcript of what was sent and receivedCaution
Netmiko masks the login password and enable secret in the session log, but everything else is recorded, including show running-config output and any keys or community strings you configure. Keep it out of shared storage and delete it when done.
Oneliners#
# Reachability sweep
fping -a -g 192.0.2.0/24 2>/dev/null | tee reachable.txt
# Run one command on every device and keep the output
while read -r h; do echo "== $h"; ssh -o ConnectTimeout=5 "$h" 'show version | include uptime'; done < hosts.txt
# Diff a device's running config against the repository copy
ssh sw-01.example.com 'show running-config' | diff -u configs/sw-01.cfg - | head -40
# Parse saved command output into JSON
python3 -c 'import sys,json;from ntc_templates.parse import parse_output;print(json.dumps(parse_output(platform="cisco_ios",command=sys.argv[1],data=sys.stdin.read())))' "show ip interface brief" < out.txt
# Interfaces that are down but not administratively down
ssh sw-01.example.com 'show ip interface brief' | awk '$5=="down" && $6=="down"'
# Back up every device before a change window
while read -r h; do ssh "$h" 'show running-config' > "backups/$h-$(date +%F).cfg"; done < hosts.txt
# Find a MAC address across the fleet
while read -r h; do ssh "$h" 'show mac address-table | include 0011.2233' | sed "s/^/$h /"; done < hosts.txt
# Count routes before and after a change
ssh rtr-01.example.com 'show ip route summary | include Total'
# Normalise a config for diffing (strip volatile lines)
sed -E '/^(Building configuration|Current configuration|! Last configuration change|! NVRAM config last updated|ntp clock-period)/d' sw-01.cfg > sw-01.norm.cfg
# Diff every backup against the previous day's copy and list devices that changed
for f in backups/*-"$(date +%F)".cfg; do h=${f#backups/}; h=${h%-*}; diff -q "backups/$h-$(date -d yesterday +%F).cfg" "$f" >/dev/null || echo "$h changed"; done
# Facts from every device as JSON with NAPALM's CLI
napalm --user "$NET_USER" --password "$NET_PASS" --vendor ios sw-01.example.com call get_facts
# Diff a candidate config against the device without applying it
napalm --user "$NET_USER" --password "$NET_PASS" --vendor ios sw-01.example.com configure candidate.cfg --strategy replace --dry-run
# Validate a device against a state file, exit non-zero on failure
napalm --user "$NET_USER" --password "$NET_PASS" --vendor ios sw-01.example.com validate validate.yml
# Ansible: run one show command everywhere and print output per host
ansible ios -m cisco.ios.ios_command -a 'commands="show ip interface brief"' | sed -n '/stdout_lines/,/]/p'
# Ansible: gather structured facts for one host into JSON
ansible sw-01.example.com -m cisco.ios.ios_facts -a 'gather_subset=min gather_network_resources=interfaces' > facts.json
# Ansible: back up every device with the vendor-neutral module
ansible ios -m ansible.netcommon.cli_backup -a 'dir_path=backups/'
# Ansible: what a playbook would change, for one site
ansible-playbook change.yml --check --diff --limit site_syd
# Which ntc-templates exist for a platform
python3 -c 'import ntc_templates,os;p=os.path.join(os.path.dirname(ntc_templates.__file__),"templates");print("\n".join(sorted(f for f in os.listdir(p) if f.startswith("cisco_ios"))))'
# Test a TextFSM template against saved output
python3 -c 'import sys,textfsm;print(textfsm.TextFSM(open(sys.argv[1])).ParseTextToDicts(open(sys.argv[2]).read()))' template.textfsm out.txt
# Devices whose SSH host key changed since the last run (possible replacement or MITM)
while read -r h; do ssh-keyscan -t ed25519 -T 5 "$h" 2>/dev/null | ssh-keygen -lf - ; done < hosts.txt | sort > keys.new; diff keys.old keys.new
# SSH algorithm negotiation failure on old IOS: allow legacy kex for one connection
ssh -o KexAlgorithms=+diffie-hellman-group14-sha1 -o HostKeyAlgorithms=+ssh-rsa sw-old.example.com
# Software version distribution across the fleet
while read -r h; do ssh "$h" 'show version | include Version' | head -1 | sed "s/^/$h /"; done < hosts.txt | awk '{print $NF}' | sort | uniq -c
# Interfaces with errors across the fleet (raw counters, no parser)
while read -r h; do ssh "$h" 'show interfaces | include line protocol|input errors' | paste - - | awk -v h="$h" '$NF+0>0 || $(NF-6)+0>0 {print h, $1}'; done < hosts.txt
# LLDP neighbour list from every device, as edges for a topology
while read -r h; do ssh "$h" 'show lldp neighbors detail | include System Name|Port id|Local Intf' | paste - - - | sed "s/^/$h /"; done < hosts.txt
# Uptime under a day (devices that reloaded recently)
while read -r h; do ssh "$h" 'show version | include uptime' | grep -vE 'week|day' | sed "s/^/$h /"; done < hosts.txt
# Serial numbers for an inventory or RMA
while read -r h; do ssh "$h" 'show inventory | include SN:' | head -1 | sed "s/^/$h /"; done < hosts.txt
# Batfish: check reachability against a directory of configs (needs a running batfish container)
python3 -c 'from pybatfish.client.session import Session;bf=Session();bf.init_snapshot("configs/",name="s",overwrite=True);print(bf.q.reachability(pathConstraints={"startLocation":"sw-01"},headers={"dstIps":"192.0.2.10","applications":["ssh"]}).answer().frame())'
# Count of lines per config file, to spot a truncated backup
wc -l backups/*.cfg | sort -n | headScripts#
Collect a state baseline from every device before a change window and compare it afterwards, reporting neighbours, routes and interfaces that differ.
#!/usr/bin/env python3
"""Snapshot or compare network state with NAPALM.
Usage: netstate.py snapshot hosts.txt before/
netstate.py compare before/ after/
Credentials from NET_USER and NET_PASS.
"""
import json
import os
import sys
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from napalm import get_network_driver
GETTERS = ["get_facts", "get_interfaces", "get_lldp_neighbors", "get_bgp_neighbors", "get_arp_table"]
def snapshot(host: str, outdir: Path) -> str:
driver = get_network_driver(os.environ.get("NET_DRIVER", "ios"))
with driver(host, os.environ["NET_USER"], os.environ["NET_PASS"], optional_args={"secret": os.environ.get("NET_ENABLE", "")}) as dev:
state = {g: getattr(dev, g)() for g in GETTERS}
routes = dev.cli(["show ip route summary"])["show ip route summary"]
state["route_summary"] = [l for l in routes.splitlines() if l.strip().startswith(("Total", "connected", "static", "ospf", "bgp"))]
state["interfaces_up"] = sorted(i for i, d in state["get_interfaces"].items() if d["is_up"])
state["bgp_up"] = sorted(p for p, d in state["get_bgp_neighbors"].get("global", {}).get("peers", {}).items() if d["is_up"])
state["lldp"] = sorted(f"{i} -> {n['hostname']}:{n['port']}" for i, ns in state["get_lldp_neighbors"].items() for n in ns)
(outdir / f"{host}.json").write_text(json.dumps(state, indent=1, sort_keys=True, default=str))
return host
def compare(a: Path, b: Path) -> int:
rc = 0
for fa in sorted(a.glob("*.json")):
fb = b / fa.name
if not fb.exists():
print(f"{fa.stem}: missing in {b}"); rc = 1; continue
sa, sb = json.loads(fa.read_text()), json.loads(fb.read_text())
for key in ("interfaces_up", "bgp_up", "lldp", "route_summary"):
lost, gained = sorted(set(sa[key]) - set(sb[key])), sorted(set(sb[key]) - set(sa[key]))
if lost or gained:
rc = 1
print(f"{fa.stem} {key}: -{lost} +{gained}")
return rc
if sys.argv[1] == "snapshot":
hosts = [h.strip() for h in open(sys.argv[2]) if h.strip() and not h.startswith("#")]
out = Path(sys.argv[3]); out.mkdir(parents=True, exist_ok=True)
with ThreadPoolExecutor(max_workers=10) as pool:
for fut in [pool.submit(snapshot, h, out) for h in hosts]:
try:
print("ok", fut.result())
except Exception as exc:
print("FAILED", exc, file=sys.stderr)
else:
sys.exit(compare(Path(sys.argv[2]), Path(sys.argv[3])))Render per-device configs from a YAML source of truth and Jinja templates, then diff each against the live device with NAPALM without applying, as a CI job that fails on drift.
#!/usr/bin/env python3
"""Render intended configs and report drift against live devices.
Usage: drift.py devices.yml templates/ [--apply]
devices.yml: a list of {host, platform, template, vars...}; templates/<template> is a Jinja file.
"""
import os
import sys
import yaml
from jinja2 import Environment, FileSystemLoader, StrictUndefined
from napalm import get_network_driver
devices = yaml.safe_load(open(sys.argv[1]))
env = Environment(loader=FileSystemLoader(sys.argv[2]), trim_blocks=True, lstrip_blocks=True, undefined=StrictUndefined)
apply = "--apply" in sys.argv
drift = 0
for d in devices:
intended = env.get_template(d["template"]).render(**d)
driver = get_network_driver(d.get("platform", "ios"))
try:
with driver(d["host"], os.environ["NET_USER"], os.environ["NET_PASS"]) as dev:
dev.load_replace_candidate(config=intended)
diff = dev.compare_config()
if not diff:
dev.discard_config(); print(f"{d['host']}: in sync"); continue
drift += 1
print(f"{d['host']}: DRIFT\n{diff}\n")
if apply:
dev.commit_config(revert_in=180)
if dev.ping(d.get("check_ip", "192.0.2.1"))["success"]["packet_loss"] == 0:
dev.confirm_commit(); print(f"{d['host']}: applied")
else:
dev.rollback(); print(f"{d['host']}: rolled back, ping failed")
else:
dev.discard_config()
except Exception as exc:
drift += 1
print(f"{d['host']}: ERROR {exc}", file=sys.stderr)
sys.exit(1 if drift else 0)Bulk-execute a read-only command across the fleet with Netmiko, parse it with TextFSM, and emit one CSV for a spreadsheet or a ticket.
#!/usr/bin/env python3
"""Run a show command on many devices and write parsed rows to CSV.
Usage: show2csv.py hosts.txt "show interfaces status" out.csv
"""
import csv
import os
import sys
from concurrent.futures import ThreadPoolExecutor
from netmiko import ConnectHandler
hosts_file, command, out = sys.argv[1:4]
hosts = [h.strip() for h in open(hosts_file) if h.strip() and not h.startswith("#")]
def run(host):
dev = {"device_type": os.environ.get("NET_TYPE", "cisco_ios"), "host": host, "username": os.environ["NET_USER"],
"password": os.environ["NET_PASS"], "secret": os.environ.get("NET_ENABLE", ""), "conn_timeout": 10, "read_timeout": 60}
try:
with ConnectHandler(**dev) as c:
if dev["secret"]:
c.enable()
rows = c.send_command(command, use_textfsm=True)
if not isinstance(rows, list):
return host, None, "no template match"
return host, rows, None
except Exception as exc:
return host, None, str(exc)
fields, records, failures = None, [], []
with ThreadPoolExecutor(max_workers=10) as pool:
for host, rows, err in pool.map(run, hosts):
if err:
failures.append((host, err)); continue
for r in rows:
fields = fields or ["host", *r.keys()]
records.append({"host": host, **r})
with open(out, "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=fields or ["host"])
w.writeheader(); w.writerows(records)
print(f"{len(records)} rows from {len(hosts) - len(failures)} devices -> {out}")
for host, err in failures:
print(f"FAILED {host}: {err}", file=sys.stderr)
sys.exit(1 if failures else 0)