An instance nobody maintains is a liability: the day it breaks is the day you find out the backup was never tested. This page sets up a nightly backup and restores it once, upgrades NetBox to a version you chose on purpose, wires the housekeeping job, and leaves a probe behind that tells you when something is wrong.
Before you start
$systemctl is-active netbox netbox-rq postgresqlactive active active
Nightly backups
$sudo mkdir -p && sudo chmod 700 #!/usr/bin/env bashset -euo pipefailstamp=$(date +%Y%m%d-%H%M)dest=sudo -u postgres pg_dump -Fc netbox > "$dest/netbox-$stamp.dump"files=(netbox/media netbox/netbox/configuration.py)for f in netbox/netbox/ldap_config.py local_requirements.txt gunicorn.py; do if [ -e "/$f" ]; then files+=("$f"); fidonetar -czf "$dest/netbox-files-$stamp.tar.gz" -C "${files[@]}"find "$dest" -name 'netbox-*' -mtime + -delete$sudo chmod 750 /usr/local/sbin/netbox-backup.sh$sudo systemctl daemon-reload$sudo systemctl enable --now netbox-backup.timer$sudo systemctl start netbox-backup.service$systemctl is-active netbox-backup.timeractive
$ls | sed 's/-[0-9]*-[0-9]*\././' | LC_ALL=C sort -unetbox-files.tar.gz netbox.dump
If the service fails
journalctl -u netbox-backup -n 20 has the reason. The usual ones: pg_dump cannot connect (PostgreSQL is stopped, or pg_hba.conf lost its local all postgres peer line); tar complains about a missing netbox/media (the folder was moved, fix the path); no space left in .
Restore drill
$latest=$(ls -t /netbox-*.dump | head -1)$sudo -u postgres createdb -O netbox netbox_drill$sudo -u postgres pg_restore --no-owner --role=netbox -d netbox_drill < "$latest"$sudo -u postgres psql -d netbox_drill -tAc 'SELECT count(*) > 0 FROM dcim_site't
$sudo -u postgres dropdb netbox_drillUpgrade NetBox
$sudo systemctl start netbox-backup.service$cd $sudo git fetch --tags$sudo git checkout $sudo ./upgrade.sh$sudo systemctl restart netbox netbox-rq$sudo git -C describe --tags --exact-match$curl -skf -H 'Authorization: Token ' https:///api/status/ | python3 -c 'import sys,json; print("v" + json.load(sys.stdin)["netbox-version"])'$curl -sk -o /dev/null -w '%{http_code}' https:///login/200
Then open https:/// and click through a device page and a search: the API answering is not the same as the UI rendering, since static files and plugins can break one and not the other.
If a plugin blocks the upgrade
- Check the plugin's compatibility table on its repository. If a compatible release exists, pin it in
local_requirements.txt(netbox-bgp==0.15.0) and re-runsudo ./upgrade.sh. - If none exists, remove the plugin from
PLUGINSinconfiguration.pyand fromlocal_requirements.txt, upgrade, and put it back when the plugin catches up. Its tables stay in the database; nothing is lost, the pages just disappear until then. - If the release notes say the plugin is now a core feature (it happens), migrate the data with the plugin's own instructions before upgrading.
Rollback
$sudo systemctl stop netbox netbox-rq$cd $sudo git checkout $(sudo git describe --tags --abbrev=0 ^)$latest=$(ls -t /netbox-*.dump | head -1)$sudo -u postgres dropdb netbox && sudo -u postgres createdb -O netbox netbox$sudo -u postgres pg_restore --no-owner --role=netbox -d netbox < "$latest"$sudo ./upgrade.sh$sudo systemctl start netbox netbox-rqHousekeeping
$sudo ln -sf /contrib/netbox-housekeeping.sh /etc/cron.daily/netbox-housekeeping$sudo /venv/bin/python /netbox/manage.py housekeeping$run-parts --test /etc/cron.daily | grep netbox/etc/cron.daily/netbox-housekeeping
Logs and disk: gunicorn writes to stdout, which systemd captures in the journal; cap it. nginx and Apache rotate their own files through logrotate, shipped with the package.
[Journal]SystemMaxUse=500MThen sudo systemctl restart systemd-journald.
Weekly, five minutes:
| Check | Command | Expected |
|---|---|---|
| Both services up | systemctl is-active netbox netbox-rq | active twice |
| Last backup ran | systemctl list-timers netbox-backup.timer | a LAST within 24h |
| Backup size sane | ls -lh ${BACKUP_DIR} | sizes in the same range as last week |
| Disk | df -h / | under 80% |
| Queue quiet | redis-cli info memory | used_memory_human stable |
| New release? | releases page | read the notes, plan the upgrade |
Monitoring basics
$systemctl status netbox netbox-rq --no-pager | grep Active$curl -skf -H "Authorization: Token " https:///api/status/ | python3 -m json.tool$curl -skf -H 'Authorization: Token ' https:///api/status/ | python3 -c 'import sys,json; print(json.load(sys.stdin)["rq-workers-running"] >= 1)'True
Done
Every night a dump and an archive land in , you have restored one by hand, NetBox runs , the change log is pruned daily, and a probe knows when the workers stop. The next page stops clicking: an automation user, pynetbox, bulk imports, webhooks and custom scripts.