# Source of "Maintenir et mettre à jour NetBox"

The files of this tutorial, as they are in the repository. The authoring brief that defines the format follows them.

````yaml title="content/netbox/series.yaml"
title:
  en: NetBox
  fr: NetBox
summary:
  en: >-
    NetBox is the source of truth for your network: devices, IPs, cables, circuits.
    This series takes it from a blank server to something you maintain and automate against.
  fr: >-
    NetBox est la source de vérité de ton réseau : équipements, IP, câbles, circuits.
    Cette série part d'un serveur vierge et va jusqu'au maintien et à l'automatisation.
order:
  - install-from-scratch
  - configure-for-your-team
  - maintain-and-upgrade
  - automate-with-the-api

# Shared by every page: the reader fills these once for the whole series.
groups:
  - id: host
    label: { en: Server, fr: Serveur }
    desc: { en: Where NetBox runs and how it is reached., fr: Où tourne NetBox et comment on l'atteint. }
  - id: auth
    label: { en: Directory, fr: Annuaire }
    desc: { en: Your LDAP or Active Directory., fr: Ton LDAP ou Active Directory. }
    when: { flag: LDAP }

vars:
  - key: NETBOX_HOST
    kind: hostname
    group: host
    default: netbox.example.com
    label: { en: NetBox hostname, fr: Nom d'hôte de NetBox }
    hint:
      en: The address users will type in their browser to reach NetBox. It must already point to this server in your DNS.
      fr: L'adresse que les utilisateurs taperont dans le navigateur pour atteindre NetBox. Elle doit déjà pointer vers ce serveur dans ton DNS.
    impact:
      en: Written into ALLOWED_HOSTS (Django refuses any other name with a 400 error), into the web server configuration and into the TLS certificate. Several names are possible, separated by commas.
      fr: Écrit dans ALLOWED_HOSTS (Django refuse tout autre nom avec une erreur 400), dans la configuration du serveur web et dans le certificat TLS. Plusieurs noms sont possibles, séparés par des virgules.
  - key: INSTALL_DIR
    kind: path
    group: host
    default: /opt/netbox
    label: { en: Install directory, fr: Répertoire d'installation }
    hint:
      en: The folder where NetBox's code and Python environment live. Keep the default unless your organisation has a rule about it.
      fr: Le dossier où vivent le code de NetBox et son environnement Python. Garde la valeur par défaut sauf règle interne.
    impact:
      en: The systemd units, the upgrade script and the web server configuration shipped with NetBox all assume /opt/netbox. Changing it means editing those files too.
      fr: Les unités systemd, le script d'upgrade et la configuration du serveur web livrés avec NetBox supposent tous /opt/netbox. Le changer oblige à éditer ces fichiers aussi.
  - key: LDAP_URI
    kind: text
    group: auth
    default: ldaps://ad.example.com:636
    label: { en: LDAP server URI, fr: URI du serveur LDAP }
    hint:
      en: Address of your directory server, with the protocol. Use ldaps:// (port 636) so credentials travel encrypted.
      fr: Adresse de ton serveur d'annuaire, avec le protocole. Utilise ldaps:// (port 636) pour que les identifiants circulent chiffrés.
    impact:
      en: Every login makes NetBox contact this address. If it is unreachable, nobody can log in except local accounts such as the first administrator.
      fr: Chaque connexion fait contacter cette adresse par NetBox. Si elle est injoignable, plus personne ne peut se connecter sauf les comptes locaux comme le premier administrateur.
  - key: LDAP_BIND_DN
    kind: text
    group: auth
    default: CN=netbox,OU=Service Accounts,DC=example,DC=com
    label: { en: Bind DN, fr: DN de bind }
    hint:
      en: The directory account NetBox uses to search for users before checking their password. A dedicated read-only service account, written as a full distinguished name.
      fr: Le compte d'annuaire que NetBox utilise pour chercher les utilisateurs avant de vérifier leur mot de passe. Un compte de service dédié en lecture seule, écrit sous forme de DN complet.
    impact:
      en: Its password goes into ldap_config.py. Give it read access to the users and groups subtrees only.
      fr: Son mot de passe va dans ldap_config.py. Donne-lui un accès en lecture aux sous-arbres utilisateurs et groupes uniquement.
  - key: LDAP_BASE_DN
    kind: text
    group: auth
    default: DC=example,DC=com
    label: { en: User search base, fr: Base de recherche utilisateurs }
    hint:
      en: Where in the directory tree NetBox looks for people. Usually the domain root, or a narrower OU if only some staff should log in.
      fr: Où, dans l'arbre de l'annuaire, NetBox cherche les personnes. En général la racine du domaine, ou une OU plus étroite si seule une partie du personnel doit se connecter.
    impact:
      en: A user outside this subtree simply cannot log in, with no explicit error. Combine with a group filter for finer control.
      fr: Un utilisateur hors de ce sous-arbre ne peut simplement pas se connecter, sans erreur explicite. Combine avec un filtre de groupe pour un contrôle plus fin.

choices:
  - key: OS
    type: select
    label: { en: Distribution, fr: Distribution }
    default: ubuntu
    options:
      - { value: ubuntu, label: { en: Ubuntu 24.04, fr: Ubuntu 24.04 } }
      - { value: rhel, label: { en: RHEL 9 / Rocky / Alma, fr: RHEL 9 / Rocky / Alma } }
  - key: WEB
    type: select
    label: { en: Web server, fr: Serveur web }
    default: nginx
    options:
      - { value: nginx, label: { en: nginx, fr: nginx } }
      - { value: apache, label: { en: Apache, fr: Apache } }
  - key: LDAP
    type: boolean
    label: { en: Authenticate against LDAP / Active Directory, fr: Authentifier via LDAP / Active Directory }
    default: false
````

````yaml title="content/netbox/maintain-and-upgrade/tuto.yaml"
# Inherits from ../series.yaml: NETBOX_HOST, INSTALL_DIR, OS, WEB, LDAP and the host/auth groups.
# NETBOX_TOKEN is also declared by configure-for-your-team: same key, the reader fills it once per series.

title:
  en: Maintain and upgrade NetBox
  fr: Maintenir et mettre à jour NetBox
summary:
  en: >-
    Nightly backups you have actually restored once, an upgrade path you can roll back,
    the housekeeping job, log rotation, and a probe that tells you NetBox is alive.
  fr: >-
    Des sauvegardes nocturnes que tu as restaurées au moins une fois, une mise à jour que tu
    sais annuler, la tâche d'entretien, la rotation des logs, et une sonde qui dit si NetBox est vivant.
difficulty: intermediate
tags: [netbox, backup, postgresql, systemd, upgrade, monitoring]
authors: [thudal]
created: 2026-09-25
minutes: 35
validated: NetBox 4.4 · Ubuntu 24.04
status: draft             # not yet run end to end by its author

groups:
  - id: maintenance
    label: { en: Backups and upgrade, fr: Sauvegardes et mise à jour }
  - id: offsite
    label: { en: Off-site copy, fr: Copie distante }
    desc: { en: Where the nightly backup is pushed over SSH., fr: Où la sauvegarde nocturne est poussée en SSH. }
    when: { flag: OFFSITE }
  - id: api
    label: { en: API access, fr: Accès API }

vars:
  - key: BACKUP_DIR
    kind: path
    group: maintenance
    default: /var/backups/netbox
    label: { en: Backup directory, fr: Répertoire des sauvegardes }
    hint:
      en: Where the database dumps and the file archives land on the server. Root only.
      fr: Là où atterrissent les dumps de base et les archives de fichiers sur le serveur. Root uniquement.
    impact:
      en: Put it on a partition with room for two weeks of dumps. A NetBox database is small (tens of MB for thousands of devices); the media folder is what grows if people attach images.
      fr: Mets-le sur une partition avec de la place pour deux semaines de dumps. Une base NetBox est petite (quelques dizaines de Mo pour des milliers d'équipements) ; c'est le dossier media qui grossit si on attache des images.
  - key: BACKUP_RETENTION_DAYS
    kind: text
    group: maintenance
    default: "14"
    label: { en: Retention (days), fr: Rétention (jours) }
    hint:
      en: Backups older than this are deleted after each nightly run. 14 covers a two-week holiday.
      fr: Les sauvegardes plus vieilles que ça sont supprimées après chaque exécution nocturne. 14 couvre deux semaines de vacances.
    impact:
      en: With the off-site copy, the same retention applies on the remote side, since rsync mirrors the folder.
      fr: Avec la copie distante, la même rétention s'applique côté distant, puisque rsync reflète le dossier.
  - key: TARGET_REF
    kind: text
    group: maintenance
    default: v4.4.1
    label: { en: Version to upgrade to, fr: Version cible }
    hint:
      en: A git tag from the NetBox releases page. Read its release notes before typing it here.
      fr: Un tag git de la page des releases de NetBox. Lis ses notes de version avant de le taper ici.
    impact:
      en: Checked out in the install directory, then applied by upgrade.sh. Going down a major version is not supported; a rollback restores the database.
      fr: Extrait dans le répertoire d'installation, puis appliqué par upgrade.sh. Redescendre d'une version majeure n'est pas supporté ; un retour arrière restaure la base.
  - key: BACKUP_HOST
    kind: hostname
    group: offsite
    default: backup.example.com
    when: { flag: OFFSITE }
    label: { en: Backup host, fr: Serveur de sauvegarde }
    hint:
      en: A machine reachable over SSH from the NetBox server, on another site or at least another rack.
      fr: Une machine joignable en SSH depuis le serveur NetBox, sur un autre site ou au moins une autre baie.
  - key: BACKUP_USER
    kind: user
    group: offsite
    default: backup
    when: { flag: OFFSITE }
    label: { en: SSH user on the backup host, fr: Utilisateur SSH sur le serveur de sauvegarde }
    hint:
      en: A dedicated account that owns the destination path and nothing else.
      fr: Un compte dédié, propriétaire du chemin de destination et de rien d'autre.
  - key: BACKUP_PATH
    kind: path
    group: offsite
    default: /srv/backups/netbox
    when: { flag: OFFSITE }
    label: { en: Destination path, fr: Chemin de destination }
    hint:
      en: Folder on the backup host that mirrors the local backup directory. It must exist and belong to the SSH user.
      fr: Dossier sur le serveur de sauvegarde qui reflète le répertoire local. Il doit exister et appartenir à l'utilisateur SSH.
  - key: NETBOX_TOKEN
    kind: secret
    group: api
    default: ""
    label: { en: API token, fr: Jeton d'API }
    hint:
      en: A token with read access, from your profile → API tokens. Used to query /api/status/ after the upgrade and from the probe.
      fr: Un jeton en lecture, depuis ton profil → jetons d'API. Sert à interroger /api/status/ après la mise à jour et depuis la sonde.
    impact:
      en: With LOGIN_REQUIRED on, /api/status/ answers 403 without a token. For a monitoring probe, create a dedicated read-only user and give the probe its token, not yours.
      fr: Avec LOGIN_REQUIRED actif, /api/status/ répond 403 sans jeton. Pour une sonde de supervision, crée un utilisateur dédié en lecture seule et donne son jeton à la sonde, pas le tien.

choices:
  - key: OFFSITE
    type: boolean
    label: { en: Copy backups to another host (rsync over SSH), fr: Copier les sauvegardes sur un autre serveur (rsync en SSH) }
    hint: { en: A backup on the same disk as the database is not a backup., fr: Une sauvegarde sur le même disque que la base n'est pas une sauvegarde. }
    default: false
````

````mdx title="content/netbox/maintain-and-upgrade/page-en.mdx"
{/* First pass — to be validated against docs.netbox.dev (installation/upgrading, administration/housekeeping, integrations/prometheus-metrics) before publishing. */}

An instance nobody maintains is a liability: the day it breaks is the day you find out the backup was never tested. This page sets up a nightly backup and restores it once, upgrades NetBox to a version you chose on purpose, wires the housekeeping job, and leaves a probe behind that tells you when something is wrong.

<Run>

The script installs the backup job on **<V name="NETBOX_HOST" />**, takes a first backup, proves it restores, then upgrades NetBox to **<V name="TARGET_REF" />**. Run it as a sudoer on the NetBox server.

```bash
#!/usr/bin/env bash
set -euo pipefail
# NetBox — backups, housekeeping and upgrade to ${TARGET_REF} on ${NETBOX_HOST}
sudo mkdir -p ${BACKUP_DIR} && sudo chmod 700 ${BACKUP_DIR}

sudo tee /usr/local/sbin/netbox-backup.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%Y%m%d-%H%M)
dest=${BACKUP_DIR}
sudo -u postgres pg_dump -Fc netbox > "$dest/netbox-$stamp.dump"
files=(netbox/media netbox/netbox/configuration.py)
for f in netbox/netbox/ldap_config.py local_requirements.txt gunicorn.py; do
  if [ -e "${INSTALL_DIR}/$f" ]; then files+=("$f"); fi
done
tar -czf "$dest/netbox-files-$stamp.tar.gz" -C ${INSTALL_DIR} "${files[@]}"
find "$dest" -name 'netbox-*' -mtime +${BACKUP_RETENTION_DAYS} -delete
EOF
sudo chmod 750 /usr/local/sbin/netbox-backup.sh
```

<When flag="OFFSITE">

```bash
sudo test -f /root/.ssh/netbox-backup || sudo ssh-keygen -t ed25519 -N '' -f /root/.ssh/netbox-backup -C netbox-backup
sudo tee -a /usr/local/sbin/netbox-backup.sh >/dev/null <<'EOF'
rsync -az --delete -e "ssh -i /root/.ssh/netbox-backup -o StrictHostKeyChecking=accept-new" \
  "$dest/" ${BACKUP_USER}@${BACKUP_HOST}:${BACKUP_PATH}/
EOF
echo "Authorise this key on ${BACKUP_HOST} for ${BACKUP_USER}:" && sudo cat /root/.ssh/netbox-backup.pub
```

</When>

```bash
sudo tee /etc/systemd/system/netbox-backup.service >/dev/null <<'EOF'
[Unit]
Description=NetBox backup (database dump + files)
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/netbox-backup.sh
EOF
sudo tee /etc/systemd/system/netbox-backup.timer >/dev/null <<'EOF'
[Unit]
Description=Nightly NetBox backup
[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now netbox-backup.timer
sudo systemctl start netbox-backup.service

# Restore drill on a scratch database
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres dropdb --if-exists netbox_drill
sudo -u postgres createdb -O netbox netbox_drill
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox_drill < "$latest"
sudo -u postgres psql -d netbox_drill -tAc 'SELECT count(*) FROM dcim_site'
sudo -u postgres dropdb netbox_drill

# Housekeeping and journal size
sudo ln -sf ${INSTALL_DIR}/contrib/netbox-housekeeping.sh /etc/cron.daily/netbox-housekeeping
sudo mkdir -p /etc/systemd/journald.conf.d
printf '[Journal]\nSystemMaxUse=500M\n' | sudo tee /etc/systemd/journald.conf.d/netbox.conf >/dev/null
sudo systemctl restart systemd-journald

# Upgrade
cd ${INSTALL_DIR}
sudo git fetch --tags
sudo git checkout ${TARGET_REF}
sudo ./upgrade.sh
sudo systemctl restart netbox netbox-rq
curl -skf -H "Authorization: Token ${NETBOX_TOKEN}" https://${NETBOX_HOST}/api/status/
```

<Warn>The upgrade part is the one you must not run blind. Read the release notes of <V name="TARGET_REF" /> first: a plugin that lags behind, or a required Python version, is what turns a five-minute upgrade into an evening.</Warn>

</Run>

## Before you start

<Guided>You need sudo on the NetBox server, the API token from the previous page (or any read token), and a look at the [NetBox releases page](https://github.com/netbox-community/netbox/releases) to pick <V name="TARGET_REF" />. Everything on this page runs on the server itself.</Guided>

<Deep>NetBox has exactly three things worth saving: the PostgreSQL database (every object, every change record), the `media` folder (image attachments), and the files you wrote by hand in <V name="INSTALL_DIR" /> (`configuration.py`, `ldap_config.py`, `local_requirements.txt`, `gunicorn.py`). Redis holds only the queue and the cache; the code comes back from git; the venv is rebuilt by `upgrade.sh`. A backup that has those three things restores an instance anywhere.</Deep>

<Check cmd="systemctl is-active netbox netbox-rq postgresql" expect="active
active
active" />

<When is="OS" equals="rhel">

<Note>On the RHEL family, `/etc/cron.daily` needs the `cronie` package to be executed. Check `systemctl is-active crond`; the timer-based alternative in the Housekeeping step avoids the question.</Note>

</When>

## Nightly backups

<Guided>One shell script does the dump and the archive; a systemd service runs it; a timer fires the service at 02:30. Retention is a `find -delete` at the end of the script. The directory is root-only because the archive contains `configuration.py`, and with it the database password and the secret key.</Guided>

```bash
sudo mkdir -p ${BACKUP_DIR} && sudo chmod 700 ${BACKUP_DIR}
```

<Tabs group="backup-files">

<Tab label="netbox-backup.sh">

<Annotated>

```bash title="/usr/local/sbin/netbox-backup.sh" {5,10-11}
#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%Y%m%d-%H%M)
dest=${BACKUP_DIR}
sudo -u postgres pg_dump -Fc netbox > "$dest/netbox-$stamp.dump"        # (1)
files=(netbox/media netbox/netbox/configuration.py)
for f in netbox/netbox/ldap_config.py local_requirements.txt gunicorn.py; do
  if [ -e "${INSTALL_DIR}/$f" ]; then files+=("$f"); fi                  # (2)
done
tar -czf "$dest/netbox-files-$stamp.tar.gz" -C ${INSTALL_DIR} "${files[@]}"
find "$dest" -name 'netbox-*' -mtime +${BACKUP_RETENTION_DAYS} -delete  # (3)
```

1. `-Fc` is PostgreSQL's custom format: compressed, and restorable table by table with `pg_restore`. The plain SQL format is bigger and all-or-nothing. The redirection is done by root, so the `postgres` user never needs write access to the backup directory.
2. `ldap_config.py` only exists with LDAP; `local_requirements.txt` only with plugins. Testing each file keeps the script valid on every instance.
3. Deletes dumps and archives older than <V name="BACKUP_RETENTION_DAYS" /> days. It runs *after* the new backup, so a failed dump never eats into the retention.

</Annotated>

</Tab>

<Tab label="netbox-backup.service">

```ini title="/etc/systemd/system/netbox-backup.service"
[Unit]
Description=NetBox backup (database dump + files)

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/netbox-backup.sh
```

</Tab>

<Tab label="netbox-backup.timer">

```ini title="/etc/systemd/system/netbox-backup.timer" {5-7}
[Unit]
Description=Nightly NetBox backup

[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=15m
Persistent=true

[Install]
WantedBy=timers.target
```

</Tab>

</Tabs>

<Deep>`Persistent=true` is what makes a timer better than cron here: if the server was off at 02:30, the job runs at boot instead of being skipped. `RandomizedDelaySec` spreads the load when several timers share the same minute. The service has no `[Install]` section on purpose: only the timer is enabled, the service is just what it starts. `journalctl -u netbox-backup` shows every run and its output.</Deep>

```bash
sudo chmod 750 /usr/local/sbin/netbox-backup.sh
sudo systemctl daemon-reload
sudo systemctl enable --now netbox-backup.timer
sudo systemctl start netbox-backup.service
```

<Check cmd="systemctl is-active netbox-backup.timer" expect="active" />

<Check cmd="ls ${BACKUP_DIR} | sed 's/-[0-9]*-[0-9]*\././' | LC_ALL=C sort -u" expect="netbox-files.tar.gz
netbox.dump" />

<Details summary="If the service fails">
`journalctl -u netbox-backup -n 20` has the reason. The usual ones: `pg_dump` cannot connect (PostgreSQL is stopped, or `pg_hba.conf` lost its `local all postgres peer` line); `tar` complains about a missing `netbox/media` (the folder was moved, fix the path); no space left in <V name="BACKUP_DIR" />.
</Details>

<When flag="OFFSITE">

### Off-site copy

<Guided>A backup on the same disk as the database only protects you from your own mistakes, not from the disk. The script pushes the folder to <V name="BACKUP_HOST" /> with rsync over SSH, using a key that belongs to root and can do nothing else.</Guided>

```bash
sudo ssh-keygen -t ed25519 -N '' -f /root/.ssh/netbox-backup -C netbox-backup
sudo cat /root/.ssh/netbox-backup.pub
```

Add the printed key to `~${BACKUP_USER}/.ssh/authorized_keys` on <V name="BACKUP_HOST" />, make sure <V name="BACKUP_PATH" /> exists and belongs to <V name="BACKUP_USER" />, then append to the backup script:

```bash title="/usr/local/sbin/netbox-backup.sh (append)"
rsync -az --delete -e "ssh -i /root/.ssh/netbox-backup -o StrictHostKeyChecking=accept-new" \
  "$dest/" ${BACKUP_USER}@${BACKUP_HOST}:${BACKUP_PATH}/
```

<Deep>`--delete` mirrors the folder, so the remote side follows the same retention and never fills up on its own. If you would rather keep more history remotely, drop `--delete` and rotate there. On the backup host, prefix the key in `authorized_keys` with `restrict,command="rrsync /the/destination/path"`: `rrsync` ships with rsync and confines the key to that folder, so a compromised NetBox server cannot read anything else there. `accept-new` records the host key at the first run and refuses a changed one afterwards, which is what you want from an unattended job.</Deep>

<Check cmd="sudo systemctl start netbox-backup.service && sudo ssh -i /root/.ssh/netbox-backup ${BACKUP_USER}@${BACKUP_HOST} ls ${BACKUP_PATH} | grep -q dump && echo synced" expect="synced" />

</When>

## Restore drill

<Guided>A backup you have never restored is a hope, not a backup. We restore the latest dump into a scratch database, count something in it, and drop it. Ten seconds, and now you know the file is usable and you know the commands.</Guided>

```bash
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres createdb -O netbox netbox_drill
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox_drill < "$latest"
```

<Check cmd="sudo -u postgres psql -d netbox_drill -tAc 'SELECT count(*) > 0 FROM dcim_site'" expect="t" />

```bash
sudo -u postgres dropdb netbox_drill
```

<Deep>`--no-owner --role=netbox` makes every restored object belong to the `netbox` role whatever the dump says, which is what lets the same file restore on a server where the roles differ. A real restore is the same three lines against the `netbox` database, with both NetBox services stopped first, then `sudo ./upgrade.sh` if the dump comes from an older NetBox version: migrations only go forward, so the code must be at least as new as the data. Restoring the files is `tar -xzf … -C ${INSTALL_DIR}`.</Deep>

<Warn>Never `pg_restore` into the live `netbox` database while the services run. Stop `netbox` and `netbox-rq` first, or you restore under the feet of the workers and end up with a mix of old and new rows.</Warn>

## Upgrade NetBox

<Guided>An upgrade is: a backup, a `git checkout` of the new tag, `upgrade.sh`, a restart. The script rebuilds the virtual environment, reinstalls the plugins listed in `local_requirements.txt`, runs the migrations and collects the static files, which is why the same script served the install. Read the release notes of <V name="TARGET_REF" /> before anything: they list breaking changes, the Python version required, and what plugins must be updated.</Guided>

```bash
sudo systemctl start netbox-backup.service
cd ${INSTALL_DIR}
sudo git fetch --tags
sudo git checkout ${TARGET_REF}
sudo ./upgrade.sh
sudo systemctl restart netbox netbox-rq
```

<Note>If `git fetch` complains about a shallow clone (the install page cloned with `--depth 1`), run `sudo git fetch --unshallow --tags` once. If `upgrade.sh` picks the wrong Python, prefix it: `sudo PYTHON=/usr/bin/python3.12 ./upgrade.sh`. Both are documented in the upgrading guide at docs.netbox.dev; confirm the flag names there.</Note>

<Deep>NetBox only supports upgrading to a newer version, and only from the latest release of the previous major: going from 3.5 to 4.4 means stopping at 3.7 first. Minor versions inside 4.x can be skipped. `upgrade.sh` removes and recreates `venv/`, so anything you `pip install`ed by hand without listing it in `local_requirements.txt` disappears at this point, which is the usual "plugin vanished after upgrade". The migrations run inside a transaction per app; if one fails, the database is left as it was before that app's migrations and the script exits non-zero, so a failed upgrade is not a corrupted database, it is simply not done.</Deep>

<Check cmd="sudo git -C ${INSTALL_DIR} describe --tags --exact-match" expect="${TARGET_REF}" />

<Check cmd={"curl -skf -H 'Authorization: Token ${NETBOX_TOKEN}' https://${NETBOX_HOST}/api/status/ | python3 -c 'import sys,json; print(\"v\" + json.load(sys.stdin)[\"netbox-version\"])'"} expect="${TARGET_REF}" />

<Check cmd={"curl -sk -o /dev/null -w '%{http_code}' https://${NETBOX_HOST}/login/"} expect="200" />

Then open **https://<V name="NETBOX_HOST" />/** and click through a device page and a search: the API answering is not the same as the UI rendering, since static files and plugins can break one and not the other.

### If a plugin blocks the upgrade

<Guided>The most common failure: `upgrade.sh` stops while installing `local_requirements.txt`, or during migrations, with an error that names a plugin. The plugin does not support the new NetBox yet, or it does but only in a version pip did not pick.</Guided>

- Check the plugin's compatibility table on its repository. If a compatible release exists, pin it in `local_requirements.txt` (`netbox-bgp==0.15.0`) and re-run `sudo ./upgrade.sh`.
- If none exists, remove the plugin from `PLUGINS` in `configuration.py` *and* from `local_requirements.txt`, upgrade, and put it back when the plugin catches up. Its tables stay in the database; nothing is lost, the pages just disappear until then.
- If the release notes say the plugin is now a core feature (it happens), migrate the data with the plugin's own instructions *before* upgrading.

<Deep>Do not `pip install` a plugin from its git `main` branch to work around it: you would be running unreleased code in your source of truth, and the next `upgrade.sh` would drop it anyway. When a plugin is essential and lags, the right move is to delay the NetBox upgrade, not to force it.</Deep>

### Rollback

<Guided>Code goes back with git; the database goes back with the dump you took right before. The two together, in that order, and you are exactly where you started.</Guided>

```bash
sudo systemctl stop netbox netbox-rq
cd ${INSTALL_DIR}
sudo git checkout $(sudo git describe --tags --abbrev=0 ${TARGET_REF}^)
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres dropdb netbox && sudo -u postgres createdb -O netbox netbox
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox < "$latest"
sudo ./upgrade.sh
sudo systemctl start netbox netbox-rq
```

<Warn>`dropdb netbox` deletes the live database. Only do this with a dump you have restored in the drill above, taken *before* the upgrade. The `git describe` line finds the tag just before <V name="TARGET_REF" />; check what it prints if you had skipped versions, and use your previous tag explicitly instead.</Warn>

<Deep>Why restore the database at all, if the code is back? Because migrations are one-way in practice: NetBox does not test reverse migrations, and a new column with data has no honest way back. The dump is the rollback. This is also why the backup runs right before the checkout in the upgrade step, not the night before.</Deep>

## Housekeeping

<Guided>NetBox ships a management command, `manage.py housekeeping`, that deletes expired sessions, change records older than `CHANGELOG_RETENTION` and job results older than `JOB_RETENTION`. The install page linked the shipped wrapper into `/etc/cron.daily`; check it is still there after the upgrade, and run it once by hand to see what it does.</Guided>

```bash
sudo ln -sf ${INSTALL_DIR}/contrib/netbox-housekeeping.sh /etc/cron.daily/netbox-housekeeping
sudo ${INSTALL_DIR}/venv/bin/python ${INSTALL_DIR}/netbox/manage.py housekeeping
```

<Check cmd="run-parts --test /etc/cron.daily | grep netbox" expect="/etc/cron.daily/netbox-housekeeping" />

<Deep>Prefer a timer over cron? Copy the two backup units, replace the script with `${INSTALL_DIR}/contrib/netbox-housekeeping.sh` and `OnCalendar=daily`, and remove the symlink. Same result, plus a journal per run. `CHANGELOG_RETENTION` is the number of days of change history kept (180 on the previous page, 90 by default, 0 for forever); the change log is the biggest table on any instance with automation writing to it, so this is the single setting that decides the size of your dumps.</Deep>

Logs and disk: gunicorn writes to stdout, which systemd captures in the journal; cap it. nginx and Apache rotate their own files through `logrotate`, shipped with the package.

<Tabs group="logs">

<Tab label="journald.conf.d/netbox.conf">

```ini title="/etc/systemd/journald.conf.d/netbox.conf"
[Journal]
SystemMaxUse=500M
```

Then `sudo systemctl restart systemd-journald`.

</Tab>

<Tab label="checks">

```bash
journalctl --disk-usage
df -h / ${BACKUP_DIR}
sudo -u postgres psql -tAc "SELECT pg_size_pretty(pg_database_size('netbox'))"
redis-cli info memory | grep used_memory_human
```

</Tab>

</Tabs>

<Deep>Redis holds the RQ queue and the cache; `used_memory_human` above a few hundred MB means either a stuck queue (webhooks piling up because a receiver is down, see `manage.py rqworker` logs in `journalctl -u netbox-rq`) or a cache never trimmed. Set `maxmemory 256mb` and `maxmemory-policy allkeys-lru` in `redis.conf` if the second: NetBox tolerates a cold cache, not a Redis that refuses writes. Confirm the parameter names in the Redis documentation for your packaged version.</Deep>

<When is="WEB" equals="nginx">

<Note>`logrotate -d /etc/logrotate.d/nginx` shows what would be rotated without doing it. The default is daily, 14 files, compressed: fine.</Note>

</When>

<When is="WEB" equals="apache">

<Note>`logrotate -d /etc/logrotate.d/apache2` (Ubuntu) or `/etc/logrotate.d/httpd` (RHEL) shows what would be rotated without doing it. The default is weekly and compressed: fine.</Note>

</When>

Weekly, five minutes:

| Check | Command | Expected |
|---|---|---|
| Both services up | `systemctl is-active netbox netbox-rq` | `active` twice |
| Last backup ran | `systemctl list-timers netbox-backup.timer` | a `LAST` within 24h |
| Backup size sane | `ls -lh ${BACKUP_DIR}` | sizes in the same range as last week |
| Disk | `df -h /` | under 80% |
| Queue quiet | `redis-cli info memory` | `used_memory_human` stable |
| New release? | releases page | read the notes, plan the upgrade |

## Monitoring basics

<Guided>Two signals are enough to start: the systemd state of both units, and the JSON answer of `/api/status/`, which reports the version and, more usefully, how many RQ workers are running. A probe that checks the second one every minute catches the failures users notice.</Guided>

```bash
systemctl status netbox netbox-rq --no-pager | grep Active
curl -skf -H "Authorization: Token ${NETBOX_TOKEN}" https://${NETBOX_HOST}/api/status/ | python3 -m json.tool
```

<Check cmd={"curl -skf -H 'Authorization: Token ${NETBOX_TOKEN}' https://${NETBOX_HOST}/api/status/ | python3 -c 'import sys,json; print(json.load(sys.stdin)[\"rq-workers-running\"] >= 1)'"} expect="True" />

<Guided>For a real probe (Zabbix, Nagios, Uptime Kuma, a cron with `curl -f`), create a dedicated user with view rights on nothing in particular and give the probe *its* token, not yours; the next page shows how to create such users and tokens. `rq-workers-running` at 0 means webhooks, scripts and reports silently stop while the UI keeps working: it is the one to alert on.</Guided>

<Deep>Prometheus: set `METRICS_ENABLED = True` in `configuration.py` and NetBox exposes `/metrics` (django-prometheus: request latencies, database and cache counters). With several gunicorn workers each process has its own counters, so the exporter needs a shared directory: add a drop-in for the `netbox` unit with `RuntimeDirectory=netbox-metrics` and `Environment=PROMETHEUS_MULTIPROC_DIR=/run/netbox-metrics`, then `daemon-reload` and restart. `/metrics` is exempt from `LOGIN_REQUIRED`, so restrict it to the Prometheus address in the web server (`location = /metrics` with `allow`/`deny` in nginx, `Require ip` in Apache). Confirm the environment variable's spelling in the "Prometheus metrics" page of the NetBox docs; older versions used the lowercase name.</Deep>

## Done

Every night a dump and an archive land in <V name="BACKUP_DIR" /><When flag="OFFSITE"> and on <V name="BACKUP_HOST" /></When>, you have restored one by hand, NetBox runs <V name="TARGET_REF" />, the change log is pruned daily, and a probe knows when the workers stop. The next page stops clicking: an automation user, pynetbox, bulk imports, webhooks and custom scripts.
````

````mdx title="content/netbox/maintain-and-upgrade/page-fr.mdx"
{/* Première passe — à valider contre docs.netbox.dev (installation/upgrading, administration/housekeeping, integrations/prometheus-metrics) avant publication. */}

Une instance que personne ne maintient est une bombe à retardement : le jour où elle casse est le jour où tu découvres que la sauvegarde n'a jamais été testée. Cette page met en place une sauvegarde nocturne et la restaure une fois, met NetBox à jour vers une version choisie exprès, branche la tâche d'entretien, et laisse derrière elle une sonde qui te prévient quand quelque chose cloche.

<Run>

Le script installe la tâche de sauvegarde sur **<V name="NETBOX_HOST" />**, prend une première sauvegarde, prouve qu'elle se restaure, puis met NetBox à jour vers **<V name="TARGET_REF" />**. Lance-le en sudoer sur le serveur NetBox.

```bash
#!/usr/bin/env bash
set -euo pipefail
# NetBox — sauvegardes, entretien et mise à jour vers ${TARGET_REF} sur ${NETBOX_HOST}
sudo mkdir -p ${BACKUP_DIR} && sudo chmod 700 ${BACKUP_DIR}

sudo tee /usr/local/sbin/netbox-backup.sh >/dev/null <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%Y%m%d-%H%M)
dest=${BACKUP_DIR}
sudo -u postgres pg_dump -Fc netbox > "$dest/netbox-$stamp.dump"
files=(netbox/media netbox/netbox/configuration.py)
for f in netbox/netbox/ldap_config.py local_requirements.txt gunicorn.py; do
  if [ -e "${INSTALL_DIR}/$f" ]; then files+=("$f"); fi
done
tar -czf "$dest/netbox-files-$stamp.tar.gz" -C ${INSTALL_DIR} "${files[@]}"
find "$dest" -name 'netbox-*' -mtime +${BACKUP_RETENTION_DAYS} -delete
EOF
sudo chmod 750 /usr/local/sbin/netbox-backup.sh
```

<When flag="OFFSITE">

```bash
sudo test -f /root/.ssh/netbox-backup || sudo ssh-keygen -t ed25519 -N '' -f /root/.ssh/netbox-backup -C netbox-backup
sudo tee -a /usr/local/sbin/netbox-backup.sh >/dev/null <<'EOF'
rsync -az --delete -e "ssh -i /root/.ssh/netbox-backup -o StrictHostKeyChecking=accept-new" \
  "$dest/" ${BACKUP_USER}@${BACKUP_HOST}:${BACKUP_PATH}/
EOF
echo "Autorise cette clé sur ${BACKUP_HOST} pour ${BACKUP_USER} :" && sudo cat /root/.ssh/netbox-backup.pub
```

</When>

```bash
sudo tee /etc/systemd/system/netbox-backup.service >/dev/null <<'EOF'
[Unit]
Description=NetBox backup (database dump + files)
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/netbox-backup.sh
EOF
sudo tee /etc/systemd/system/netbox-backup.timer >/dev/null <<'EOF'
[Unit]
Description=Nightly NetBox backup
[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
EOF
sudo systemctl daemon-reload
sudo systemctl enable --now netbox-backup.timer
sudo systemctl start netbox-backup.service

# Exercice de restauration sur une base jetable
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres dropdb --if-exists netbox_drill
sudo -u postgres createdb -O netbox netbox_drill
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox_drill < "$latest"
sudo -u postgres psql -d netbox_drill -tAc 'SELECT count(*) FROM dcim_site'
sudo -u postgres dropdb netbox_drill

# Entretien et taille du journal
sudo ln -sf ${INSTALL_DIR}/contrib/netbox-housekeeping.sh /etc/cron.daily/netbox-housekeeping
sudo mkdir -p /etc/systemd/journald.conf.d
printf '[Journal]\nSystemMaxUse=500M\n' | sudo tee /etc/systemd/journald.conf.d/netbox.conf >/dev/null
sudo systemctl restart systemd-journald

# Mise à jour
cd ${INSTALL_DIR}
sudo git fetch --tags
sudo git checkout ${TARGET_REF}
sudo ./upgrade.sh
sudo systemctl restart netbox netbox-rq
curl -skf -H "Authorization: Token ${NETBOX_TOKEN}" https://${NETBOX_HOST}/api/status/
```

<Warn>La partie mise à jour est celle qu'il ne faut pas lancer à l'aveugle. Lis d'abord les notes de version de <V name="TARGET_REF" /> : un plugin en retard, ou une version de Python exigée, c'est ce qui transforme une mise à jour de cinq minutes en soirée.</Warn>

</Run>

## Avant de commencer

<Guided>Il te faut sudo sur le serveur NetBox, le jeton d'API de la page précédente (ou n'importe quel jeton en lecture), et un passage par la [page des releases de NetBox](https://github.com/netbox-community/netbox/releases) pour choisir <V name="TARGET_REF" />. Tout ce qui est sur cette page s'exécute sur le serveur lui-même.</Guided>

<Deep>NetBox a exactement trois choses qui méritent d'être sauvegardées : la base PostgreSQL (chaque objet, chaque enregistrement de changement), le dossier `media` (les images attachées), et les fichiers que tu as écrits à la main dans <V name="INSTALL_DIR" /> (`configuration.py`, `ldap_config.py`, `local_requirements.txt`, `gunicorn.py`). Redis ne contient que la file et le cache ; le code revient de git ; le venv est reconstruit par `upgrade.sh`. Une sauvegarde qui contient ces trois choses restaure une instance n'importe où.</Deep>

<Check cmd="systemctl is-active netbox netbox-rq postgresql" expect="active
active
active" />

<When is="OS" equals="rhel">

<Note>Sur la famille RHEL, `/etc/cron.daily` n'est exécuté que si le paquet `cronie` est installé. Vérifie `systemctl is-active crond` ; l'alternative par timer dans l'étape Entretien évite la question.</Note>

</When>

## Sauvegardes nocturnes

<Guided>Un script shell fait le dump et l'archive ; un service systemd l'exécute ; un timer déclenche le service à 02 h 30. La rétention est un `find -delete` à la fin du script. Le répertoire est réservé à root parce que l'archive contient `configuration.py`, et avec lui le mot de passe de la base et la clé secrète.</Guided>

```bash
sudo mkdir -p ${BACKUP_DIR} && sudo chmod 700 ${BACKUP_DIR}
```

<Tabs group="backup-files">

<Tab label="netbox-backup.sh">

<Annotated>

```bash title="/usr/local/sbin/netbox-backup.sh" {5,10-11}
#!/usr/bin/env bash
set -euo pipefail
stamp=$(date +%Y%m%d-%H%M)
dest=${BACKUP_DIR}
sudo -u postgres pg_dump -Fc netbox > "$dest/netbox-$stamp.dump"        # (1)
files=(netbox/media netbox/netbox/configuration.py)
for f in netbox/netbox/ldap_config.py local_requirements.txt gunicorn.py; do
  if [ -e "${INSTALL_DIR}/$f" ]; then files+=("$f"); fi                  # (2)
done
tar -czf "$dest/netbox-files-$stamp.tar.gz" -C ${INSTALL_DIR} "${files[@]}"
find "$dest" -name 'netbox-*' -mtime +${BACKUP_RETENTION_DAYS} -delete  # (3)
```

1. `-Fc` est le format « custom » de PostgreSQL : compressé, et restaurable table par table avec `pg_restore`. Le format SQL brut est plus gros et tout-ou-rien. La redirection est faite par root, donc l'utilisateur `postgres` n'a jamais besoin d'écrire dans le répertoire de sauvegarde.
2. `ldap_config.py` n'existe qu'avec LDAP ; `local_requirements.txt` qu'avec des plugins. Tester chaque fichier garde le script valable sur toutes les instances.
3. Supprime les dumps et archives de plus de <V name="BACKUP_RETENTION_DAYS" /> jours. Ça tourne *après* la nouvelle sauvegarde, donc un dump raté n'entame jamais la rétention.

</Annotated>

</Tab>

<Tab label="netbox-backup.service">

```ini title="/etc/systemd/system/netbox-backup.service"
[Unit]
Description=NetBox backup (database dump + files)

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/netbox-backup.sh
```

</Tab>

<Tab label="netbox-backup.timer">

```ini title="/etc/systemd/system/netbox-backup.timer" {5-7}
[Unit]
Description=Nightly NetBox backup

[Timer]
OnCalendar=*-*-* 02:30:00
RandomizedDelaySec=15m
Persistent=true

[Install]
WantedBy=timers.target
```

</Tab>

</Tabs>

<Deep>`Persistent=true` est ce qui rend un timer meilleur que cron ici : si le serveur était éteint à 02 h 30, la tâche tourne au démarrage au lieu d'être sautée. `RandomizedDelaySec` étale la charge quand plusieurs timers partagent la même minute. Le service n'a pas de section `[Install]`, c'est voulu : seul le timer est activé, le service est juste ce qu'il démarre. `journalctl -u netbox-backup` montre chaque exécution et sa sortie.</Deep>

```bash
sudo chmod 750 /usr/local/sbin/netbox-backup.sh
sudo systemctl daemon-reload
sudo systemctl enable --now netbox-backup.timer
sudo systemctl start netbox-backup.service
```

<Check cmd="systemctl is-active netbox-backup.timer" expect="active" />

<Check cmd="ls ${BACKUP_DIR} | sed 's/-[0-9]*-[0-9]*\././' | LC_ALL=C sort -u" expect="netbox-files.tar.gz
netbox.dump" />

<Details summary="Si le service échoue">
`journalctl -u netbox-backup -n 20` donne la raison. Les habituelles : `pg_dump` ne peut pas se connecter (PostgreSQL est arrêté, ou `pg_hba.conf` a perdu sa ligne `local all postgres peer`) ; `tar` se plaint d'un `netbox/media` manquant (le dossier a été déplacé, corrige le chemin) ; plus de place dans <V name="BACKUP_DIR" />.
</Details>

<When flag="OFFSITE">

### Copie distante

<Guided>Une sauvegarde sur le même disque que la base ne te protège que de tes propres erreurs, pas du disque. Le script pousse le dossier vers <V name="BACKUP_HOST" /> avec rsync en SSH, avec une clé qui appartient à root et ne sait rien faire d'autre.</Guided>

```bash
sudo ssh-keygen -t ed25519 -N '' -f /root/.ssh/netbox-backup -C netbox-backup
sudo cat /root/.ssh/netbox-backup.pub
```

Ajoute la clé affichée dans `~${BACKUP_USER}/.ssh/authorized_keys` sur <V name="BACKUP_HOST" />, vérifie que <V name="BACKUP_PATH" /> existe et appartient à <V name="BACKUP_USER" />, puis ajoute à la fin du script de sauvegarde :

```bash title="/usr/local/sbin/netbox-backup.sh (à la fin)"
rsync -az --delete -e "ssh -i /root/.ssh/netbox-backup -o StrictHostKeyChecking=accept-new" \
  "$dest/" ${BACKUP_USER}@${BACKUP_HOST}:${BACKUP_PATH}/
```

<Deep>`--delete` reflète le dossier, donc le côté distant suit la même rétention et ne se remplit jamais tout seul. Si tu préfères garder plus d'historique à distance, retire `--delete` et fais la rotation là-bas. Sur le serveur de sauvegarde, préfixe la clé dans `authorized_keys` avec `restrict,command="rrsync /le/chemin/de/destination"` : `rrsync` est livré avec rsync et confine la clé à ce dossier, donc un serveur NetBox compromis ne peut rien lire d'autre là-bas. `accept-new` enregistre la clé d'hôte à la première exécution et refuse une clé changée ensuite, exactement ce qu'on attend d'une tâche sans surveillance.</Deep>

<Check cmd="sudo systemctl start netbox-backup.service && sudo ssh -i /root/.ssh/netbox-backup ${BACKUP_USER}@${BACKUP_HOST} ls ${BACKUP_PATH} | grep -q dump && echo synced" expect="synced" />

</When>

## Exercice de restauration

<Guided>Une sauvegarde jamais restaurée est un espoir, pas une sauvegarde. On restaure le dernier dump dans une base jetable, on y compte quelque chose, et on la supprime. Dix secondes, et maintenant tu sais que le fichier est utilisable et tu connais les commandes.</Guided>

```bash
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres createdb -O netbox netbox_drill
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox_drill < "$latest"
```

<Check cmd="sudo -u postgres psql -d netbox_drill -tAc 'SELECT count(*) > 0 FROM dcim_site'" expect="t" />

```bash
sudo -u postgres dropdb netbox_drill
```

<Deep>`--no-owner --role=netbox` fait appartenir chaque objet restauré au rôle `netbox` quoi que dise le dump, ce qui permet au même fichier de se restaurer sur un serveur où les rôles diffèrent. Une vraie restauration, ce sont les mêmes trois lignes vers la base `netbox`, les deux services NetBox arrêtés avant, puis `sudo ./upgrade.sh` si le dump vient d'une version plus ancienne de NetBox : les migrations ne vont que vers l'avant, donc le code doit être au moins aussi récent que les données. Restaurer les fichiers, c'est `tar -xzf … -C ${INSTALL_DIR}`.</Deep>

<Warn>Ne fais jamais de `pg_restore` dans la base `netbox` en production pendant que les services tournent. Arrête `netbox` et `netbox-rq` d'abord, sinon tu restaures sous les pieds des workers et tu te retrouves avec un mélange de lignes anciennes et nouvelles.</Warn>

## Mettre NetBox à jour

<Guided>Une mise à jour, c'est : une sauvegarde, un `git checkout` du nouveau tag, `upgrade.sh`, un redémarrage. Le script reconstruit l'environnement virtuel, réinstalle les plugins listés dans `local_requirements.txt`, exécute les migrations et collecte les fichiers statiques, raison pour laquelle le même script a servi à l'installation. Lis les notes de version de <V name="TARGET_REF" /> avant tout : elles listent les changements cassants, la version de Python exigée, et les plugins à mettre à jour.</Guided>

```bash
sudo systemctl start netbox-backup.service
cd ${INSTALL_DIR}
sudo git fetch --tags
sudo git checkout ${TARGET_REF}
sudo ./upgrade.sh
sudo systemctl restart netbox netbox-rq
```

<Note>Si `git fetch` se plaint d'un clone superficiel (la page d'installation a cloné avec `--depth 1`), lance `sudo git fetch --unshallow --tags` une fois. Si `upgrade.sh` choisit le mauvais Python, préfixe-le : `sudo PYTHON=/usr/bin/python3.12 ./upgrade.sh`. Les deux sont documentés dans le guide de mise à jour sur docs.netbox.dev ; confirme les noms des options là-bas.</Note>

<Deep>NetBox ne supporte que la montée vers une version plus récente, et seulement depuis la dernière release de la majeure précédente : passer de 3.5 à 4.4 impose une étape en 3.7. Les versions mineures à l'intérieur de 4.x peuvent être sautées. `upgrade.sh` supprime et recrée `venv/`, donc tout ce que tu as installé avec `pip` à la main sans le lister dans `local_requirements.txt` disparaît à ce moment-là, c'est le classique « le plugin a disparu après la mise à jour ». Les migrations tournent dans une transaction par application ; si l'une échoue, la base reste dans l'état d'avant les migrations de cette application et le script sort en erreur, donc une mise à jour ratée n'est pas une base corrompue, elle est simplement pas faite.</Deep>

<Check cmd="sudo git -C ${INSTALL_DIR} describe --tags --exact-match" expect="${TARGET_REF}" />

<Check cmd={"curl -skf -H 'Authorization: Token ${NETBOX_TOKEN}' https://${NETBOX_HOST}/api/status/ | python3 -c 'import sys,json; print(\"v\" + json.load(sys.stdin)[\"netbox-version\"])'"} expect="${TARGET_REF}" />

<Check cmd={"curl -sk -o /dev/null -w '%{http_code}' https://${NETBOX_HOST}/login/"} expect="200" />

Ensuite, ouvre **https://<V name="NETBOX_HOST" />/** et clique sur une page d'équipement et une recherche : que l'API réponde ne veut pas dire que l'interface s'affiche, puisque les fichiers statiques et les plugins peuvent casser l'un sans l'autre.

### Si un plugin bloque la mise à jour

<Guided>La panne la plus fréquente : `upgrade.sh` s'arrête pendant l'installation de `local_requirements.txt`, ou pendant les migrations, avec une erreur qui nomme un plugin. Le plugin ne supporte pas encore le nouveau NetBox, ou il le supporte mais seulement dans une version que pip n'a pas choisie.</Guided>

- Regarde le tableau de compatibilité du plugin sur son dépôt. S'il existe une release compatible, fige-la dans `local_requirements.txt` (`netbox-bgp==0.15.0`) et relance `sudo ./upgrade.sh`.
- S'il n'en existe pas, retire le plugin de `PLUGINS` dans `configuration.py` *et* de `local_requirements.txt`, mets à jour, et remets-le quand le plugin aura rattrapé. Ses tables restent dans la base ; rien n'est perdu, les pages disparaissent juste en attendant.
- Si les notes de version disent que le plugin est devenu une fonction native (ça arrive), migre les données avec les instructions du plugin *avant* la mise à jour.

<Deep>Ne fais pas de `pip install` d'un plugin depuis sa branche git `main` pour contourner : tu ferais tourner du code non publié dans ta source de vérité, et le prochain `upgrade.sh` le jetterait de toute façon. Quand un plugin est indispensable et en retard, la bonne décision est de retarder la mise à jour de NetBox, pas de la forcer.</Deep>

### Retour arrière

<Guided>Le code revient avec git ; la base revient avec le dump pris juste avant. Les deux ensemble, dans cet ordre, et tu es exactement là où tu étais.</Guided>

```bash
sudo systemctl stop netbox netbox-rq
cd ${INSTALL_DIR}
sudo git checkout $(sudo git describe --tags --abbrev=0 ${TARGET_REF}^)
latest=$(ls -t ${BACKUP_DIR}/netbox-*.dump | head -1)
sudo -u postgres dropdb netbox && sudo -u postgres createdb -O netbox netbox
sudo -u postgres pg_restore --no-owner --role=netbox -d netbox < "$latest"
sudo ./upgrade.sh
sudo systemctl start netbox netbox-rq
```

<Warn>`dropdb netbox` supprime la base en production. Ne le fais qu'avec un dump que tu as restauré dans l'exercice ci-dessus, pris *avant* la mise à jour. La ligne `git describe` trouve le tag juste avant <V name="TARGET_REF" /> ; vérifie ce qu'elle affiche si tu avais sauté des versions, et utilise ton tag précédent explicitement à la place.</Warn>

<Deep>Pourquoi restaurer la base, si le code est revenu ? Parce que les migrations sont à sens unique en pratique : NetBox ne teste pas les migrations inverses, et une nouvelle colonne avec des données n'a pas de chemin honnête vers l'arrière. Le dump est le retour arrière. C'est aussi pour ça que la sauvegarde tourne juste avant le checkout dans l'étape de mise à jour, et pas la nuit d'avant.</Deep>

## Entretien

<Guided>NetBox livre une commande de gestion, `manage.py housekeeping`, qui supprime les sessions expirées, les enregistrements de changement plus vieux que `CHANGELOG_RETENTION` et les résultats de tâches plus vieux que `JOB_RETENTION`. La page d'installation a lié le wrapper livré dans `/etc/cron.daily` ; vérifie qu'il est toujours là après la mise à jour, et lance-le une fois à la main pour voir ce qu'il fait.</Guided>

```bash
sudo ln -sf ${INSTALL_DIR}/contrib/netbox-housekeeping.sh /etc/cron.daily/netbox-housekeeping
sudo ${INSTALL_DIR}/venv/bin/python ${INSTALL_DIR}/netbox/manage.py housekeeping
```

<Check cmd="run-parts --test /etc/cron.daily | grep netbox" expect="/etc/cron.daily/netbox-housekeeping" />

<Deep>Tu préfères un timer à cron ? Copie les deux unités de sauvegarde, remplace le script par `${INSTALL_DIR}/contrib/netbox-housekeeping.sh` et mets `OnCalendar=daily`, puis retire le lien symbolique. Même résultat, plus un journal par exécution. `CHANGELOG_RETENTION` est le nombre de jours d'historique conservés (180 à la page précédente, 90 par défaut, 0 pour toujours) ; le journal des changements est la plus grosse table de toute instance où de l'automatisation écrit, donc c'est le réglage qui décide de la taille de tes dumps.</Deep>

Logs et disque : gunicorn écrit sur la sortie standard, que systemd capture dans le journal ; plafonne-le. nginx et Apache font tourner leurs propres fichiers via `logrotate`, livré avec le paquet.

<Tabs group="logs">

<Tab label="journald.conf.d/netbox.conf">

```ini title="/etc/systemd/journald.conf.d/netbox.conf"
[Journal]
SystemMaxUse=500M
```

Puis `sudo systemctl restart systemd-journald`.

</Tab>

<Tab label="vérifications">

```bash
journalctl --disk-usage
df -h / ${BACKUP_DIR}
sudo -u postgres psql -tAc "SELECT pg_size_pretty(pg_database_size('netbox'))"
redis-cli info memory | grep used_memory_human
```

</Tab>

</Tabs>

<Deep>Redis contient la file RQ et le cache ; un `used_memory_human` au-dessus de quelques centaines de Mo signifie soit une file bloquée (des webhooks qui s'empilent parce qu'un récepteur est tombé, voir les logs de `manage.py rqworker` dans `journalctl -u netbox-rq`), soit un cache jamais purgé. Mets `maxmemory 256mb` et `maxmemory-policy allkeys-lru` dans `redis.conf` dans le second cas : NetBox tolère un cache froid, pas un Redis qui refuse d'écrire. Confirme les noms des paramètres dans la documentation Redis de la version de ton paquet.</Deep>

<When is="WEB" equals="nginx">

<Note>`logrotate -d /etc/logrotate.d/nginx` montre ce qui serait tourné sans le faire. Par défaut : quotidien, 14 fichiers, compressés. Très bien.</Note>

</When>

<When is="WEB" equals="apache">

<Note>`logrotate -d /etc/logrotate.d/apache2` (Ubuntu) ou `/etc/logrotate.d/httpd` (RHEL) montre ce qui serait tourné sans le faire. Par défaut : hebdomadaire et compressé. Très bien.</Note>

</When>

Chaque semaine, cinq minutes :

| Vérification | Commande | Attendu |
|---|---|---|
| Les deux services tournent | `systemctl is-active netbox netbox-rq` | `active` deux fois |
| La dernière sauvegarde a tourné | `systemctl list-timers netbox-backup.timer` | un `LAST` de moins de 24 h |
| Taille de sauvegarde cohérente | `ls -lh ${BACKUP_DIR}` | des tailles du même ordre que la semaine dernière |
| Disque | `df -h /` | sous 80 % |
| File calme | `redis-cli info memory` | `used_memory_human` stable |
| Nouvelle release ? | page des releases | lire les notes, planifier la mise à jour |

## Supervision de base

<Guided>Deux signaux suffisent pour commencer : l'état systemd des deux unités, et la réponse JSON de `/api/status/`, qui donne la version et, plus utile, le nombre de workers RQ en marche. Une sonde qui vérifie le second toutes les minutes attrape les pannes que les utilisateurs remarquent.</Guided>

```bash
systemctl status netbox netbox-rq --no-pager | grep Active
curl -skf -H "Authorization: Token ${NETBOX_TOKEN}" https://${NETBOX_HOST}/api/status/ | python3 -m json.tool
```

<Check cmd={"curl -skf -H 'Authorization: Token ${NETBOX_TOKEN}' https://${NETBOX_HOST}/api/status/ | python3 -c 'import sys,json; print(json.load(sys.stdin)[\"rq-workers-running\"] >= 1)'"} expect="True" />

<Guided>Pour une vraie sonde (Zabbix, Nagios, Uptime Kuma, un cron avec `curl -f`), crée un utilisateur dédié avec des droits de lecture sur rien en particulier et donne à la sonde *son* jeton, pas le tien ; la page suivante montre comment créer ces utilisateurs et jetons. `rq-workers-running` à 0 veut dire que webhooks, scripts et rapports s'arrêtent en silence pendant que l'interface continue de marcher : c'est celui sur lequel alerter.</Guided>

<Deep>Prometheus : mets `METRICS_ENABLED = True` dans `configuration.py` et NetBox expose `/metrics` (django-prometheus : latences des requêtes, compteurs base et cache). Avec plusieurs workers gunicorn, chaque processus a ses propres compteurs, donc l'exporteur a besoin d'un répertoire partagé : ajoute un drop-in à l'unité `netbox` avec `RuntimeDirectory=netbox-metrics` et `Environment=PROMETHEUS_MULTIPROC_DIR=/run/netbox-metrics`, puis `daemon-reload` et redémarre. `/metrics` est exempté de `LOGIN_REQUIRED`, donc restreins-le à l'adresse de Prometheus dans le serveur web (`location = /metrics` avec `allow`/`deny` sous nginx, `Require ip` sous Apache). Confirme l'orthographe de la variable d'environnement dans la page « Prometheus metrics » de la doc NetBox ; les anciennes versions utilisaient le nom en minuscules.</Deep>

## Terminé

Chaque nuit un dump et une archive atterrissent dans <V name="BACKUP_DIR" /><When flag="OFFSITE"> et sur <V name="BACKUP_HOST" /></When>, tu en as restauré un à la main, NetBox tourne en <V name="TARGET_REF" />, le journal des changements est purgé chaque jour, et une sonde sait quand les workers s'arrêtent. La page suivante arrête de cliquer : un utilisateur d'automatisation, pynetbox, des imports en masse, des webhooks et des scripts personnalisés.
````

````yaml title="content/netbox/maintain-and-upgrade/diagram.yaml"
# Two flows on one server: the nightly backup (timer → script → dumps → off-site)
# and the upgrade path (tags → upgrade.sh → services → probe). Quick shows the
# two chains; Guided adds what the backup reads and the housekeeping job; Deep
# adds the journal, the logs and every path and unit.
# `desc` is what the reader gets when hovering a box or an arrow.
title: { en: "What runs every night, and what you run on purpose", fr: "Ce qui tourne chaque nuit, et ce que tu lances exprès" }
caption:
  en: "A timer dumps the database and archives the files into the backup directory; with the off-site option they are pushed elsewhere. An upgrade is a tag checked out from git, applied by upgrade.sh, verified by a probe on /api/status/."
  fr: "Un timer dumpe la base et archive les fichiers dans le répertoire de sauvegarde ; avec l'option distante ils sont poussés ailleurs. Une mise à jour, c'est un tag extrait de git, appliqué par upgrade.sh, vérifié par une sonde sur /api/status/."

groups:
  - id: timers
    label: { en: "Scheduled on the server", fr: "Planifié sur le serveur" }
    desc:
      en: "What systemd and cron start on their own on the NetBox server, with nobody logged in."
      fr: "Ce que systemd et cron lancent tout seuls sur le serveur NetBox, sans personne de connecté."
  - id: server
    label: { en: "NetBox server", fr: "Serveur NetBox" }
    desc:
      en: "The instance from the first page. Everything inside runs as root from systemd, except the probe."
      fr: "L'instance de la première page. Tout ce qui est dedans tourne en root via systemd, sauf la sonde."

nodes:
  - id: timer
    kind: service
    label: { en: "Nightly timer", fr: "Timer nocturne" }
    sub: "02:30 · netbox-backup.timer"
    in: timers
    desc:
      en: "A systemd timer fires the backup service at 02:30. Persistent: a run missed while the server was off happens at boot."
      fr: "Un timer systemd déclenche le service de sauvegarde à 02 h 30. Persistant : une exécution ratée pendant que le serveur était éteint a lieu au démarrage."
    deep:
      sub: "netbox-backup.timer · OnCalendar 02:30 · Persistent=true → netbox-backup.service"
  - id: backup
    kind: service
    label: { en: "Backup script", fr: "Script de sauvegarde" }
    sub: "pg_dump -Fc · tar"
    in: server
    focus: true
    desc:
      en: "One shell script: a custom-format dump of the netbox database, an archive of media and the config files, then retention of ${BACKUP_RETENTION_DAYS} days."
      fr: "Un script shell : un dump au format custom de la base netbox, une archive de media et des fichiers de config, puis rétention de ${BACKUP_RETENTION_DAYS} jours."
    deep:
      sub: "/usr/local/sbin/netbox-backup.sh · pg_dump -Fc · tar -czf · find -mtime +${BACKUP_RETENTION_DAYS} -delete"
  - id: postgres
    kind: store
    label: { en: "PostgreSQL", fr: "PostgreSQL" }
    sub: "netbox database"
    in: server
    level: guided
    desc:
      en: "Every object and every change record. The dump is the backup; the restore drill proves it on a scratch database."
      fr: "Chaque objet et chaque enregistrement de changement. Le dump est la sauvegarde ; l'exercice de restauration le prouve sur une base jetable."
    deep:
      sub: "db netbox · role netbox · dump via unix socket as postgres"
  - id: files
    kind: file
    label: { en: "Files", fr: "Fichiers" }
    sub: "media/ · configuration.py"
    in: server
    level: guided
    desc:
      en: "Image attachments and the files you wrote by hand. Code and venv are not saved: git and upgrade.sh bring them back."
      fr: "Les images attachées et les fichiers écrits à la main. Code et venv ne sont pas sauvegardés : git et upgrade.sh les ramènent."
    deep:
      sub: "${INSTALL_DIR}/netbox/media · netbox/netbox/configuration.py · ldap_config.py · local_requirements.txt"
  - id: bdir
    kind: store
    label: { en: "Backup directory", fr: "Répertoire de sauvegarde" }
    sub: "${BACKUP_DIR}"
    in: server
    desc:
      en: "Root-only, since the archive holds the database password and the secret key. Dumps and archives older than ${BACKUP_RETENTION_DAYS} days are deleted."
      fr: "Réservé à root, puisque l'archive contient le mot de passe de la base et la clé secrète. Dumps et archives de plus de ${BACKUP_RETENTION_DAYS} jours sont supprimés."
    deep:
      sub: "${BACKUP_DIR} · 0700 · netbox-<stamp>.dump · netbox-files-<stamp>.tar.gz"
  - id: offsite
    kind: cloud
    label: { en: "Backup host", fr: "Serveur de sauvegarde" }
    sub: "${BACKUP_HOST}"
    when: { flag: OFFSITE }
    desc:
      en: "Another machine, another site. rsync mirrors the backup directory there over SSH with a key that can do nothing else."
      fr: "Une autre machine, un autre site. rsync y reflète le répertoire de sauvegarde en SSH avec une clé qui ne sait rien faire d'autre."
    deep:
      sub: "${BACKUP_USER}@${BACKUP_HOST}:${BACKUP_PATH} · rsync -az --delete · ed25519 key"
  - id: github
    kind: cloud
    label: { en: "NetBox releases", fr: "Releases NetBox" }
    sub: "git tags · ${TARGET_REF}"
    desc:
      en: "The upstream repository. An upgrade starts by reading the release notes of ${TARGET_REF}, then fetching its tag."
      fr: "Le dépôt amont. Une mise à jour commence par lire les notes de version de ${TARGET_REF}, puis récupérer son tag."
    deep:
      sub: "github.com/netbox-community/netbox · git fetch --tags · checkout ${TARGET_REF}"
  - id: upgrade
    kind: service
    label: { en: "upgrade.sh", fr: "upgrade.sh" }
    sub: "venv · migrations · static"
    in: server
    desc:
      en: "NetBox's own script: rebuilds the venv, reinstalls plugins from local_requirements.txt, migrates the database, collects static files."
      fr: "Le script de NetBox lui-même : reconstruit le venv, réinstalle les plugins depuis local_requirements.txt, migre la base, collecte les fichiers statiques."
    deep:
      sub: "${INSTALL_DIR}/upgrade.sh · pip install -r local_requirements.txt · migrate · collectstatic"
  - id: services
    kind: server
    label: { en: "netbox · netbox-rq", fr: "netbox · netbox-rq" }
    sub: "systemctl restart"
    in: server
    desc:
      en: "Both units restart after the upgrade: gunicorn serves the new code, the RQ worker runs the new jobs."
      fr: "Les deux unités redémarrent après la mise à jour : gunicorn sert le nouveau code, le worker RQ exécute les nouvelles tâches."
    deep:
      sub: "netbox.service (gunicorn :8001) · netbox-rq.service (rqworker) · restart after upgrade"
  - id: housekeeping
    kind: service
    label: { en: "Housekeeping", fr: "Entretien" }
    sub: "cron.daily · manage.py housekeeping"
    in: timers
    level: guided
    desc:
      en: "Once a day: expired sessions, change records older than CHANGELOG_RETENTION and old job results are deleted. Keeps the dumps small."
      fr: "Une fois par jour : sessions expirées, changements plus vieux que CHANGELOG_RETENTION et anciens résultats de tâches sont supprimés. Garde les dumps petits."
    deep:
      sub: "/etc/cron.daily/netbox-housekeeping → contrib/netbox-housekeeping.sh · CHANGELOG_RETENTION · JOB_RETENTION"
  - id: journal
    kind: file
    label: { en: "Logs", fr: "Logs" }
    sub: "journald · logrotate"
    in: server
    level: deep
    desc:
      en: "gunicorn and the worker log to the journal, capped by SystemMaxUse; the web server rotates its own files through logrotate."
      fr: "gunicorn et le worker écrivent dans le journal, plafonné par SystemMaxUse ; le serveur web fait tourner ses propres fichiers via logrotate."
    deep:
      sub: "journalctl -u netbox / netbox-rq · SystemMaxUse=500M · /etc/logrotate.d/"
  - id: probe
    kind: client
    label: { en: "Probe", fr: "Sonde" }
    sub: "GET /api/status/"
    level: guided
    desc:
      en: "A curl every minute with a read-only token. rq-workers-running at 0 is the alert that matters: the UI keeps working while webhooks and scripts silently stop."
      fr: "Un curl toutes les minutes avec un jeton en lecture. rq-workers-running à 0 est l'alerte qui compte : l'interface continue pendant que webhooks et scripts s'arrêtent en silence."
    deep:
      sub: "https://${NETBOX_HOST}/api/status/ · Authorization: Token · netbox-version · rq-workers-running · optional /metrics"

edges:
  - from: timer
    to: backup
    label: "02:30"
    deep: { label: "starts netbox-backup.service" }
    desc: { en: "The timer starts the oneshot service; the service runs the script and records its output in the journal.", fr: "Le timer démarre le service oneshot ; le service exécute le script et enregistre sa sortie dans le journal." }
  - from: backup
    to: postgres
    label: "pg_dump"
    dashed: true
    level: guided
    deep: { label: "pg_dump -Fc netbox" }
    desc: { en: "A consistent snapshot of the whole database, taken while NetBox keeps running.", fr: "Un instantané cohérent de toute la base, pris pendant que NetBox continue de tourner." }
  - from: backup
    to: files
    label: "tar"
    dashed: true
    level: guided
    deep: { label: "tar -czf · media + config" }
    desc: { en: "The media folder and the hand-written config files, in one compressed archive.", fr: "Le dossier media et les fichiers de config écrits à la main, dans une archive compressée." }
  - from: backup
    to: bdir
    label: "writes"
    deep: { label: "dump + tar.gz · then retention" }
    desc: { en: "Two files per night, named by date. Retention runs after the write, never before.", fr: "Deux fichiers par nuit, nommés par date. La rétention tourne après l'écriture, jamais avant." }
  - from: bdir
    to: offsite
    label: "rsync over SSH"
    when: { flag: OFFSITE }
    deep: { label: "rsync -az --delete · ssh -i netbox-backup" }
    desc: { en: "The folder is mirrored, so the remote side follows the same retention.", fr: "Le dossier est reflété, donc le côté distant suit la même rétention." }
  - from: github
    to: upgrade
    label: "git checkout ${TARGET_REF}"
    deep: { label: "git fetch --tags · git checkout ${TARGET_REF}" }
    desc: { en: "The code moves to the target tag; nothing runs yet. Rollback is a checkout of the previous tag plus a database restore.", fr: "Le code passe au tag cible ; rien ne tourne encore. Le retour arrière, c'est un checkout du tag précédent plus une restauration de base." }
  - from: upgrade
    to: services
    label: "restart"
    deep: { label: "systemctl restart netbox netbox-rq" }
    desc: { en: "Only after upgrade.sh exits cleanly. If it fails, the services still run the old code: fix or roll back first.", fr: "Seulement après une sortie propre d'upgrade.sh. S'il échoue, les services font encore tourner l'ancien code : corrige ou reviens en arrière d'abord." }
  - from: upgrade
    to: postgres
    label: "migrate"
    dashed: true
    level: guided
    desc: { en: "Migrations only go forward: this is why the backup is taken right before, not the night before.", fr: "Les migrations ne vont que vers l'avant : c'est pour ça que la sauvegarde est prise juste avant, pas la nuit d'avant." }
  - from: services
    to: probe
    label: "/api/status/"
    level: guided
    deep: { label: "JSON · netbox-version · rq-workers-running" }
    desc: { en: "The probe reads the version and the worker count after every upgrade, then every minute for good.", fr: "La sonde lit la version et le nombre de workers après chaque mise à jour, puis toutes les minutes pour de bon." }
  - from: housekeeping
    to: postgres
    label: "prunes"
    dashed: true
    level: guided
    deep: { label: "DELETE changelog > CHANGELOG_RETENTION · old jobs · sessions" }
    desc: { en: "Deletes old change records, job results and sessions once a day.", fr: "Supprime une fois par jour les anciens changements, résultats de tâches et sessions." }
  - from: services
    to: journal
    label: "stdout"
    dashed: true
    level: deep
    desc: { en: "gunicorn and the worker write to stdout; systemd stores it in the journal, capped at 500 MB.", fr: "gunicorn et le worker écrivent sur stdout ; systemd le range dans le journal, plafonné à 500 Mo." }
````

---

# Writing a page for Runfold

A self-contained brief. Hand it to a person or a model with a subject ("set up a VPS to host a Next.js site behind nginx with TLS") and you get back the files a page is made of. Nothing else is needed to write; the site's test suite then checks the result.

## 1. What the reader gets, and why it shapes the writing

A page is one tutorial. The reader fills in **their context once** — hostnames, ports, paths, and a few **choices** of stack (Debian or RHEL, nginx or Apache, TLS from Let's Encrypt or self-signed) — and every command on the page is rewritten with their values. Sections that do not apply to their choices disappear.

The reader also picks a **reading level**, once for the whole site:

| Level | EN / FR | What it shows |
|---|---|---|
| Run | Run / Automatique | Only the `<Run>` block: a script that does the whole page. "Too lazy to read? Run it." |
| Quick | Quick / Express | The plain paragraphs and the commands. "It worked? Fine." |
| Guided | Guided / Détaillé | Quick + the `<Guided>` explanations: why, and how, for someone who has never done it. |
| Deep | Deep / Exhaustif | Everything, plus `<Deep>`: mechanism, alternatives, gotchas, what happens under the hood. |

So a page is **written once, in layers**, not four times. Every paragraph you write belongs to a layer. Plain text is Quick; wrap the rest.

Each page also carries a **gist**: a small diagram at the top (boxes, containers, arrows, no coordinates) that follows the reader's choices and level. And each page exists in **English and French**, each with its own file, same structure.

## 2. The files

```
content/<series>/series.yaml                 what all pages of the series share (title, order, shared variables/choices)
content/<series>/<page>/tuto.yaml            this page's contract: metadata, its own variables/choices
content/<series>/<page>/page-en.mdx          the prose, English
content/<series>/<page>/page-fr.mdx          the prose, French (same heading skeleton)
content/<series>/<page>/diagram.yaml         the gist (optional but expected)
```

Slugs are kebab-case (`secure-ssh-on-a-fresh-server`). Variable and choice keys are `UPPER_SNAKE_CASE`.

A page **inherits** the series' groups, variables and choices and may add its own; it must **never redeclare** a key the series declares. Two different pages may declare the same key (the reader's value is shared across the series).

Deliver every file in full, each in its own fenced block with the path as title.

## 3. `series.yaml` (only when creating a series)

```yaml
title:   { en: NetBox, fr: NetBox }
summary:
  en: >-
    One or two sentences: what the series takes the reader from and to.
  fr: >-
    Une ou deux phrases : d'où part la série et où elle mène.
order: [install-from-scratch, configure-for-your-team]   # page slugs, reading order

groups:                        # sections of the "Your values" panel
  - id: host
    label: { en: Server, fr: Serveur }
    desc:  { en: Where it runs and how it is reached., fr: Où ça tourne et comment on l'atteint. }
  - id: auth
    label: { en: Directory, fr: Annuaire }
    when:  { flag: LDAP }      # the whole group disappears when the choice is off

vars:                          # same shape as in tuto.yaml, see below
  - key: NETBOX_HOST
    kind: hostname
    group: host
    default: netbox.example.com
    label:  { en: NetBox hostname, fr: Nom d'hôte de NetBox }
    hint:   { en: …, fr: … }
    impact: { en: …, fr: … }

choices:
  - key: OS
    type: select
    label: { en: Distribution, fr: Distribution }
    default: ubuntu
    options:
      - { value: ubuntu, label: { en: Ubuntu 24.04, fr: Ubuntu 24.04 } }
      - { value: rhel,   label: { en: RHEL 9 / Rocky / Alma, fr: RHEL 9 / Rocky / Alma } }
  - key: LDAP
    type: boolean
    label: { en: Authenticate against LDAP, fr: Authentifier via LDAP }
    default: false
```

## 4. `tuto.yaml`

```yaml
# Inherits from ../series.yaml: <list the inherited keys here as a comment>.
title:   { en: Secure SSH on a fresh server, fr: Sécuriser l'accès SSH d'un serveur neuf }
summary:
  en: >-
    Two or three lines. What the reader has at the end. Concrete.
  fr: >-
    Deux ou trois lignes. Ce que le lecteur a à la fin. Concret.
difficulty: beginner           # beginner | intermediate | advanced
tags: [ssh, linux, security]   # lowercase, used by search filters
authors: [thudal]
created: 2026-09-26            # YYYY-MM-DD
minutes: 10                    # hands-on time
validated: Debian 12 · Rocky 9 # upstream version / platform it was written for

groups:
  - id: access
    label: { en: Access, fr: Accès }

vars:
  - key: SSH_PORT
    kind: port                 # text | ip | cidr | port | hostname | domain | user | email | secret | path | url | sshkey
    group: access
    default: "1234"            # always a string; "" for none (secrets)
    label: { en: SSH port, fr: Port SSH }
    hint:                      # what it is, one or two sentences — shown in the ? tip
      en: The port SSH listens on after hardening. Anything but 22 removes most automated noise.
      fr: Le port sur lequel SSH écoutera après durcissement. Tout sauf 22 supprime l'essentiel du bruit automatisé.
    impact:                    # consequences of getting it wrong, where it is reused — also in the tip
      en: Reused by the firewall and fail2ban. Open it in the firewall before restarting SSH, or you lock yourself out.
      fr: Réutilisé par le pare-feu et fail2ban. Ouvre-le dans le pare-feu avant de redémarrer SSH, sinon tu te verrouilles dehors.
    when: { flag: FAIL2BAN }   # optional: only relevant when a choice holds

choices:
  - key: FAIL2BAN
    type: boolean
    label: { en: Ban brute-force attempts (fail2ban), fr: Bannir le brute force (fail2ban) }
    hint:  { en: Recommended on any server reachable from the internet., fr: Recommandé sur tout serveur joignable depuis internet. }
    default: true
```

Rules:
- `kind: secret` for passwords and tokens: masked in the panel, never printed in the print view, never carried by share links. `default: ""`.
- Every value a reader could reasonably change is a variable. Hard-code only true constants (a well-known port of a protocol, a package name).
- `hint` and `impact` are the **whole explanation** of a variable; the page prose should not repeat them.
- Anything the reader **must** provide (their public key, their token) is a variable with `required: true` and no example default (marked in the panel). A block that uses an empty or invalid value is marked red and cannot be copied. Never write an example that looks like code in a block (`ssh-ed25519 AAAA… user@host`, `YOUR.IP`, `<user>`): readers paste it as is (the test refuses these).
- `kind: sshkey` checks a whole OpenSSH public key line (`ssh-ed25519 AAAAC3…`).
- `${EDITOR}` is reserved: the app fills it with the reader's editor (vim by default, nano in the panel). Write `sudo ${EDITOR} /etc/x` whenever the reader edits a file by hand. Never declare it.
- Conditions (`when`, and `<When>` in prose): `{ flag: KEY }`, `{ notFlag: KEY }`, `{ is: KEY, equals: value }`, `{ is: KEY, oneOf: [a, b] }`. They may only reference choices.

## 5. `page-en.mdx` / `page-fr.mdx`

### Skeleton

```mdx
{/* First pass — to be validated against <official docs URL> before publishing. */}

One paragraph (Quick): what this page does, in what order, for whom.

<Run>

One sentence saying what the script does, on which machine, as which user, and what must be in place first.

```bash
#!/usr/bin/env bash
set -euo pipefail
# <Title> — ${MAIN_VAR}
… the whole page as a script, using ${VARS} …
```

</Run>

## Before you start

<Guided>What you need on hand: accounts, access, prerequisites, and where the commands run.</Guided>

<Check cmd="…" expect="…" />

## First step

Plain sentence saying what to do.

```bash
command with ${VARS}
```

<Guided>Why this step, what the command does, what to look at.</Guided>

<Deep>The mechanism, the alternative, the gotcha, the thing that bites in production.</Deep>

<Check cmd="…" expect="…" />

## Second step
…

## Done

What the reader now has. What the next page of the series does.
```

The **French file has the same headings, in the same order, the same count** (the test checks it). It is written in natural French with "tu" (tutoiement), not a word-for-word translation.

### The syntax

| Syntax | Meaning |
|---|---|
| `${SSH_PORT}` inside a code fence | replaced by the reader's value; the copy button copies the filled command |
| `<V name="SSH_PORT" />` | the same, inline in prose |
| plain markdown | visible from **Quick** |
| `<Guided>…</Guided>` | visible from **Guided** |
| `<Deep>…</Deep>` | visible at **Deep** only |
| `<Note>…</Note>` | an aside, from Guided |
| `<Warn>…</Warn>` | always visible: anything that can lock you out, lose data or cost money |
| `<When is="OS" equals="debian">…</When>` | conditional on a select choice; also `oneOf="a,b"` |
| `<When flag="FAIL2BAN">…</When>` / `<When notFlag="…">` | conditional on a boolean choice |
| `<Run>…</Run>` | the only thing shown at Run: a script or a single file that does the whole page |
| `<Check cmd="…" expect="…" />` | a verification with a checkbox, tracked per page. Also `id="…"` to name it. |
| `<Details summary="If it fails">…</Details>` | collapsible, closed by default |
| `<Tabs group="x"><Tab label="/etc/a.conf">…</Tab><Tab label="/etc/b.conf">…</Tab></Tabs>` | several files of one feature, side by side |
| ` ```ini title="/etc/x.conf" {1,3-5} ` | caption, highlighted lines |
| ` ```bash on="mac" ` · ` ```bash on="server" as="debian" ` · `as="${USERNAME}"` | where the block is typed and as whom: a coloured badge and edge, one colour per terminal |
| ` ```caddy file="/etc/caddy/Caddyfile" on="server" as="${USERNAME}" ` | **the whole content** of a file: shown with its path, copied as a ready `sudo tee … <<'EOF'` command (or as content only, for an editor). `sudo` is added outside /home and /tmp; force with `sudo` / `sudo="false"` |
| ` ```bash interactive ` | the command asks something (password, confirmation): badge, and it must be alone in its block |
| `$${X}` | a literal `${X}` (GitHub Actions `$${{ … }}`, JS template literals) |
| `<Annotated>` + `# (1)` markers at line ends + a numbered list after the fence | annotations: stripped at Quick, badges at Guided, inline at Deep |
| `## Heading` | a step: numbered automatically, listed in the outline. `###` for sub-steps. |

Details that matter:
- A `<Check>` is a real command whose expected output is **stable**: a version line, `active (running)`, an HTTP `200`. Two to six per page. `expect` may contain `${VARS}` but never a secret. If the command mixes `'` and `"`, write `cmd={"…"}`.
- Fences: ` ```bash ` for shell (a `$` prompt is drawn in front of what the reader types, not in front of comments or continuation lines), ` ```yaml `, ` ```python `, ` ```ini `, ` ```nginx `. One fence per command group; keep them short enough to read.
- Inside `<Run>`, the script must work with the variables and the choices: branch it with `<When>` blocks around separate fences when the stack differs. It runs unattended, so `set -euo pipefail`, no prompts, no `sudo` password expectations that the intro does not state.
- Do not put MDX components inside a fence, and do not put `\"` inside a component attribute (use `cmd={"…"}`).
- Blank line before and after every component, and around fences inside components.
- A component whose content spans several paragraphs or holds a list must be written as a block: the opening tag alone on its line, a blank line, the content, a blank line, the closing tag alone on its line. `<Guided>One line.</Guided>` is fine; `<Guided>Intro:\n\n1. item</Guided>` does not compile.
- New pages carry `status: draft` in `tuto.yaml` until their author has run them end to end.

Rules that come from a real run (the test suite checks the first four):
- **Where and as whom.** When a page uses more than one terminal (the laptop, the server as `debian`, the server as yourself), every block says so with `on=` / `as=`. A tired reader follows the badges, not the sentences.
- **One file, one block, complete.** A file is shown once, whole, with `file="…"`. Never "now add this line before `log {`". When a choice changes the file, `<When>` picks between complete versions of it. The Run script writes exactly the file the page shows.
- **Interactive commands alone.** `adduser`, `passwd`, `ssh-keygen` without `-N`, `ssh-copy-id`: alone in their block, marked `interactive`, or the rest of a paste lands in the password prompt.
- **No fake examples** in blocks (see `required` above).
- **A check proves the right thing.** A `<Check>` that tests an SSH login forces the method under test: `ssh -i KEY -o IdentitiesOnly=yes -o PasswordAuthentication=no -o BatchMode=yes …`. A check that passes thanks to a password before the step that forbids passwords is how people lock themselves out.
- **No unexplained jargon** in the Quick text: "the bare domain, `example.com`, nothing in front" rather than "the apex".

### Voice

Direct, concrete, an experienced engineer talking to a colleague. Short sentences. No marketing, no "simply", no "just". Say what a command does before showing it when it is not obvious. Say when a step is dangerous before the reader runs it (`<Warn>`). Name the file being edited. When a fact is from memory and must be confirmed, keep the common well-known form and add a `<Note>` saying where to confirm it.

Length: 200–400 lines per language file. Accuracy over volume.

## 6. `diagram.yaml` — the gist

```yaml
title:   { en: "One server, four pieces", fr: "Un serveur, quatre briques" }
caption:
  en: "One or two sentences read under the drawing."
  fr: "Une ou deux phrases lues sous le schéma."

groups:                                       # dashed containers
  - id: server
    label: { en: "Your server", fr: "Ton serveur" }
    desc:  { en: "…", fr: "…" }               # shown on hover

nodes:
  - id: users                                 # kebab-case
    kind: user                                # user | client | server | service | store | file | net | cloud (the icon)
    label: { en: "Your team", fr: "Ton équipe" }
    sub: "https://${NETBOX_HOST}"             # second line, may use ${VARS}
    desc:  { en: "What this piece does.", fr: "Ce que fait cette brique." }
  - id: nginx
    kind: net
    label: { en: "nginx", fr: "nginx" }
    sub: "443 → 8001"
    in: server                                # container
    when: { is: WEB, equals: nginx }          # follows the reader's choices
    desc:  { en: "…", fr: "…" }
    guided: { sub: ":443 → 127.0.0.1:8001" }  # overrides from Guided up
    deep:   { sub: "TLS ends here · :443 → 127.0.0.1:8001 · HTTP/1.1" }
  - id: firewall
    kind: net
    label: { en: "Firewall", fr: "Pare-feu" }
    in: server
    level: guided                             # appears from Guided up (level: deep for Deep only)
    desc:  { en: "…", fr: "…" }
  - id: netbox
    kind: server
    label: { en: "NetBox", fr: "NetBox" }
    in: server
    focus: true                               # the thing the page is about, drawn in the accent
    desc:  { en: "…", fr: "…" }

edges:
  - { from: users, to: nginx, label: "HTTPS", max: quick, desc: { en: "…", fr: "…" } }          # a simplification replaced from Guided
  - { from: users, to: firewall, label: "HTTPS · TCP 443", level: guided, deep: { label: "L7 HTTPS · L4 TCP 443 · L3 IP" }, desc: { en: "…", fr: "…" } }
  - { from: firewall, to: nginx, level: guided, desc: { en: "…", fr: "…" } }
  - { from: nginx, to: netbox, desc: { en: "…", fr: "…" } }
  - { from: netbox, to: ldap, label: "bind", dashed: true, when: { flag: LDAP }, desc: { en: "…", fr: "…" } }   # dashed = secondary relation
```

Rules:
- Solid arrows lay the boxes out left to right; dashed arrows are secondary and do not move anything. No coordinates ever.
- Quick shows 4–7 boxes; Guided adds the pieces a first-timer should know exist; Deep shows every layer (DNS, TLS, ports, sockets, units, log files).
- `desc` on every node, edge and group, both languages; it is what the reader gets on hover.
- Only reference variables and choices declared for the page (series + page).

## 7. Checklist before delivering

- [ ] Every user-facing string exists in `en` and `fr`.
- [ ] Every `${VAR}` and `<V name>` in the prose is declared (series or page); every `<When>` references a declared choice.
- [ ] No key redeclared that the series already declares.
- [ ] Same `##`/`###` headings, same order, same count, in both language files.
- [ ] A `<Run>` block that does the whole page with the variables.
- [ ] 2–6 `<Check>`s with stable expected output; no secret in `expect`.
- [ ] `<Warn>` before anything that locks out, deletes or costs.
- [ ] The first-pass comment at the top of each `.mdx`, naming the docs to validate against.
- [ ] `diagram.yaml` with `desc` everywhere, `focus` on one node, levels used.
- [ ] Every file delivered once, in full, as a `file="…"` block; the Run script writes the same content.
- [ ] `on=` / `as=` on every block when the page uses more than one terminal; `interactive` commands alone.
- [ ] What the reader must provide is a `required` variable, never an example in a block.

## 8. A worked request

> Write the page `host-a-nextjs-site` for a new series `vps` (title "A VPS from scratch"): from a freshly delivered Debian 12 VPS to a Next.js site served by nginx with a Let's Encrypt certificate, the app run by systemd, deployed by `git pull` + `npm run build`. Series variables: SERVER_IP (ip), USERNAME (user, the deploy user), DOMAIN (domain). Page variables: APP_DIR (path, default /srv/site), REPO_URL (url), NODE_MAJOR (text, default "22"), APP_PORT (port, default "3000"). Choices: TLS (select letsencrypt | selfsigned, default letsencrypt), PM (select systemd | pm2, default systemd). Deliver series.yaml, tuto.yaml, page-en.mdx, page-fr.mdx, diagram.yaml.

The answer is five fenced blocks, complete, following sections 3–6 and passing the checklist in 7.
