The site is live. This page makes sure that when something breaks, you hear about it first and you can put it back. You decide what is worth saving, find OVH's own backups, send an encrypted copy off-site every night with restic, restore from it once to prove it works, put an external monitor on the site, cap the logs, and end with a ten-minute routine to run once a month.
Before you start
Pages 3 to 6 are done: you log in with ssh vps, Caddy serves the site over HTTPS, and deploys go through vps-deploy. Everything below runs on the server after ssh vps, as , except the web dashboards and the blocks marked for the Mac.
$systemctl is-active caddyactive
What to back up
The site's source lives only in the project folder on your Mac, : the site is rebuilt and redeployed from there in minutes (pages 1 and 6). What the Mac cannot rebuild is the server's own state.
| Path | What is in it | If you lose it |
|---|---|---|
/etc | SSH drop-in, Caddyfile, ufw rules, fail2ban jails, apt sources, sudoers, users and groups | Redo pages 3 to 5 by hand: an hour or two, with room for mistakes |
/var/lib/caddy | TLS certificates and the ACME account | Caddy gets new ones on its own, within Let's Encrypt's rate limits |
/home | Your dotfiles, both users' authorized_keys, scripts you wrote | Re-add keys through the KVM console |
| The built releases, not the source: no code, no originals | Redeploy from the Mac; this copy only matters if the Mac and its backup are both gone |
/var/lib/<app> (later) | A backend's data, once you add one | Irreplaceable: the real reason backups exist |
OVH backups and snapshots
In the OVH control panel, open your VPS (Bare Metal Cloud → Virtual Private Servers → the VPS). The Home tab lists the options: Automated backup and Snapshot.
- Automated backup: included with the VPS, one per day, one day kept. The
...next to it → Restore. - Snapshot: a paid option. You take it by hand, before a risky change: a Debian major upgrade, a Caddy config you are unsure of, a firewall change.
Create the bucket and keys
The off-site copy goes to S3 object storage: a bucket, plus an S3 user whose keys restic uses. At OVH, in the control panel:
- Public Cloud → create a project if you have none (it needs a payment method).
- Object Storage → Create an object container: S3 API, standard class, region GRA (Gravelines), name
. - Object Storage → S3 users → create a user, give it read/write on the container, and copy its access key and secret key into the values panel and your password manager.
Set up restic
restic makes encrypted, deduplicated snapshots of directories into a "repository", here the bucket. Install it, then write its credentials in a file only root can read.
$sudo apt install -y restic$sudo install -d -m 700 /etc/restic$sudo install -m 600 /dev/null /etc/restic/envThen write the credentials into it: copy the command and paste it in the terminal. The file already exists with mode 600, and tee keeps that mode, so only root can read the secrets.
RESTIC_REPOSITORY=s3:/AWS_ACCESS_KEY_ID=AWS_SECRET_ACCESS_KEY=RESTIC_PASSWORD=Then the list of what to skip, /etc/restic/excludes:
/home/*/.cache/var/lib/caddy/.local/share/caddy/locksAnd the repository itself:
$sudo bash -c 'set -a; . /etc/restic/env; restic init'If restic init fails
The specified bucket does not exist: the bucket name or the endpoint region does not match. Compare with the control panel.Access DeniedorSignatureDoesNotMatch: wrong keys, or the S3 user has no rights on the container.- A region error: add
AWS_DEFAULT_REGION=gra(your region, lowercase) to/etc/restic/env. config file already exists: the repository is already initialized; nothing to do.
To edit /etc/restic/env by hand:
$sudo vim /etc/restic/envNow the nightly job: a systemd service that makes the backup and prunes old snapshots, and a timer that starts it.
With healthchecks.io, the last line pings it after a successful backup. Its URL comes from the check you create further down, in "Get told when something breaks": the block can be copied once is filled, so create the check first, or come back to this block then.
[Unit]Description=Nightly restic backup to off-site S3Wants=network-online.targetAfter=network-online.target[Service]Type=oneshotEnvironmentFile=/etc/restic/envEnvironment=RESTIC_CACHE_DIR=/var/cache/resticCacheDirectory=resticNice=10ExecStart=/usr/bin/restic backup /etc /var/lib/caddy /home --exclude-file=/etc/restic/excludes --tag nightlyExecStartPost=/usr/bin/restic forget --prune --keep-daily 7 --keep-weekly 4 --keep-monthly 6ExecStartPost=-/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null Copy each block and paste it in the terminal on the server: each one writes its file. Then enable the timer and run one backup now:
$sudo systemctl daemon-reload$sudo systemctl enable --now restic-backup.timer$sudo systemctl start restic-backup.service$systemctl is-active restic-backup.timeractive
$sudo systemctl show -p Result restic-backup.serviceResult=success
Test a restore
A backup you never restored is a hope, not a backup. Restore the Caddy configuration into a scratch folder and compare it with the live one:
$sudo bash -c 'set -a; . /etc/restic/env; restic restore latest --target /tmp/restore-test --include /etc/caddy'$sudo diff -r /etc/caddy /tmp/restore-test/etc/caddy && echo "restore OK"$sudo diff -r /etc/caddy /tmp/restore-test/etc/caddy && echo 'restore OK'restore OK
Then remove the scratch copy with sudo rm -rf /tmp/restore-test.
Get told when something breaks
An external service checks the site from the internet and emails you when it stops answering. UptimeRobot is one example; Better Stack and others work the same way. Create a free account and add a monitor:
- Type Keyword, URL https://
/, every 5 minutes. - Keyword: a word that only your real home page contains, such as your name in the title.
- Alert contact:
.
The uptime monitor watches the site; healthchecks.io watches the backup. It works the other way round: your server pings it after each backup, and it emails you when the ping does not come. On healthchecks.io, create an account and Add Check:
- Period 1 day, grace time 2 hours.
- Integrations: email to
(on by default for the account address). - Copy the ping URL, shown on the check's page, into
. If you already wroterestic-backup.serviceabove, copy its block again (it now carries the URL) and paste it on the server, then reload systemd and run one backup:
$sudo systemctl daemon-reload$sudo systemctl start restic-backup.serviceKeep logs in check
journald keeps up to 10% of the disk by default, up to 4 GB. On a 40 GB VPS, cap it at 500 MB with a drop-in, /etc/systemd/journald.conf.d/size.conf:
$sudo mkdir -p /etc/systemd/journald.conf.d[Journal]SystemMaxUse=500MCopy the command and paste it in the terminal, then restart journald and look at its size:
$sudo systemctl restart systemd-journald$journalctl --disk-usage$systemd-analyze cat-config systemd/journald.conf | grep '^SystemMaxUse'SystemMaxUse=500M
The monthly routine
Ten minutes, once a month, same day each month. From your laptop, ssh vps, then:
$# 1. Packages to upgrade: install them with sudo apt upgrade$sudo apt update && apt list --upgradable$# 2. After a kernel update: sudo reboot, then check the site$[ -f /var/run/reboot-required ] && echo "reboot needed"$# 3. Disk use should stay under 80%$df -h /$# 4. Banned addresses (if fail2ban was chosen on page 3): nothing to do unless one is yours$sudo fail2ban-client status sshd 2>/dev/null || echo "fail2ban not installed"$# 5. Expect "0 loaded units listed"$systemctl --failed$# 6. The month's errors: read them once, search the new ones$sudo journalctl -p err --since '30 days ago' | tail -n 30Then the backups: the timer shows the next run, and the last snapshot should be from last night.
$systemctl list-timers restic-backup.timer$sudo bash -c 'set -a; . /etc/restic/env; restic snapshots --latest 3'Last, open the uptime monitor's dashboard and look at the month's uptime and incidents. Then, on the Mac, check that Time Machine saved the project folder recently:
$tmutil latestbackupDone
Your server's state is saved three ways, and a restore is proven. You get an email when the site stops answering, the logs cannot fill the disk, and a monthly routine keeps it all honest.
The next page, Launch checklist, goes through everything once more before you announce the site.