Backups, monitoring and the monthly routine

Know what is worth saving on the server and what is not. An encrypted off-site copy made every night by restic, with a restore you have actually tested; an email when the site goes down or the backup did not run; logs that cannot fill the disk; and a ten-minute check once a month.

intermediate~40 min hands-on
#backup#restic#s3#monitoring#systemd#ovh#debian

Not validated end to end yet — be the first.Report a problem

Draft — not yet run end to end. This page was written but its author has not yet run it on a real machine. Commands may be wrong: read before you run, and tell us what breaks.

The gistSaved three ways, watched from outside
Your VPS
readsencrypted, nightlywhole disk, dailyHTTPS every 5 minemail if down
Config, certs, web root/etc · /var/lib/caddy · /home ·
resticnightly · 7d / 4w / 6m
S3 bucket
OVH backupdaily · 1 day kept
Your sitehttps://
Uptime monitorevery 5 min · keyword
You

Every night restic reads the server's configuration and certificates, encrypts them and sends them to a bucket outside the VPS; OVH keeps its own copy of the whole disk. An external monitor checks the site every five minutes and emails you when it stops answering.

The site is live. This page makes sure that when something breaks, you hear about it first and you can put it back. You decide what is worth saving, find OVH's own backups, send an encrypted copy off-site every night with restic, restore from it once to prove it works, put an external monitor on the site, cap the logs, and end with a ten-minute routine to run once a month.

Before you start

Pages 3 to 6 are done: you log in with ssh vps, Caddy serves the site over HTTPS, and deploys go through vps-deploy. Everything below runs on the server after ssh vps, as , except the web dashboards and the blocks marked for the Mac.

Check
$systemctl is-active caddy
Expected output
active

What to back up

The site's source lives only in the project folder on your Mac, : the site is rebuilt and redeployed from there in minutes (pages 1 and 6). What the Mac cannot rebuild is the server's own state.

PathWhat is in itIf you lose it
/etcSSH drop-in, Caddyfile, ufw rules, fail2ban jails, apt sources, sudoers, users and groupsRedo pages 3 to 5 by hand: an hour or two, with room for mistakes
/var/lib/caddyTLS certificates and the ACME accountCaddy gets new ones on its own, within Let's Encrypt's rate limits
/homeYour dotfiles, both users' authorized_keys, scripts you wroteRe-add keys through the KVM console
The built releases, not the source: no code, no originalsRedeploy from the Mac; this copy only matters if the Mac and its backup are both gone
/var/lib/<app> (later)A backend's data, once you add oneIrreplaceable: the real reason backups exist

OVH backups and snapshots

In the OVH control panel, open your VPS (Bare Metal Cloud → Virtual Private Servers → the VPS). The Home tab lists the options: Automated backup and Snapshot.

  • Automated backup: included with the VPS, one per day, one day kept. The ... next to it → Restore.
  • Snapshot: a paid option. You take it by hand, before a risky change: a Debian major upgrade, a Caddy config you are unsure of, a firewall change.

Create the bucket and keys

The off-site copy goes to S3 object storage: a bucket, plus an S3 user whose keys restic uses. At OVH, in the control panel:

  1. Public Cloud → create a project if you have none (it needs a payment method).
  2. Object Storage → Create an object container: S3 API, standard class, region GRA (Gravelines), name .
  3. Object Storage → S3 users → create a user, give it read/write on the container, and copy its access key and secret key into the values panel and your password manager.

Set up restic

restic makes encrypted, deduplicated snapshots of directories into a "repository", here the bucket. Install it, then write its credentials in a file only root can read.

Server·
$sudo apt install -y restic
$sudo install -d -m 700 /etc/restic
$sudo install -m 600 /dev/null /etc/restic/env

Then write the credentials into it: copy the command and paste it in the terminal. The file already exists with mode 600, and tee keeps that mode, so only root can read the secrets.

Server·writes a file/etc/restic/env
RESTIC_REPOSITORY=s3:/
AWS_ACCESS_KEY_ID=
AWS_SECRET_ACCESS_KEY=
RESTIC_PASSWORD=

Then the list of what to skip, /etc/restic/excludes:

Server·writes a file/etc/restic/excludes
/home/*/.cache
/var/lib/caddy/.local/share/caddy/locks

And the repository itself:

Server·
$sudo bash -c 'set -a; . /etc/restic/env; restic init'
If restic init fails
  • The specified bucket does not exist: the bucket name or the endpoint region does not match. Compare with the control panel.
  • Access Denied or SignatureDoesNotMatch: wrong keys, or the S3 user has no rights on the container.
  • A region error: add AWS_DEFAULT_REGION=gra (your region, lowercase) to /etc/restic/env.
  • config file already exists: the repository is already initialized; nothing to do.

To edit /etc/restic/env by hand:

Server·
$sudo vim /etc/restic/env

Now the nightly job: a systemd service that makes the backup and prunes old snapshots, and a timer that starts it.

With healthchecks.io, the last line pings it after a successful backup. Its URL comes from the check you create further down, in "Get told when something breaks": the block can be copied once is filled, so create the check first, or come back to this block then.

Server·writes a file/etc/systemd/system/restic-backup.service
[Unit]
Description=Nightly restic backup to off-site S3
Wants=network-online.target
After=network-online.target
[Service]
Type=oneshot
EnvironmentFile=/etc/restic/env
Environment=RESTIC_CACHE_DIR=/var/cache/restic
CacheDirectory=restic
Nice=10
ExecStart=/usr/bin/restic backup /etc /var/lib/caddy /home --exclude-file=/etc/restic/excludes --tag nightly
ExecStartPost=/usr/bin/restic forget --prune --keep-daily 7 --keep-weekly 4 --keep-monthly 6
ExecStartPost=-/usr/bin/curl -fsS -m 10 --retry 3 -o /dev/null

Copy each block and paste it in the terminal on the server: each one writes its file. Then enable the timer and run one backup now:

Server·
$sudo systemctl daemon-reload
$sudo systemctl enable --now restic-backup.timer
$sudo systemctl start restic-backup.service
Check
$systemctl is-active restic-backup.timer
Expected output
active
Check
$sudo systemctl show -p Result restic-backup.service
Expected output
Result=success

Test a restore

A backup you never restored is a hope, not a backup. Restore the Caddy configuration into a scratch folder and compare it with the live one:

Server·
$sudo bash -c 'set -a; . /etc/restic/env; restic restore latest --target /tmp/restore-test --include /etc/caddy'
$sudo diff -r /etc/caddy /tmp/restore-test/etc/caddy && echo "restore OK"
Check
$sudo diff -r /etc/caddy /tmp/restore-test/etc/caddy && echo 'restore OK'
Expected output
restore OK

Then remove the scratch copy with sudo rm -rf /tmp/restore-test.

Get told when something breaks

An external service checks the site from the internet and emails you when it stops answering. UptimeRobot is one example; Better Stack and others work the same way. Create a free account and add a monitor:

  • Type Keyword, URL https:///, every 5 minutes.
  • Keyword: a word that only your real home page contains, such as your name in the title.
  • Alert contact: .

The uptime monitor watches the site; healthchecks.io watches the backup. It works the other way round: your server pings it after each backup, and it emails you when the ping does not come. On healthchecks.io, create an account and Add Check:

  • Period 1 day, grace time 2 hours.
  • Integrations: email to (on by default for the account address).
  • Copy the ping URL, shown on the check's page, into . If you already wrote restic-backup.service above, copy its block again (it now carries the URL) and paste it on the server, then reload systemd and run one backup:
Server·
$sudo systemctl daemon-reload
$sudo systemctl start restic-backup.service

Keep logs in check

journald keeps up to 10% of the disk by default, up to 4 GB. On a 40 GB VPS, cap it at 500 MB with a drop-in, /etc/systemd/journald.conf.d/size.conf:

Server·
$sudo mkdir -p /etc/systemd/journald.conf.d
Server·writes a file/etc/systemd/journald.conf.d/size.conf
[Journal]
SystemMaxUse=500M

Copy the command and paste it in the terminal, then restart journald and look at its size:

Server·
$sudo systemctl restart systemd-journald
$journalctl --disk-usage
Check
$systemd-analyze cat-config systemd/journald.conf | grep '^SystemMaxUse'
Expected output
SystemMaxUse=500M

The monthly routine

Ten minutes, once a month, same day each month. From your laptop, ssh vps, then:

Server·
Sensitive command — reboots or powers off the machine. Review before running.
$# 1. Packages to upgrade: install them with sudo apt upgrade
$sudo apt update && apt list --upgradable
$# 2. After a kernel update: sudo reboot, then check the site
$[ -f /var/run/reboot-required ] && echo "reboot needed"
$# 3. Disk use should stay under 80%
$df -h /
$# 4. Banned addresses (if fail2ban was chosen on page 3): nothing to do unless one is yours
$sudo fail2ban-client status sshd 2>/dev/null || echo "fail2ban not installed"
$# 5. Expect "0 loaded units listed"
$systemctl --failed
$# 6. The month's errors: read them once, search the new ones
$sudo journalctl -p err --since '30 days ago' | tail -n 30

Then the backups: the timer shows the next run, and the last snapshot should be from last night.

Server·
$systemctl list-timers restic-backup.timer
$sudo bash -c 'set -a; . /etc/restic/env; restic snapshots --latest 3'

Last, open the uptime monitor's dashboard and look at the month's uptime and incidents. Then, on the Mac, check that Time Machine saved the project folder recently:

Mac
$tmutil latestbackup

Done

Your server's state is saved three ways, and a restore is proven. You get an email when the site stops answering, the logs cannot fill the disk, and a monthly routine keeps it all honest.

The next page, Launch checklist, goes through everything once more before you announce the site.

Did everything work?

If you followed this page to the end on a real machine, say so. Your validation is dated and records your stack, so the next reader on the same path knows it still works.

This copy is read-only. To report that it works, or that it does not, open an issue

Only your stack choices are recorded, never your values. The pseudonym stays on this browser.