# Source of "Monter un cluster k3s"

The files of this tutorial, as they are in the repository. The authoring brief that defines the format follows them.

````yaml title="content/kubernetes/series.yaml"
title:
  en: Kubernetes
  fr: Kubernetes
summary:
  en: >-
    From a few Linux machines to a small production-grade k3s cluster, then a real web
    application on it — indicat, thudal's SaaS prototype — with a database, TLS and
    continuous deployment.
  fr: >-
    De quelques machines Linux à un petit cluster k3s digne de la prod, puis une vraie
    application web dessus — indicat, le prototype SaaS de thudal — avec base de données,
    TLS et déploiement continu.
order:
  - bootstrap-a-k3s-cluster
  - deploy-indicat
  - deploy-on-push-with-ci

# Shared by every page: the reader fills these once for the whole series.
groups:
  - id: cluster
    label: { en: Cluster, fr: Cluster }
    desc: { en: The k3s nodes and how you reach them., fr: Les nœuds k3s et comment tu les atteins. }
  - id: app
    label: { en: Application, fr: Application }
    desc: { en: Where indicat is published and how it is named inside the cluster., fr: Où indicat est publié et comment il est nommé dans le cluster. }

vars:
  - key: NODE_IP
    kind: ip
    group: cluster
    default: 203.0.113.20
    label: { en: First server node IP, fr: IP du premier nœud serveur }
    hint:
      en: The address of the first k3s server, reachable from your laptop. The API and the kubeconfig point at it.
      fr: L'adresse du premier serveur k3s, joignable depuis ton portable. L'API et le kubeconfig pointent dessus.
    impact:
      en: Baked into the API certificate (--tls-san) and into the kubeconfig you copy to your laptop and to CI. Change it later and both must be redone.
      fr: Inscrite dans le certificat de l'API (--tls-san) et dans le kubeconfig que tu copies sur ton portable et dans la CI. La changer plus tard oblige à refaire les deux.
  - key: K3S_VERSION
    kind: text
    group: cluster
    default: v1.31.4+k3s1
    label: { en: k3s version, fr: Version de k3s }
    hint:
      en: The release the install script pins. Every node of the cluster must run the same one.
      fr: La version que fige le script d'installation. Tous les nœuds du cluster doivent avoir la même.
    impact:
      en: Pinning keeps a re-run of the install script from silently upgrading a node. Upgrades are a deliberate step, one node at a time.
      fr: Figer la version évite qu'une relance du script d'installation mette un nœud à jour en silence. Les mises à jour sont un acte volontaire, nœud par nœud.
  - key: APP_DOMAIN
    kind: domain
    group: app
    default: indicat.example.com
    label: { en: Application domain, fr: Domaine de l'application }
    hint:
      en: The name users type. Its A record must point at the node(s) before the certificate can be issued.
      fr: Le nom que tapent les utilisateurs. Son enregistrement A doit pointer vers le(s) nœud(s) avant que le certificat puisse être émis.
    impact:
      en: Used as the Ingress host and as the certificate's common name. With Let's Encrypt the name must resolve publicly, or the HTTP-01 challenge never completes.
      fr: Utilisé comme hôte de l'Ingress et comme nom du certificat. Avec Let's Encrypt, le nom doit résoudre publiquement, sinon le défi HTTP-01 n'aboutit jamais.
  - key: NAMESPACE
    kind: text
    group: app
    default: indicat
    label: { en: Namespace, fr: Namespace }
    hint:
      en: The Kubernetes namespace everything of the application lives in. Lowercase, no spaces.
      fr: Le namespace Kubernetes où vit tout ce qui touche à l'application. Minuscules, sans espace.
    impact:
      en: Every kubectl command of the series carries -n with this value, and the CI ServiceAccount is only allowed inside it.
      fr: Chaque commande kubectl de la série porte -n avec cette valeur, et le ServiceAccount de la CI n'a de droits que dedans.
  - key: ACME_EMAIL
    kind: email
    group: cluster
    default: ops@example.com
    label: { en: Let's Encrypt email, fr: Email Let's Encrypt }
    when: { is: TLS, equals: letsencrypt }
    hint:
      en: The address Let's Encrypt writes to when a certificate is about to expire without renewal.
      fr: L'adresse à laquelle Let's Encrypt écrit quand un certificat va expirer sans avoir été renouvelé.
    impact:
      en: Stored in the ClusterIssuer. Not published; it only matters if cert-manager stops renewing.
      fr: Stockée dans le ClusterIssuer. Non publiée ; elle ne sert que si cert-manager cesse de renouveler.

choices:
  - key: TOPOLOGY
    type: select
    label: { en: Topology, fr: Topologie }
    default: single
    hint:
      en: One node is enough to learn and for a lab. Three servers survive the loss of one.
      fr: Un nœud suffit pour apprendre et pour un lab. Trois serveurs survivent à la perte de l'un d'eux.
    options:
      - { value: single, label: { en: One node (lab), fr: Un nœud (lab) } }
      - { value: ha, label: { en: Three server nodes (HA, embedded etcd), fr: Trois nœuds serveur (HA, etcd embarqué) } }
  - key: TLS
    type: select
    label: { en: Certificates, fr: Certificats }
    default: letsencrypt
    options:
      - { value: letsencrypt, label: { en: Let's Encrypt, fr: Let's Encrypt } }
      - { value: selfsigned, label: { en: Self-signed, fr: Auto-signés } }
  - key: DB
    type: select
    label: { en: PostgreSQL, fr: PostgreSQL }
    default: cnpg
    options:
      - { value: cnpg, label: { en: PostgreSQL by CloudNativePG operator, fr: PostgreSQL par l'opérateur CloudNativePG } }
      - { value: statefulset, label: { en: PostgreSQL as a plain StatefulSet, fr: PostgreSQL en simple StatefulSet } }
````

````yaml title="content/kubernetes/bootstrap-a-k3s-cluster/tuto.yaml"
# Contract for this page. Inherits NODE_IP, K3S_VERSION, APP_DOMAIN, NAMESPACE, ACME_EMAIL
# and the TOPOLOGY / TLS / DB choices from ../series.yaml.

title:
  en: Bootstrap a k3s cluster
  fr: Monter un cluster k3s
summary:
  en: >-
    From one to three Ubuntu machines to a working Kubernetes: k3s pinned to a version,
    kubectl and helm on your laptop, the bundled Traefik ingress, cert-manager with a
    ClusterIssuer, and a hello-world behind HTTPS to prove the whole chain.
  fr: >-
    D'une à trois machines Ubuntu à un Kubernetes qui marche : k3s figé sur une version,
    kubectl et helm sur ton portable, l'ingress Traefik livré avec, cert-manager et un
    ClusterIssuer, puis un hello-world derrière HTTPS pour prouver toute la chaîne.
difficulty: intermediate
tags: [kubernetes, k3s, traefik, cert-manager, helm]
authors: [thudal]
created: 2026-09-25
minutes: 40
validated: k3s v1.31 · Ubuntu 24.04
status: draft             # not yet run end to end by its author

groups:
  - id: ha
    label: { en: Extra server nodes, fr: Nœuds serveur supplémentaires }
    desc: { en: The two other members of the etcd quorum., fr: Les deux autres membres du quorum etcd. }
    when: { is: TOPOLOGY, equals: ha }

vars:
  - key: NODE2_IP
    kind: ip
    group: ha
    default: 203.0.113.21
    label: { en: Second server node IP, fr: IP du deuxième nœud serveur }
    when: { is: TOPOLOGY, equals: ha }
    hint:
      en: Reachable from the first node on 6443, 2379-2380 and 8472/udp.
      fr: Joignable depuis le premier nœud sur 6443, 2379-2380 et 8472/udp.
    impact:
      en: Joins the embedded etcd. With three members the cluster keeps working when any one of them is down.
      fr: Rejoint l'etcd embarqué. Avec trois membres, le cluster continue de fonctionner quand l'un d'eux est arrêté.
  - key: NODE3_IP
    kind: ip
    group: ha
    default: 203.0.113.22
    label: { en: Third server node IP, fr: IP du troisième nœud serveur }
    when: { is: TOPOLOGY, equals: ha }
    hint:
      en: Same requirements as the second one.
      fr: Mêmes prérequis que le deuxième.
    impact:
      en: Third etcd member. Two servers only would be worse than one, so it is three or one, never two.
      fr: "Troisième membre etcd. Deux serveurs seulement seraient pires qu'un seul : c'est trois ou un, jamais deux."
````

````mdx title="content/kubernetes/bootstrap-a-k3s-cluster/page-en.mdx"
{/* First pass — to be validated against docs.k3s.io (installation, requirements, HA embedded etcd) and cert-manager.io/docs before publishing. */}

k3s is Kubernetes in one binary: the API server, the scheduler, the kubelet, a datastore and an ingress controller, installed by one script in under a minute. This page takes <When is="TOPOLOGY" equals="single">one Ubuntu machine</When><When is="TOPOLOGY" equals="ha">three Ubuntu machines</When> to a cluster you drive from your laptop, with certificates issued automatically, and proves it with a hello-world behind HTTPS at <V name="APP_DOMAIN" />.

<Run>

The script does every step of this page from **your laptop**. It needs root SSH access to <When is="TOPOLOGY" equals="single">the node</When><When is="TOPOLOGY" equals="ha">the three nodes</When>, `kubectl` and `helm` already installed locally, and the DNS record for <V name="APP_DOMAIN" /> already pointing at <V name="NODE_IP" />.

<When is="TOPOLOGY" equals="single">

```bash
#!/usr/bin/env bash
set -euo pipefail
# k3s ${K3S_VERSION} — one server node at ${NODE_IP}
ssh root@${NODE_IP} "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
  curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --tls-san ${NODE_IP} --write-kubeconfig-mode 644"
```

</When>

<When is="TOPOLOGY" equals="ha">

```bash
#!/usr/bin/env bash
set -euo pipefail
# k3s ${K3S_VERSION} — three server nodes, embedded etcd, first one at ${NODE_IP}
ssh root@${NODE_IP} "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
  curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --cluster-init --tls-san ${NODE_IP} --write-kubeconfig-mode 644"
TOKEN=$(ssh root@${NODE_IP} cat /var/lib/rancher/k3s/server/node-token)
for ip in ${NODE2_IP} ${NODE3_IP}; do
  ssh root@$ip "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
    curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} K3S_TOKEN=$TOKEN sh -s - server \
    --server https://${NODE_IP}:6443 --tls-san $ip"
done
```

</When>

```bash
mkdir -p ~/.kube
ssh root@${NODE_IP} cat /etc/rancher/k3s/k3s.yaml | sed "s/127.0.0.1/${NODE_IP}/" > ~/.kube/config
chmod 600 ~/.kube/config
kubectl wait --for=condition=Ready node --all --timeout=120s
helm repo add jetstack https://charts.jetstack.io --force-update
helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager --create-namespace --set crds.enabled=true --wait
```

<When is="TLS" equals="letsencrypt">

```bash
kubectl apply -f - <<EOF
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: ${ACME_EMAIL}
    privateKeySecretRef:
      name: letsencrypt-account-key
    solvers:
      - http01:
          ingress:
            ingressClassName: traefik
EOF
ISSUER=letsencrypt
```

</When>

<When is="TLS" equals="selfsigned">

```bash
kubectl apply -f - <<EOF
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: selfsigned
spec:
  selfSigned: {}
EOF
ISSUER=selfsigned
```

</When>

```bash
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata: { name: hello }
spec:
  replicas: 1
  selector: { matchLabels: { app: hello } }
  template:
    metadata: { labels: { app: hello } }
    spec:
      containers:
        - name: whoami
          image: traefik/whoami:v1.10
          ports: [{ containerPort: 80 }]
---
apiVersion: v1
kind: Service
metadata: { name: hello }
spec:
  selector: { app: hello }
  ports: [{ port: 80, targetPort: 80 }]
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations: { cert-manager.io/cluster-issuer: $ISSUER }
spec:
  ingressClassName: traefik
  tls: [{ hosts: ["${APP_DOMAIN}"], secretName: hello-tls }]
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend: { service: { name: hello, port: { number: 80 } } }
EOF
kubectl wait --for=condition=Ready certificate/hello-tls --timeout=180s
echo "Done. Open https://${APP_DOMAIN}"
```

<Warn>The script disables swap and rewrites `/etc/fstab` on every node. Run it on fresh machines dedicated to the cluster, not on a server that does something else.</Warn>

</Run>

## Before you start

<Guided>You need <When is="TOPOLOGY" equals="single">one Ubuntu 24.04 machine</When><When is="TOPOLOGY" equals="ha">three Ubuntu 24.04 machines</When> with 2 CPU and 4 GB of RAM or more, root access over SSH, and a DNS record for <V name="APP_DOMAIN" /> pointing at <V name="NODE_IP" />. Kubernetes wants three things from the OS before anything else: no swap, a synchronised clock, and a handful of open ports.</Guided>

On each node, as root:

```bash
swapoff -a
sed -i '/ swap / s/^/#/' /etc/fstab
timedatectl set-ntp true
```

<Deep>The kubelet refuses to start with swap enabled by default (`failSwapOn`), because it cannot reason about memory limits when the kernel may page a container out. The clock matters for TLS: certificates carry a validity window and etcd members compare timestamps; a node a few minutes off gets rejected with confusing errors.</Deep>

<Guided>If `ufw` is enabled on the nodes, open what k3s needs before installing. The pod and service networks (`10.42.0.0/16` and `10.43.0.0/16`) must be allowed as well, or pods cannot reach each other across nodes.</Guided>

```bash
ufw allow 22/tcp
ufw allow 80/tcp
ufw allow 443/tcp
ufw allow 6443/tcp
ufw allow from 10.42.0.0/16 to any
ufw allow from 10.43.0.0/16 to any
```

<When is="TOPOLOGY" equals="ha">

```bash
ufw allow 10250/tcp
ufw allow 8472/udp
ufw allow 2379:2380/tcp
```

</When>

<Deep>Port by port: 6443 is the API (your laptop, the other nodes); 10250 is the kubelet, used by `kubectl logs` and `exec` and by metrics-server; 8472/udp is flannel's VXLAN tunnel between nodes; 2379-2380 is etcd, only in the HA topology. 80 and 443 are Traefik. The k3s docs list a few more (5001 for the embedded registry, 51820/udp for WireGuard) that this page does not use.</Deep>

## Install the first server

<Guided>The official script downloads the binary for the version you pin, writes a systemd unit and starts it. `--tls-san` adds <V name="NODE_IP" /> to the API certificate, so a kubeconfig pointing at that address is trusted. `--write-kubeconfig-mode 644` lets a non-root user on the node read the kubeconfig.</Guided>

<When is="TOPOLOGY" equals="single">

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --tls-san ${NODE_IP} --write-kubeconfig-mode 644
```

</When>

<When is="TOPOLOGY" equals="ha">

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --cluster-init --tls-san ${NODE_IP} --write-kubeconfig-mode 644
```

<Deep>`--cluster-init` starts an embedded etcd instead of the default SQLite datastore. Only the first node gets it; the others join with `--server`. Without it, k3s stores its state in SQLite, which is fine for one node and impossible to replicate.</Deep>

</When>

<Deep>Everything lands in `/var/lib/rancher/k3s/`: the datastore under `server/db`, the node token under `server/node-token`, and the `server/manifests/` directory, where any YAML you drop is applied automatically (that is how Traefik and CoreDNS get installed). The unit is `k3s.service`; `journalctl -u k3s -f` is where you look when something is off.</Deep>

<Details summary="If the node stays NotReady">
Give it a minute: the node is Ready once flannel and CoreDNS run. If it stays NotReady, `journalctl -u k3s --no-pager | tail -50` usually names the cause: swap still on, a firewall dropping 8472/udp, or a leftover Docker or containerd install on the machine. Ubuntu cloud images with `apparmor` are fine; Raspberry Pi images need cgroups enabled in `cmdline.txt`, see the k3s docs.
</Details>

<When is="TOPOLOGY" equals="ha">

## Join two more servers

<Guided>The first node holds a token that authorises new members. Copy it, then run the same script on the two other nodes with `--server` pointing at the first one. They join as full servers: API, scheduler and an etcd member each.</Guided>

On the first node:

```bash
cat /var/lib/rancher/k3s/server/node-token
```

On <V name="NODE2_IP" /> and <V name="NODE3_IP" />, with the token pasted in `K3S_TOKEN`:

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} K3S_TOKEN=K10…::server:… sh -s - server \
  --server https://${NODE_IP}:6443 --tls-san $(hostname -I | awk '{print $1}')
```

<Warn>Three servers, not two. etcd needs a majority: with two members, losing either one freezes the cluster. With three, any one can go down.</Warn>

<Deep>Every server also runs the kubelet, so the three nodes take workloads. For a cluster that grows, add *agents* (`sh -s - agent --server … --token …`): they run pods and nothing else. The `--server` address is only used to join; afterwards each node talks to the local API. A fixed address in front of the three servers (a load balancer or a floating IP) is what a real HA setup adds, so that a kubeconfig does not depend on <V name="NODE_IP" /> being up.</Deep>

<Check cmd="k3s kubectl get nodes --no-headers | wc -l" expect="3" />

</When>

## Kubeconfig, kubectl and helm on the laptop

<Guided>k3s writes an admin kubeconfig on the node. It points at `127.0.0.1`; copy it to your laptop and replace the address with <V name="NODE_IP" />. From then on, everything happens from your machine.</Guided>

```bash
mkdir -p ~/.kube
ssh root@${NODE_IP} cat /etc/rancher/k3s/k3s.yaml | sed "s/127.0.0.1/${NODE_IP}/" > ~/.kube/config
chmod 600 ~/.kube/config
```

<Guided>If `kubectl` and `helm` are not installed yet:</Guided>

```bash
curl -LO "https://dl.k8s.io/release/$(curl -Ls https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl"
sudo install -m 755 kubectl /usr/local/bin/kubectl
curl -fsSL https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
```

<Note>On macOS, `brew install kubectl helm` does the same. Keep kubectl within one minor version of the cluster (<V name="K3S_VERSION" /> is a 1.31 API): further apart, some commands misbehave silently.</Note>

<Deep>The kubeconfig holds a client certificate signed by the cluster's CA: it is `cluster-admin`, with no expiry worth mentioning. Treat the file like a root password. The next pages create a narrower ServiceAccount for CI precisely so this file never leaves your laptop. If you manage several clusters, rename the context: `kubectl config rename-context default k3s-lab`.</Deep>

<Check cmd="kubectl get nodes --no-headers | awk '{print $2}' | sort -u" expect="Ready" />

<Details summary="If kubectl says x509: certificate is valid for …">
The API certificate does not include <V name="NODE_IP" />: the `--tls-san` flag was missing or given another address. Re-run the install command on the first node with the right `--tls-san`; the script is idempotent and only rewrites the unit and the certificate.
</Details>

## Ingress: Traefik and ServiceLB

<Guided>k3s ships an ingress controller, Traefik, and a small load-balancer, ServiceLB, that publishes any `LoadBalancer` Service on the nodes' own addresses. Together they mean: an `Ingress` object with a host name, and that host name answers on port 443 of every node. Nothing to install.</Guided>

```bash
kubectl -n kube-system get svc traefik
```

<Deep>ServiceLB (formerly klipper-lb) runs a `svclb-traefik` DaemonSet whose pods bind ports 80 and 443 on each node with `hostPort`, and the Service's `EXTERNAL-IP` column lists the node addresses. It is not a real load balancer: DNS pointing at one node is a single point of entry. In the HA topology, put the three addresses in DNS or a load balancer in front. Traefik itself is deployed by the HelmChart controller from `/var/lib/rancher/k3s/server/manifests/traefik.yaml`; to change its settings, do not edit that file (it is overwritten on upgrade) but drop a `HelmChartConfig` next to it. To use another controller, install k3s with `--disable traefik`.</Deep>

<Check cmd="kubectl -n kube-system rollout status deploy/traefik" expect={'deployment "traefik" successfully rolled out'} />

## cert-manager and a ClusterIssuer

<Guided>cert-manager watches `Ingress` objects: when one carries its annotation, it obtains a certificate from an issuer and stores it in the Secret the Ingress names. A `ClusterIssuer` is an issuer available in every namespace: you define it once.</Guided>

```bash
helm repo add jetstack https://charts.jetstack.io --force-update
helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager --create-namespace --set crds.enabled=true --wait
```

<Note>`crds.enabled=true` is the chart option from cert-manager 1.15 on; older charts used `installCRDs=true`. Check the version the repo serves with `helm search repo jetstack`.</Note>

<Check cmd="kubectl -n cert-manager rollout status deploy/cert-manager-webhook" expect={'deployment "cert-manager-webhook" successfully rolled out'} />

<When is="TLS" equals="letsencrypt">

<Annotated>

```yaml title="clusterissuer.yaml" {8,13}
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory   # (1)
    email: ${ACME_EMAIL}                                        # (2)
    privateKeySecretRef:
      name: letsencrypt-account-key                             # (3)
    solvers:
      - http01:
          ingress:
            ingressClassName: traefik                           # (4)
```

1. The production endpoint. It rate-limits (five identical certificates a week): while experimenting, point a second issuer at `https://acme-staging-v02.api.letsencrypt.org/directory`, whose certificates browsers reject but which never runs out.
2. Only used for expiry warnings.
3. The ACME account key, generated by cert-manager on first use and kept in this Secret in the `cert-manager` namespace.
4. HTTP-01: cert-manager creates a temporary Ingress on Traefik serving `/.well-known/acme-challenge/…`, Let's Encrypt fetches it over port 80, then signs. The name must resolve publicly to the node(s).

</Annotated>

```bash
kubectl apply -f clusterissuer.yaml
```

</When>

<When is="TLS" equals="selfsigned">

```yaml title="clusterissuer.yaml"
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: selfsigned
spec:
  selfSigned: {}
```

```bash
kubectl apply -f clusterissuer.yaml
```

<Note>Browsers will warn on every certificate this issuer signs. Fine for a lab without a public name. For something users open, switch the Certificates choice to Let's Encrypt.</Note>

<Deep>A nicer lab setup is a private CA: this self-signed issuer signs one `Certificate` marked `isCA: true`, a second `ClusterIssuer` of type `ca` uses that Secret, and every application certificate chains to it. Import the CA once in your browsers and the warnings go away. The cert-manager docs have the three manifests under "CA".</Deep>

</When>

## Storage: the local-path class

<Guided>k3s comes with one StorageClass, `local-path`, marked as default: a `PersistentVolumeClaim` gets a directory on the node where the pod is scheduled. Nothing to do here; know its limits before the database of the next page relies on it.</Guided>

```bash
kubectl get storageclass
```

<Deep>Volumes live under `/var/lib/rancher/k3s/storage/` on the node. Two consequences: a pod using one is pinned to that node (the scheduler will not move it), and losing the node's disk loses the data. No snapshots, no replication, no resizing. It is perfect for a lab and acceptable for a single-node prod with backups. For anything that must survive a node, the usual next step on k3s is Longhorn, which replicates volumes across nodes; that is beyond this series.</Deep>

## Hello world behind HTTPS

<Guided>One Deployment, one Service, one Ingress with the cert-manager annotation: if <V name="APP_DOMAIN" /> answers with a valid certificate, the whole chain (DNS, Traefik, cert-manager, the issuer) works, and the next page can focus on the application.</Guided>

```yaml title="hello.yaml"
apiVersion: apps/v1
kind: Deployment
metadata:
  name: hello
spec:
  replicas: 1
  selector:
    matchLabels: { app: hello }
  template:
    metadata:
      labels: { app: hello }
    spec:
      containers:
        - name: whoami
          image: traefik/whoami:v1.10
          ports:
            - containerPort: 80
---
apiVersion: v1
kind: Service
metadata:
  name: hello
spec:
  selector: { app: hello }
  ports:
    - port: 80
      targetPort: 80
```

<When is="TLS" equals="letsencrypt">

```yaml title="hello-ingress.yaml" {6,10-11}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt
spec:
  ingressClassName: traefik
  tls:
    - hosts: ["${APP_DOMAIN}"]
      secretName: hello-tls
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: hello
                port: { number: 80 }
```

</When>

<When is="TLS" equals="selfsigned">

```yaml title="hello-ingress.yaml" {6,10-11}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations:
    cert-manager.io/cluster-issuer: selfsigned
spec:
  ingressClassName: traefik
  tls:
    - hosts: ["${APP_DOMAIN}"]
      secretName: hello-tls
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: hello
                port: { number: 80 }
```

</When>

```bash
kubectl apply -f hello.yaml -f hello-ingress.yaml
kubectl get certificate hello-tls -w
```

<Deep>The annotation makes cert-manager create a `Certificate` named after the Secret, then an `Order`, a `Challenge`, and finally the Secret `hello-tls` with `tls.crt` and `tls.key`. Traefik watches Secrets referenced by Ingresses and serves the new certificate without a restart. `kubectl describe challenge` is the place to look when the certificate stays not Ready: it prints the exact URL Let's Encrypt failed to fetch.</Deep>

<Check cmd="kubectl wait --for=condition=Ready certificate/hello-tls --timeout=120s" expect="certificate.cert-manager.io/hello-tls condition met" />

<When is="TLS" equals="letsencrypt">

<Check cmd="curl -sI https://${APP_DOMAIN} | head -1" expect="HTTP/2 200" />

</When>

<When is="TLS" equals="selfsigned">

<Check cmd="curl -skI https://${APP_DOMAIN} | head -1" expect="HTTP/2 200" />

</When>

<Details summary="If the certificate stays not Ready">
In order of likelihood:

- <V name="APP_DOMAIN" /> does not resolve to a node yet, or resolves to a private address Let's Encrypt cannot reach. `dig +short` the name from outside.
- Port 80 is closed somewhere (ufw, the provider's firewall): HTTP-01 needs it, even though users only use 443.
- Rate limit hit after too many attempts: `kubectl describe order` says so. Use the staging endpoint until the setup is right.
- The `Challenge` is `pending` with a `wrong status code 404`: Traefik is not serving the temporary Ingress, check `kubectl -n kube-system logs deploy/traefik`.
</Details>

## Upgrade, uninstall

Upgrading k3s is re-running the install script with a newer version; uninstalling is one script per node.

<Deep>To upgrade, change <V name="K3S_VERSION" /> and re-run the exact install command on each server, one at a time, the first node first; the script replaces the binary and restarts the unit, pods keep running. Stay within one minor version per hop (1.31 → 1.32, never 1.31 → 1.33), and read the k3s release notes for the datastore. For a cluster with many nodes, the system-upgrade-controller does this from a `Plan` object. `/usr/local/bin/k3s-uninstall.sh` removes everything from a server (`k3s-agent-uninstall.sh` on an agent), including `/var/lib/rancher/k3s` and its volumes: back up first. To move to another datastore or reset etcd, `k3s server --cluster-reset` on one node rebuilds a single-member cluster from its own copy.</Deep>

## Done

<When is="TOPOLOGY" equals="single">One node</When><When is="TOPOLOGY" equals="ha">Three servers with a replicated datastore</When>, driven from your laptop, an ingress on every node, certificates that issue and renew themselves, and a StorageClass for the database. The hello-world can go:

```bash
kubectl delete -f hello-ingress.yaml -f hello.yaml
```

The next page deploys indicat on this cluster: namespace, secrets, PostgreSQL, the application with its Ingress at <V name="APP_DOMAIN" />, and an autoscaler.
````

````mdx title="content/kubernetes/bootstrap-a-k3s-cluster/page-fr.mdx"
{/* Première passe — à valider contre docs.k3s.io (installation, prérequis, HA etcd embarqué) et cert-manager.io/docs avant publication. */}

k3s, c'est Kubernetes en un seul binaire : l'API server, le scheduler, le kubelet, un datastore et un ingress controller, installés par un script en moins d'une minute. Cette page part <When is="TOPOLOGY" equals="single">d'une machine Ubuntu</When><When is="TOPOLOGY" equals="ha">de trois machines Ubuntu</When> et arrive à un cluster que tu pilotes depuis ton portable, avec des certificats émis tout seuls, prouvé par un hello-world derrière HTTPS sur <V name="APP_DOMAIN" />.

<Run>

Le script fait toutes les étapes de cette page depuis **ton portable**. Il lui faut un accès SSH root <When is="TOPOLOGY" equals="single">au nœud</When><When is="TOPOLOGY" equals="ha">aux trois nœuds</When>, `kubectl` et `helm` déjà installés en local, et l'enregistrement DNS de <V name="APP_DOMAIN" /> qui pointe déjà vers <V name="NODE_IP" />.

<When is="TOPOLOGY" equals="single">

```bash
#!/usr/bin/env bash
set -euo pipefail
# k3s ${K3S_VERSION} — un nœud serveur sur ${NODE_IP}
ssh root@${NODE_IP} "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
  curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --tls-san ${NODE_IP} --write-kubeconfig-mode 644"
```

</When>

<When is="TOPOLOGY" equals="ha">

```bash
#!/usr/bin/env bash
set -euo pipefail
# k3s ${K3S_VERSION} — trois nœuds serveur, etcd embarqué, le premier sur ${NODE_IP}
ssh root@${NODE_IP} "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
  curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --cluster-init --tls-san ${NODE_IP} --write-kubeconfig-mode 644"
TOKEN=$(ssh root@${NODE_IP} cat /var/lib/rancher/k3s/server/node-token)
for ip in ${NODE2_IP} ${NODE3_IP}; do
  ssh root@$ip "swapoff -a; sed -i '/ swap / s/^/#/' /etc/fstab; \
    curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} K3S_TOKEN=$TOKEN sh -s - server \
    --server https://${NODE_IP}:6443 --tls-san $ip"
done
```

</When>

```bash
mkdir -p ~/.kube
ssh root@${NODE_IP} cat /etc/rancher/k3s/k3s.yaml | sed "s/127.0.0.1/${NODE_IP}/" > ~/.kube/config
chmod 600 ~/.kube/config
kubectl wait --for=condition=Ready node --all --timeout=120s
helm repo add jetstack https://charts.jetstack.io --force-update
helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager --create-namespace --set crds.enabled=true --wait
```

<When is="TLS" equals="letsencrypt">

```bash
kubectl apply -f - <<EOF
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: ${ACME_EMAIL}
    privateKeySecretRef:
      name: letsencrypt-account-key
    solvers:
      - http01:
          ingress:
            ingressClassName: traefik
EOF
ISSUER=letsencrypt
```

</When>

<When is="TLS" equals="selfsigned">

```bash
kubectl apply -f - <<EOF
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: selfsigned
spec:
  selfSigned: {}
EOF
ISSUER=selfsigned
```

</When>

```bash
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata: { name: hello }
spec:
  replicas: 1
  selector: { matchLabels: { app: hello } }
  template:
    metadata: { labels: { app: hello } }
    spec:
      containers:
        - name: whoami
          image: traefik/whoami:v1.10
          ports: [{ containerPort: 80 }]
---
apiVersion: v1
kind: Service
metadata: { name: hello }
spec:
  selector: { app: hello }
  ports: [{ port: 80, targetPort: 80 }]
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations: { cert-manager.io/cluster-issuer: $ISSUER }
spec:
  ingressClassName: traefik
  tls: [{ hosts: ["${APP_DOMAIN}"], secretName: hello-tls }]
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend: { service: { name: hello, port: { number: 80 } } }
EOF
kubectl wait --for=condition=Ready certificate/hello-tls --timeout=180s
echo "Terminé. Ouvre https://${APP_DOMAIN}"
```

<Warn>Le script désactive le swap et réécrit `/etc/fstab` sur chaque nœud. Lance-le sur des machines neuves dédiées au cluster, pas sur un serveur qui fait autre chose.</Warn>

</Run>

## Avant de commencer

<Guided>Il te faut <When is="TOPOLOGY" equals="single">une machine Ubuntu 24.04</When><When is="TOPOLOGY" equals="ha">trois machines Ubuntu 24.04</When> avec au moins 2 CPU et 4 Go de RAM, un accès root en SSH, et un enregistrement DNS pour <V name="APP_DOMAIN" /> qui pointe vers <V name="NODE_IP" />. Kubernetes attend trois choses de l'OS avant tout : pas de swap, une horloge synchronisée, et une poignée de ports ouverts.</Guided>

Sur chaque nœud, en root :

```bash
swapoff -a
sed -i '/ swap / s/^/#/' /etc/fstab
timedatectl set-ntp true
```

<Deep>Le kubelet refuse de démarrer avec du swap par défaut (`failSwapOn`) : il ne peut pas raisonner sur les limites mémoire si le noyau peut paginer un conteneur. L'horloge compte pour le TLS : les certificats ont une fenêtre de validité et les membres etcd comparent des horodatages ; un nœud décalé de quelques minutes se fait rejeter avec des erreurs déroutantes.</Deep>

<Guided>Si `ufw` est actif sur les nœuds, ouvre ce dont k3s a besoin avant d'installer. Les réseaux des pods et des services (`10.42.0.0/16` et `10.43.0.0/16`) doivent aussi être autorisés, sinon les pods ne se joignent pas d'un nœud à l'autre.</Guided>

```bash
ufw allow 22/tcp
ufw allow 80/tcp
ufw allow 443/tcp
ufw allow 6443/tcp
ufw allow from 10.42.0.0/16 to any
ufw allow from 10.43.0.0/16 to any
```

<When is="TOPOLOGY" equals="ha">

```bash
ufw allow 10250/tcp
ufw allow 8472/udp
ufw allow 2379:2380/tcp
```

</When>

<Deep>Port par port : 6443 c'est l'API (ton portable, les autres nœuds) ; 10250 le kubelet, utilisé par `kubectl logs`, `exec` et par metrics-server ; 8472/udp le tunnel VXLAN de flannel entre nœuds ; 2379-2380 etcd, seulement en topologie HA. 80 et 443, c'est Traefik. La doc k3s en liste quelques autres (5001 pour le registre embarqué, 51820/udp pour WireGuard) que cette page n'utilise pas.</Deep>

## Installer le premier serveur

<Guided>Le script officiel télécharge le binaire de la version que tu figes, écrit une unité systemd et la démarre. `--tls-san` ajoute <V name="NODE_IP" /> au certificat de l'API, pour qu'un kubeconfig qui pointe vers cette adresse soit accepté. `--write-kubeconfig-mode 644` laisse un utilisateur non-root du nœud lire le kubeconfig.</Guided>

<When is="TOPOLOGY" equals="single">

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --tls-san ${NODE_IP} --write-kubeconfig-mode 644
```

</When>

<When is="TOPOLOGY" equals="ha">

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} sh -s - server \
  --cluster-init --tls-san ${NODE_IP} --write-kubeconfig-mode 644
```

<Deep>`--cluster-init` démarre un etcd embarqué à la place du datastore SQLite par défaut. Seul le premier nœud le reçoit ; les autres rejoignent avec `--server`. Sans lui, k3s range son état dans SQLite, très bien pour un nœud et impossible à répliquer.</Deep>

</When>

<Deep>Tout atterrit dans `/var/lib/rancher/k3s/` : le datastore sous `server/db`, le token sous `server/node-token`, et le répertoire `server/manifests/`, où tout YAML déposé est appliqué automatiquement (c'est comme ça que Traefik et CoreDNS s'installent). L'unité s'appelle `k3s.service` ; `journalctl -u k3s -f` est l'endroit où regarder quand quelque chose cloche.</Deep>

<Details summary="Si le nœud reste NotReady">
Laisse-lui une minute : le nœud passe Ready une fois flannel et CoreDNS lancés. S'il reste NotReady, `journalctl -u k3s --no-pager | tail -50` nomme en général la cause : swap encore actif, un pare-feu qui bloque 8472/udp, ou un reste de Docker ou de containerd sur la machine. Les images cloud Ubuntu avec `apparmor` passent sans problème ; les images Raspberry Pi demandent d'activer les cgroups dans `cmdline.txt`, voir la doc k3s.
</Details>

<When is="TOPOLOGY" equals="ha">

## Joindre deux serveurs de plus

<Guided>Le premier nœud détient un token qui autorise les nouveaux membres. Copie-le, puis lance le même script sur les deux autres nœuds avec `--server` pointant vers le premier. Ils rejoignent en serveurs complets : API, scheduler et un membre etcd chacun.</Guided>

Sur le premier nœud :

```bash
cat /var/lib/rancher/k3s/server/node-token
```

Sur <V name="NODE2_IP" /> et <V name="NODE3_IP" />, avec le token collé dans `K3S_TOKEN` :

```bash
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=${K3S_VERSION} K3S_TOKEN=K10…::server:… sh -s - server \
  --server https://${NODE_IP}:6443 --tls-san $(hostname -I | awk '{print $1}')
```

<Warn>Trois serveurs, pas deux. etcd a besoin d'une majorité : à deux membres, perdre l'un ou l'autre gèle le cluster. À trois, n'importe lequel peut tomber.</Warn>

<Deep>Chaque serveur fait aussi tourner le kubelet, donc les trois nœuds prennent des charges. Pour un cluster qui grossit, ajoute des *agents* (`sh -s - agent --server … --token …`) : ils exécutent des pods et rien d'autre. L'adresse `--server` ne sert qu'à rejoindre ; ensuite chaque nœud parle à l'API locale. Une adresse fixe devant les trois serveurs (un load balancer ou une IP flottante), c'est ce qu'ajoute une vraie HA, pour qu'un kubeconfig ne dépende pas de la disponibilité de <V name="NODE_IP" />.</Deep>

<Check cmd="k3s kubectl get nodes --no-headers | wc -l" expect="3" />

</When>

## Kubeconfig, kubectl et helm sur le portable

<Guided>k3s écrit un kubeconfig admin sur le nœud. Il pointe vers `127.0.0.1` ; copie-le sur ton portable et remplace l'adresse par <V name="NODE_IP" />. À partir de là, tout se passe depuis ta machine.</Guided>

```bash
mkdir -p ~/.kube
ssh root@${NODE_IP} cat /etc/rancher/k3s/k3s.yaml | sed "s/127.0.0.1/${NODE_IP}/" > ~/.kube/config
chmod 600 ~/.kube/config
```

<Guided>Si `kubectl` et `helm` ne sont pas encore installés :</Guided>

```bash
curl -LO "https://dl.k8s.io/release/$(curl -Ls https://dl.k8s.io/release/stable.txt)/bin/linux/amd64/kubectl"
sudo install -m 755 kubectl /usr/local/bin/kubectl
curl -fsSL https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
```

<Note>Sur macOS, `brew install kubectl helm` fait la même chose. Garde kubectl à une version mineure près du cluster (<V name="K3S_VERSION" /> est une API 1.31) : plus loin, certaines commandes se comportent mal sans prévenir.</Note>

<Deep>Le kubeconfig contient un certificat client signé par la CA du cluster : c'est `cluster-admin`, sans expiration digne de ce nom. Traite ce fichier comme un mot de passe root. Les pages suivantes créent un ServiceAccount plus étroit pour la CI justement pour que ce fichier ne quitte jamais ton portable. Si tu gères plusieurs clusters, renomme le contexte : `kubectl config rename-context default k3s-lab`.</Deep>

<Check cmd="kubectl get nodes --no-headers | awk '{print $2}' | sort -u" expect="Ready" />

<Details summary="Si kubectl dit x509: certificate is valid for …">
Le certificat de l'API n'inclut pas <V name="NODE_IP" /> : l'option `--tls-san` manquait ou portait une autre adresse. Relance la commande d'installation sur le premier nœud avec le bon `--tls-san` ; le script est idempotent et ne réécrit que l'unité et le certificat.
</Details>

## Ingress : Traefik et ServiceLB

<Guided>k3s livre un ingress controller, Traefik, et un petit load-balancer, ServiceLB, qui publie tout Service `LoadBalancer` sur les adresses des nœuds eux-mêmes. À eux deux, ça veut dire : un objet `Ingress` avec un nom d'hôte, et ce nom répond sur le port 443 de chaque nœud. Rien à installer.</Guided>

```bash
kubectl -n kube-system get svc traefik
```

<Deep>ServiceLB (ex-klipper-lb) fait tourner un DaemonSet `svclb-traefik` dont les pods prennent les ports 80 et 443 de chaque nœud en `hostPort`, et la colonne `EXTERNAL-IP` du Service liste les adresses des nœuds. Ce n'est pas un vrai load balancer : un DNS qui pointe vers un nœud est un point d'entrée unique. En topologie HA, mets les trois adresses dans le DNS ou un load balancer devant. Traefik lui-même est déployé par le contrôleur HelmChart depuis `/var/lib/rancher/k3s/server/manifests/traefik.yaml` ; pour changer ses réglages, n'édite pas ce fichier (il est écrasé à la mise à jour) mais dépose un `HelmChartConfig` à côté. Pour utiliser un autre contrôleur, installe k3s avec `--disable traefik`.</Deep>

<Check cmd="kubectl -n kube-system rollout status deploy/traefik" expect={'deployment "traefik" successfully rolled out'} />

## cert-manager et un ClusterIssuer

<Guided>cert-manager surveille les objets `Ingress` : quand l'un porte son annotation, il obtient un certificat auprès d'un issuer et le range dans le Secret que nomme l'Ingress. Un `ClusterIssuer` est un issuer disponible dans tous les namespaces : tu le définis une fois.</Guided>

```bash
helm repo add jetstack https://charts.jetstack.io --force-update
helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager --create-namespace --set crds.enabled=true --wait
```

<Note>`crds.enabled=true` est l'option du chart depuis cert-manager 1.15 ; les charts plus anciens utilisaient `installCRDs=true`. Vérifie la version servie par le dépôt avec `helm search repo jetstack`.</Note>

<Check cmd="kubectl -n cert-manager rollout status deploy/cert-manager-webhook" expect={'deployment "cert-manager-webhook" successfully rolled out'} />

<When is="TLS" equals="letsencrypt">

<Annotated>

```yaml title="clusterissuer.yaml" {8,13}
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory   # (1)
    email: ${ACME_EMAIL}                                        # (2)
    privateKeySecretRef:
      name: letsencrypt-account-key                             # (3)
    solvers:
      - http01:
          ingress:
            ingressClassName: traefik                           # (4)
```

1. L'endpoint de production. Il est limité (cinq certificats identiques par semaine) : le temps des essais, pointe un second issuer vers `https://acme-staging-v02.api.letsencrypt.org/directory`, dont les certificats sont refusés par les navigateurs mais qui ne s'épuise jamais.
2. Ne sert qu'aux avertissements d'expiration.
3. La clé du compte ACME, générée par cert-manager au premier usage et gardée dans ce Secret du namespace `cert-manager`.
4. HTTP-01 : cert-manager crée un Ingress temporaire sur Traefik qui sert `/.well-known/acme-challenge/…`, Let's Encrypt le récupère sur le port 80, puis signe. Le nom doit résoudre publiquement vers le(s) nœud(s).

</Annotated>

```bash
kubectl apply -f clusterissuer.yaml
```

</When>

<When is="TLS" equals="selfsigned">

```yaml title="clusterissuer.yaml"
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: selfsigned
spec:
  selfSigned: {}
```

```bash
kubectl apply -f clusterissuer.yaml
```

<Note>Les navigateurs avertiront sur chaque certificat signé par cet issuer. Bien pour un lab sans nom public. Pour quelque chose que des utilisateurs ouvrent, passe le choix Certificats sur Let's Encrypt.</Note>

<Deep>Un montage de lab plus agréable, c'est une CA privée : cet issuer auto-signé signe un `Certificate` marqué `isCA: true`, un second `ClusterIssuer` de type `ca` utilise ce Secret, et chaque certificat applicatif chaîne dessus. Importe la CA une fois dans tes navigateurs et les avertissements disparaissent. La doc cert-manager a les trois manifestes sous « CA ».</Deep>

</When>

## Stockage : la classe local-path

<Guided>k3s arrive avec une StorageClass, `local-path`, marquée par défaut : un `PersistentVolumeClaim` reçoit un répertoire sur le nœud où le pod est placé. Rien à faire ici ; connais ses limites avant que la base de la page suivante repose dessus.</Guided>

```bash
kubectl get storageclass
```

<Deep>Les volumes vivent sous `/var/lib/rancher/k3s/storage/` sur le nœud. Deux conséquences : un pod qui en utilise un est cloué à ce nœud (le scheduler ne le déplacera pas), et perdre le disque du nœud perd les données. Pas de snapshot, pas de réplication, pas de redimensionnement. Parfait pour un lab, acceptable pour une prod à un nœud avec des sauvegardes. Pour ce qui doit survivre à un nœud, l'étape suivante habituelle sur k3s est Longhorn, qui réplique les volumes entre nœuds ; c'est hors de cette série.</Deep>

## Hello world derrière HTTPS

<Guided>Un Deployment, un Service, un Ingress avec l'annotation cert-manager : si <V name="APP_DOMAIN" /> répond avec un certificat valide, toute la chaîne (DNS, Traefik, cert-manager, l'issuer) fonctionne, et la page suivante peut se concentrer sur l'application.</Guided>

```yaml title="hello.yaml"
apiVersion: apps/v1
kind: Deployment
metadata:
  name: hello
spec:
  replicas: 1
  selector:
    matchLabels: { app: hello }
  template:
    metadata:
      labels: { app: hello }
    spec:
      containers:
        - name: whoami
          image: traefik/whoami:v1.10
          ports:
            - containerPort: 80
---
apiVersion: v1
kind: Service
metadata:
  name: hello
spec:
  selector: { app: hello }
  ports:
    - port: 80
      targetPort: 80
```

<When is="TLS" equals="letsencrypt">

```yaml title="hello-ingress.yaml" {6,10-11}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt
spec:
  ingressClassName: traefik
  tls:
    - hosts: ["${APP_DOMAIN}"]
      secretName: hello-tls
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: hello
                port: { number: 80 }
```

</When>

<When is="TLS" equals="selfsigned">

```yaml title="hello-ingress.yaml" {6,10-11}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello
  annotations:
    cert-manager.io/cluster-issuer: selfsigned
spec:
  ingressClassName: traefik
  tls:
    - hosts: ["${APP_DOMAIN}"]
      secretName: hello-tls
  rules:
    - host: ${APP_DOMAIN}
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: hello
                port: { number: 80 }
```

</When>

```bash
kubectl apply -f hello.yaml -f hello-ingress.yaml
kubectl get certificate hello-tls -w
```

<Deep>L'annotation fait créer par cert-manager un `Certificate` nommé d'après le Secret, puis un `Order`, un `Challenge`, et enfin le Secret `hello-tls` avec `tls.crt` et `tls.key`. Traefik surveille les Secrets référencés par les Ingress et sert le nouveau certificat sans redémarrer. `kubectl describe challenge` est l'endroit où regarder quand le certificat reste non Ready : il affiche l'URL exacte que Let's Encrypt n'a pas réussi à récupérer.</Deep>

<Check cmd="kubectl wait --for=condition=Ready certificate/hello-tls --timeout=120s" expect="certificate.cert-manager.io/hello-tls condition met" />

<When is="TLS" equals="letsencrypt">

<Check cmd="curl -sI https://${APP_DOMAIN} | head -1" expect="HTTP/2 200" />

</When>

<When is="TLS" equals="selfsigned">

<Check cmd="curl -skI https://${APP_DOMAIN} | head -1" expect="HTTP/2 200" />

</When>

<Details summary="Si le certificat reste non Ready">
Par ordre de probabilité :

- <V name="APP_DOMAIN" /> ne résout pas encore vers un nœud, ou résout vers une adresse privée que Let's Encrypt ne peut pas joindre. Fais un `dig +short` du nom depuis l'extérieur.
- Le port 80 est fermé quelque part (ufw, le pare-feu de l'hébergeur) : HTTP-01 en a besoin, même si les utilisateurs n'utilisent que le 443.
- Limite de débit atteinte après trop d'essais : `kubectl describe order` le dit. Utilise l'endpoint de staging jusqu'à ce que le montage soit bon.
- Le `Challenge` est `pending` avec un `wrong status code 404` : Traefik ne sert pas l'Ingress temporaire, regarde `kubectl -n kube-system logs deploy/traefik`.
</Details>

## Mettre à jour, désinstaller

Mettre k3s à jour, c'est relancer le script d'installation avec une version plus récente ; désinstaller, c'est un script par nœud.

<Deep>Pour mettre à jour, change <V name="K3S_VERSION" /> et relance exactement la commande d'installation sur chaque serveur, un à la fois, le premier nœud d'abord ; le script remplace le binaire et redémarre l'unité, les pods continuent de tourner. Reste à une version mineure par saut (1.31 → 1.32, jamais 1.31 → 1.33), et lis les notes de version k3s pour le datastore. Pour un cluster à beaucoup de nœuds, le system-upgrade-controller fait ça depuis un objet `Plan`. `/usr/local/bin/k3s-uninstall.sh` retire tout d'un serveur (`k3s-agent-uninstall.sh` sur un agent), y compris `/var/lib/rancher/k3s` et ses volumes : sauvegarde d'abord. Pour changer de datastore ou réinitialiser etcd, `k3s server --cluster-reset` sur un nœud reconstruit un cluster à un membre depuis sa propre copie.</Deep>

## Terminé

<When is="TOPOLOGY" equals="single">Un nœud</When><When is="TOPOLOGY" equals="ha">Trois serveurs avec un datastore répliqué</When>, piloté depuis ton portable, un ingress sur chaque nœud, des certificats qui s'émettent et se renouvellent seuls, et une StorageClass pour la base. Le hello-world peut partir :

```bash
kubectl delete -f hello-ingress.yaml -f hello.yaml
```

La page suivante déploie indicat sur ce cluster : namespace, secrets, PostgreSQL, l'application avec son Ingress sur <V name="APP_DOMAIN" />, et un autoscaler.
````

````yaml title="content/kubernetes/bootstrap-a-k3s-cluster/diagram.yaml"
# The gist: your laptop drives the API on the first node; browsers reach
# Traefik, which routes to the hello-world; cert-manager gets the certificate.
# Quick: five or six boxes. Guided adds cert-manager and the challenge.
# Deep adds ports, the storage class and etcd.
title: { en: "One API, one ingress, one certificate", fr: "Une API, un ingress, un certificat" }
caption:
  en: "kubectl on your laptop talks to the k3s API on the first node. Browsers reach Traefik on 443 on any node; it routes by host name to the hello-world Service. cert-manager gets the certificate and stores it in a Secret Traefik reads."
  fr: "kubectl sur ton portable parle à l'API k3s du premier nœud. Les navigateurs atteignent Traefik sur le 443 de n'importe quel nœud ; il route par nom d'hôte vers le Service hello-world. cert-manager obtient le certificat et le range dans un Secret que lit Traefik."

groups:
  - id: cluster
    label: { en: "Your cluster", fr: "Ton cluster" }
    desc:
      en: "Every box in here runs on the k3s nodes. One node in the lab topology; three servers sharing an embedded etcd in HA."
      fr: "Tout ce qui est ici tourne sur les nœuds k3s. Un nœud en topologie lab ; trois serveurs partageant un etcd embarqué en HA."

nodes:
  - id: laptop
    kind: client
    label: { en: "Your laptop", fr: "Ton portable" }
    sub: "kubectl · helm"
    desc:
      en: "Where you work. kubectl and helm read ~/.kube/config, which points at ${NODE_IP}:6443 with the cluster's admin credentials."
      fr: "Là où tu travailles. kubectl et helm lisent ~/.kube/config, qui pointe vers ${NODE_IP}:6443 avec les identifiants admin du cluster."
    deep:
      sub: "kubectl · helm · ~/.kube/config → https://${NODE_IP}:6443"
  - id: browser
    kind: user
    label: { en: "A browser", fr: "Un navigateur" }
    sub: "https://${APP_DOMAIN}"
    desc:
      en: "Anyone resolving ${APP_DOMAIN}. DNS must point the name at the node(s) before the certificate can be issued."
      fr: "Quiconque résout ${APP_DOMAIN}. Le DNS doit faire pointer le nom vers le(s) nœud(s) avant que le certificat puisse être émis."
  - id: node1
    kind: server
    label: { en: "Server node 1", fr: "Nœud serveur 1" }
    sub: "${NODE_IP} · k3s ${K3S_VERSION}"
    in: cluster
    focus: true
    desc:
      en: "The first k3s server: API server, scheduler, controllers, kubelet and the datastore, all in one binary. In the lab topology it is the whole cluster."
      fr: "Le premier serveur k3s : API server, scheduler, contrôleurs, kubelet et datastore, tout dans un binaire. En topologie lab, c'est tout le cluster."
    guided:
      sub: "${NODE_IP} · k3s ${K3S_VERSION} · API :6443"
    deep:
      sub: "${NODE_IP} · k3s ${K3S_VERSION} · API 6443 · kubelet 10250 · flannel VXLAN 8472/udp"
  - id: node2
    kind: server
    label: { en: "Server node 2", fr: "Nœud serveur 2" }
    sub: "${NODE2_IP}"
    in: cluster
    when: { is: TOPOLOGY, equals: ha }
    desc:
      en: "Joined with the token from node 1. Runs the same components and a second etcd member."
      fr: "Rejoint avec le token du nœud 1. Fait tourner les mêmes composants et un second membre etcd."
    deep:
      sub: "${NODE2_IP} · etcd member · 2379-2380"
  - id: node3
    kind: server
    label: { en: "Server node 3", fr: "Nœud serveur 3" }
    sub: "${NODE3_IP}"
    in: cluster
    when: { is: TOPOLOGY, equals: ha }
    desc:
      en: "Third etcd member. With three, the cluster keeps a quorum when one node is down."
      fr: "Troisième membre etcd. À trois, le cluster garde un quorum quand un nœud est arrêté."
    deep:
      sub: "${NODE3_IP} · etcd member · 2379-2380"
  - id: traefik
    kind: net
    label: { en: "Traefik", fr: "Traefik" }
    sub: ":80 :443 · Ingress"
    in: cluster
    desc:
      en: "The ingress controller k3s ships. ServiceLB publishes it on ports 80 and 443 of every node; it routes each request by host name."
      fr: "L'ingress controller livré avec k3s. ServiceLB le publie sur les ports 80 et 443 de chaque nœud ; il route chaque requête par nom d'hôte."
    deep:
      sub: "kube-system · Service LoadBalancer via ServiceLB · :80 :443 on every node · TLS ends here"
  - id: certmanager
    kind: service
    label: { en: "cert-manager", fr: "cert-manager" }
    sub: "ClusterIssuer"
    in: cluster
    level: guided
    desc:
      en: "Watches Ingresses carrying its annotation, orders a certificate from the ClusterIssuer and writes it into the Secret named in the Ingress. Renews it by itself."
      fr: "Surveille les Ingress qui portent son annotation, commande un certificat au ClusterIssuer et l'écrit dans le Secret nommé dans l'Ingress. Le renouvelle tout seul."
    deep:
      sub: "cert-manager ns · ClusterIssuer · Certificate → Secret hello-tls"
  - id: letsencrypt
    kind: cloud
    label: { en: "Let's Encrypt", fr: "Let's Encrypt" }
    sub: "ACME · HTTP-01"
    when: { is: TLS, equals: letsencrypt }
    desc:
      en: "The public CA. It checks that you control ${APP_DOMAIN} by fetching a token over plain HTTP on port 80, then signs the certificate."
      fr: "L'AC publique. Elle vérifie que tu contrôles ${APP_DOMAIN} en récupérant un jeton en HTTP sur le port 80, puis signe le certificat."
    deep:
      sub: "acme-v02.api.letsencrypt.org · HTTP-01 on :80 · 90-day certs"
  - id: hello
    kind: service
    label: { en: "hello-world", fr: "hello-world" }
    sub: "Deployment · Service · Ingress"
    in: cluster
    desc:
      en: "The smallest possible application, there only to prove that ingress and TLS work end to end. Deleted at the end."
      fr: "La plus petite application possible, là seulement pour prouver qu'ingress et TLS marchent de bout en bout. Supprimée à la fin."
    deep:
      sub: "Deployment hello (whoami) · Service :80 · Ingress host ${APP_DOMAIN} · Secret hello-tls"
  - id: storage
    kind: store
    label: { en: "local-path", fr: "local-path" }
    sub: "/var/lib/rancher/k3s/storage"
    in: cluster
    level: deep
    desc:
      en: "The default StorageClass: a directory on the node that holds the pod. Simple, fast, and tied to that node — the next page's database lives on it."
      fr: "La StorageClass par défaut : un répertoire sur le nœud qui héberge le pod. Simple, rapide, et lié à ce nœud — la base de la page suivante vit dessus."

edges:
  - from: laptop
    to: node1
    label: "kubectl"
    guided: { label: "kubectl · TCP 6443" }
    deep: { label: "L7 HTTPS (Kubernetes API) · L4 TCP 6443 · client cert" }
    desc:
      en: "Every kubectl and helm command is an HTTPS call to the API server, authenticated by the client certificate in your kubeconfig."
      fr: "Chaque commande kubectl ou helm est un appel HTTPS à l'API server, authentifié par le certificat client de ton kubeconfig."
  - from: node1
    to: node2
    label: "etcd"
    dashed: true
    when: { is: TOPOLOGY, equals: ha }
    deep: { label: "etcd raft · TCP 2379-2380" }
    desc:
      en: "The three servers replicate the cluster state between them. Writes need two of three to agree."
      fr: "Les trois serveurs répliquent l'état du cluster entre eux. Une écriture demande l'accord de deux sur trois."
  - from: node1
    to: node3
    label: "etcd"
    dashed: true
    when: { is: TOPOLOGY, equals: ha }
    deep: { label: "etcd raft · TCP 2379-2380" }
    desc:
      en: "Same replication to the third member."
      fr: "Même réplication vers le troisième membre."
  - from: browser
    to: traefik
    label: "HTTPS"
    deep: { label: "L7 HTTPS · L4 TCP 443 · SNI ${APP_DOMAIN}" }
    desc:
      en: "The browser connects to any node on 443; Traefik picks the certificate by the requested name."
      fr: "Le navigateur se connecte à n'importe quel nœud sur le 443 ; Traefik choisit le certificat d'après le nom demandé."
  - from: traefik
    to: hello
    label: "Ingress"
    guided: { label: "host ${APP_DOMAIN} → Service" }
    deep: { label: "Ingress host ${APP_DOMAIN} → Service hello :80 → pod :80" }
    desc:
      en: "The Ingress rule: this host name goes to this Service, which load-balances across the pods."
      fr: "La règle d'Ingress : ce nom d'hôte va vers ce Service, qui répartit sur les pods."
  - from: certmanager
    to: letsencrypt
    label: "ACME order"
    dashed: true
    level: guided
    when: { is: TLS, equals: letsencrypt }
    desc:
      en: "cert-manager asks for a certificate for ${APP_DOMAIN} and gets a challenge back."
      fr: "cert-manager demande un certificat pour ${APP_DOMAIN} et reçoit un défi en retour."
  - from: letsencrypt
    to: traefik
    label: "HTTP-01"
    dashed: true
    level: guided
    when: { is: TLS, equals: letsencrypt }
    deep: { label: "GET /.well-known/acme-challenge/… · TCP 80" }
    desc:
      en: "Let's Encrypt fetches the challenge token over port 80; cert-manager created a temporary Ingress to serve it."
      fr: "Let's Encrypt récupère le jeton du défi sur le port 80 ; cert-manager a créé un Ingress temporaire pour le servir."
  - from: certmanager
    to: hello
    label: "tls Secret"
    dashed: true
    level: guided
    desc:
      en: "The signed certificate lands in the Secret the Ingress names; Traefik picks it up without a restart."
      fr: "Le certificat signé atterrit dans le Secret nommé par l'Ingress ; Traefik le prend en compte sans redémarrer."
````

---

# Writing a page for Runfold

A self-contained brief. Hand it to a person or a model with a subject ("set up a VPS to host a Next.js site behind nginx with TLS") and you get back the files a page is made of. Nothing else is needed to write; the site's test suite then checks the result.

## 1. What the reader gets, and why it shapes the writing

A page is one tutorial. The reader fills in **their context once** — hostnames, ports, paths, and a few **choices** of stack (Debian or RHEL, nginx or Apache, TLS from Let's Encrypt or self-signed) — and every command on the page is rewritten with their values. Sections that do not apply to their choices disappear.

The reader also picks a **reading level**, once for the whole site:

| Level | EN / FR | What it shows |
|---|---|---|
| Run | Run / Automatique | Only the `<Run>` block: a script that does the whole page. "Too lazy to read? Run it." |
| Quick | Quick / Express | The plain paragraphs and the commands. "It worked? Fine." |
| Guided | Guided / Détaillé | Quick + the `<Guided>` explanations: why, and how, for someone who has never done it. |
| Deep | Deep / Exhaustif | Everything, plus `<Deep>`: mechanism, alternatives, gotchas, what happens under the hood. |

So a page is **written once, in layers**, not four times. Every paragraph you write belongs to a layer. Plain text is Quick; wrap the rest.

Each page also carries a **gist**: a small diagram at the top (boxes, containers, arrows, no coordinates) that follows the reader's choices and level. And each page exists in **English and French**, each with its own file, same structure.

## 2. The files

```
content/<series>/series.yaml                 what all pages of the series share (title, order, shared variables/choices)
content/<series>/<page>/tuto.yaml            this page's contract: metadata, its own variables/choices
content/<series>/<page>/page-en.mdx          the prose, English
content/<series>/<page>/page-fr.mdx          the prose, French (same heading skeleton)
content/<series>/<page>/diagram.yaml         the gist (optional but expected)
```

Slugs are kebab-case (`secure-ssh-on-a-fresh-server`). Variable and choice keys are `UPPER_SNAKE_CASE`.

A page **inherits** the series' groups, variables and choices and may add its own; it must **never redeclare** a key the series declares. Two different pages may declare the same key (the reader's value is shared across the series).

Deliver every file in full, each in its own fenced block with the path as title.

## 3. `series.yaml` (only when creating a series)

```yaml
title:   { en: NetBox, fr: NetBox }
summary:
  en: >-
    One or two sentences: what the series takes the reader from and to.
  fr: >-
    Une ou deux phrases : d'où part la série et où elle mène.
order: [install-from-scratch, configure-for-your-team]   # page slugs, reading order

groups:                        # sections of the "Your values" panel
  - id: host
    label: { en: Server, fr: Serveur }
    desc:  { en: Where it runs and how it is reached., fr: Où ça tourne et comment on l'atteint. }
  - id: auth
    label: { en: Directory, fr: Annuaire }
    when:  { flag: LDAP }      # the whole group disappears when the choice is off

vars:                          # same shape as in tuto.yaml, see below
  - key: NETBOX_HOST
    kind: hostname
    group: host
    default: netbox.example.com
    label:  { en: NetBox hostname, fr: Nom d'hôte de NetBox }
    hint:   { en: …, fr: … }
    impact: { en: …, fr: … }

choices:
  - key: OS
    type: select
    label: { en: Distribution, fr: Distribution }
    default: ubuntu
    options:
      - { value: ubuntu, label: { en: Ubuntu 24.04, fr: Ubuntu 24.04 } }
      - { value: rhel,   label: { en: RHEL 9 / Rocky / Alma, fr: RHEL 9 / Rocky / Alma } }
  - key: LDAP
    type: boolean
    label: { en: Authenticate against LDAP, fr: Authentifier via LDAP }
    default: false
```

## 4. `tuto.yaml`

```yaml
# Inherits from ../series.yaml: <list the inherited keys here as a comment>.
title:   { en: Secure SSH on a fresh server, fr: Sécuriser l'accès SSH d'un serveur neuf }
summary:
  en: >-
    Two or three lines. What the reader has at the end. Concrete.
  fr: >-
    Deux ou trois lignes. Ce que le lecteur a à la fin. Concret.
difficulty: beginner           # beginner | intermediate | advanced
tags: [ssh, linux, security]   # lowercase, used by search filters
authors: [thudal]
created: 2026-09-26            # YYYY-MM-DD
minutes: 10                    # hands-on time
validated: Debian 12 · Rocky 9 # upstream version / platform it was written for

groups:
  - id: access
    label: { en: Access, fr: Accès }

vars:
  - key: SSH_PORT
    kind: port                 # text | ip | cidr | port | hostname | domain | user | email | secret | path | url | sshkey
    group: access
    default: "1234"            # always a string; "" for none (secrets)
    label: { en: SSH port, fr: Port SSH }
    hint:                      # what it is, one or two sentences — shown in the ? tip
      en: The port SSH listens on after hardening. Anything but 22 removes most automated noise.
      fr: Le port sur lequel SSH écoutera après durcissement. Tout sauf 22 supprime l'essentiel du bruit automatisé.
    impact:                    # consequences of getting it wrong, where it is reused — also in the tip
      en: Reused by the firewall and fail2ban. Open it in the firewall before restarting SSH, or you lock yourself out.
      fr: Réutilisé par le pare-feu et fail2ban. Ouvre-le dans le pare-feu avant de redémarrer SSH, sinon tu te verrouilles dehors.
    when: { flag: FAIL2BAN }   # optional: only relevant when a choice holds

choices:
  - key: FAIL2BAN
    type: boolean
    label: { en: Ban brute-force attempts (fail2ban), fr: Bannir le brute force (fail2ban) }
    hint:  { en: Recommended on any server reachable from the internet., fr: Recommandé sur tout serveur joignable depuis internet. }
    default: true
```

Rules:
- `kind: secret` for passwords and tokens: masked in the panel, never printed in the print view, never carried by share links. `default: ""`.
- Every value a reader could reasonably change is a variable. Hard-code only true constants (a well-known port of a protocol, a package name).
- `hint` and `impact` are the **whole explanation** of a variable; the page prose should not repeat them.
- Anything the reader **must** provide (their public key, their token) is a variable with `required: true` and no example default (marked in the panel). A block that uses an empty or invalid value is marked red and cannot be copied. Never write an example that looks like code in a block (`ssh-ed25519 AAAA… user@host`, `YOUR.IP`, `<user>`): readers paste it as is (the test refuses these).
- `kind: sshkey` checks a whole OpenSSH public key line (`ssh-ed25519 AAAAC3…`).
- `${EDITOR}` is reserved: the app fills it with the reader's editor (vim by default, nano in the panel). Write `sudo ${EDITOR} /etc/x` whenever the reader edits a file by hand. Never declare it.
- Conditions (`when`, and `<When>` in prose): `{ flag: KEY }`, `{ notFlag: KEY }`, `{ is: KEY, equals: value }`, `{ is: KEY, oneOf: [a, b] }`. They may only reference choices.

## 5. `page-en.mdx` / `page-fr.mdx`

### Skeleton

```mdx
{/* First pass — to be validated against <official docs URL> before publishing. */}

One paragraph (Quick): what this page does, in what order, for whom.

<Run>

One sentence saying what the script does, on which machine, as which user, and what must be in place first.

```bash
#!/usr/bin/env bash
set -euo pipefail
# <Title> — ${MAIN_VAR}
… the whole page as a script, using ${VARS} …
```

</Run>

## Before you start

<Guided>What you need on hand: accounts, access, prerequisites, and where the commands run.</Guided>

<Check cmd="…" expect="…" />

## First step

Plain sentence saying what to do.

```bash
command with ${VARS}
```

<Guided>Why this step, what the command does, what to look at.</Guided>

<Deep>The mechanism, the alternative, the gotcha, the thing that bites in production.</Deep>

<Check cmd="…" expect="…" />

## Second step
…

## Done

What the reader now has. What the next page of the series does.
```

The **French file has the same headings, in the same order, the same count** (the test checks it). It is written in natural French with "tu" (tutoiement), not a word-for-word translation.

### The syntax

| Syntax | Meaning |
|---|---|
| `${SSH_PORT}` inside a code fence | replaced by the reader's value; the copy button copies the filled command |
| `<V name="SSH_PORT" />` | the same, inline in prose |
| plain markdown | visible from **Quick** |
| `<Guided>…</Guided>` | visible from **Guided** |
| `<Deep>…</Deep>` | visible at **Deep** only |
| `<Note>…</Note>` | an aside, from Guided |
| `<Warn>…</Warn>` | always visible: anything that can lock you out, lose data or cost money |
| `<When is="OS" equals="debian">…</When>` | conditional on a select choice; also `oneOf="a,b"` |
| `<When flag="FAIL2BAN">…</When>` / `<When notFlag="…">` | conditional on a boolean choice |
| `<Run>…</Run>` | the only thing shown at Run: a script or a single file that does the whole page |
| `<Check cmd="…" expect="…" />` | a verification with a checkbox, tracked per page. Also `id="…"` to name it. |
| `<Details summary="If it fails">…</Details>` | collapsible, closed by default |
| `<Tabs group="x"><Tab label="/etc/a.conf">…</Tab><Tab label="/etc/b.conf">…</Tab></Tabs>` | several files of one feature, side by side |
| ` ```ini title="/etc/x.conf" {1,3-5} ` | caption, highlighted lines |
| ` ```bash on="mac" ` · ` ```bash on="server" as="debian" ` · `as="${USERNAME}"` | where the block is typed and as whom: a coloured badge and edge, one colour per terminal |
| ` ```caddy file="/etc/caddy/Caddyfile" on="server" as="${USERNAME}" ` | **the whole content** of a file: shown with its path, copied as a ready `sudo tee … <<'EOF'` command (or as content only, for an editor). `sudo` is added outside /home and /tmp; force with `sudo` / `sudo="false"` |
| ` ```bash interactive ` | the command asks something (password, confirmation): badge, and it must be alone in its block |
| `$${X}` | a literal `${X}` (GitHub Actions `$${{ … }}`, JS template literals) |
| `<Annotated>` + `# (1)` markers at line ends + a numbered list after the fence | annotations: stripped at Quick, badges at Guided, inline at Deep |
| `## Heading` | a step: numbered automatically, listed in the outline. `###` for sub-steps. |

Details that matter:
- A `<Check>` is a real command whose expected output is **stable**: a version line, `active (running)`, an HTTP `200`. Two to six per page. `expect` may contain `${VARS}` but never a secret. If the command mixes `'` and `"`, write `cmd={"…"}`.
- Fences: ` ```bash ` for shell (a `$` prompt is drawn in front of what the reader types, not in front of comments or continuation lines), ` ```yaml `, ` ```python `, ` ```ini `, ` ```nginx `. One fence per command group; keep them short enough to read.
- Inside `<Run>`, the script must work with the variables and the choices: branch it with `<When>` blocks around separate fences when the stack differs. It runs unattended, so `set -euo pipefail`, no prompts, no `sudo` password expectations that the intro does not state.
- Do not put MDX components inside a fence, and do not put `\"` inside a component attribute (use `cmd={"…"}`).
- Blank line before and after every component, and around fences inside components.
- A component whose content spans several paragraphs or holds a list must be written as a block: the opening tag alone on its line, a blank line, the content, a blank line, the closing tag alone on its line. `<Guided>One line.</Guided>` is fine; `<Guided>Intro:\n\n1. item</Guided>` does not compile.
- New pages carry `status: draft` in `tuto.yaml` until their author has run them end to end.

Rules that come from a real run (the test suite checks the first four):
- **Where and as whom.** When a page uses more than one terminal (the laptop, the server as `debian`, the server as yourself), every block says so with `on=` / `as=`. A tired reader follows the badges, not the sentences.
- **One file, one block, complete.** A file is shown once, whole, with `file="…"`. Never "now add this line before `log {`". When a choice changes the file, `<When>` picks between complete versions of it. The Run script writes exactly the file the page shows.
- **Interactive commands alone.** `adduser`, `passwd`, `ssh-keygen` without `-N`, `ssh-copy-id`: alone in their block, marked `interactive`, or the rest of a paste lands in the password prompt.
- **No fake examples** in blocks (see `required` above).
- **A check proves the right thing.** A `<Check>` that tests an SSH login forces the method under test: `ssh -i KEY -o IdentitiesOnly=yes -o PasswordAuthentication=no -o BatchMode=yes …`. A check that passes thanks to a password before the step that forbids passwords is how people lock themselves out.
- **No unexplained jargon** in the Quick text: "the bare domain, `example.com`, nothing in front" rather than "the apex".

### Voice

Direct, concrete, an experienced engineer talking to a colleague. Short sentences. No marketing, no "simply", no "just". Say what a command does before showing it when it is not obvious. Say when a step is dangerous before the reader runs it (`<Warn>`). Name the file being edited. When a fact is from memory and must be confirmed, keep the common well-known form and add a `<Note>` saying where to confirm it.

Length: 200–400 lines per language file. Accuracy over volume.

## 6. `diagram.yaml` — the gist

```yaml
title:   { en: "One server, four pieces", fr: "Un serveur, quatre briques" }
caption:
  en: "One or two sentences read under the drawing."
  fr: "Une ou deux phrases lues sous le schéma."

groups:                                       # dashed containers
  - id: server
    label: { en: "Your server", fr: "Ton serveur" }
    desc:  { en: "…", fr: "…" }               # shown on hover

nodes:
  - id: users                                 # kebab-case
    kind: user                                # user | client | server | service | store | file | net | cloud (the icon)
    label: { en: "Your team", fr: "Ton équipe" }
    sub: "https://${NETBOX_HOST}"             # second line, may use ${VARS}
    desc:  { en: "What this piece does.", fr: "Ce que fait cette brique." }
  - id: nginx
    kind: net
    label: { en: "nginx", fr: "nginx" }
    sub: "443 → 8001"
    in: server                                # container
    when: { is: WEB, equals: nginx }          # follows the reader's choices
    desc:  { en: "…", fr: "…" }
    guided: { sub: ":443 → 127.0.0.1:8001" }  # overrides from Guided up
    deep:   { sub: "TLS ends here · :443 → 127.0.0.1:8001 · HTTP/1.1" }
  - id: firewall
    kind: net
    label: { en: "Firewall", fr: "Pare-feu" }
    in: server
    level: guided                             # appears from Guided up (level: deep for Deep only)
    desc:  { en: "…", fr: "…" }
  - id: netbox
    kind: server
    label: { en: "NetBox", fr: "NetBox" }
    in: server
    focus: true                               # the thing the page is about, drawn in the accent
    desc:  { en: "…", fr: "…" }

edges:
  - { from: users, to: nginx, label: "HTTPS", max: quick, desc: { en: "…", fr: "…" } }          # a simplification replaced from Guided
  - { from: users, to: firewall, label: "HTTPS · TCP 443", level: guided, deep: { label: "L7 HTTPS · L4 TCP 443 · L3 IP" }, desc: { en: "…", fr: "…" } }
  - { from: firewall, to: nginx, level: guided, desc: { en: "…", fr: "…" } }
  - { from: nginx, to: netbox, desc: { en: "…", fr: "…" } }
  - { from: netbox, to: ldap, label: "bind", dashed: true, when: { flag: LDAP }, desc: { en: "…", fr: "…" } }   # dashed = secondary relation
```

Rules:
- Solid arrows lay the boxes out left to right; dashed arrows are secondary and do not move anything. No coordinates ever.
- Quick shows 4–7 boxes; Guided adds the pieces a first-timer should know exist; Deep shows every layer (DNS, TLS, ports, sockets, units, log files).
- `desc` on every node, edge and group, both languages; it is what the reader gets on hover.
- Only reference variables and choices declared for the page (series + page).

## 7. Checklist before delivering

- [ ] Every user-facing string exists in `en` and `fr`.
- [ ] Every `${VAR}` and `<V name>` in the prose is declared (series or page); every `<When>` references a declared choice.
- [ ] No key redeclared that the series already declares.
- [ ] Same `##`/`###` headings, same order, same count, in both language files.
- [ ] A `<Run>` block that does the whole page with the variables.
- [ ] 2–6 `<Check>`s with stable expected output; no secret in `expect`.
- [ ] `<Warn>` before anything that locks out, deletes or costs.
- [ ] The first-pass comment at the top of each `.mdx`, naming the docs to validate against.
- [ ] `diagram.yaml` with `desc` everywhere, `focus` on one node, levels used.
- [ ] Every file delivered once, in full, as a `file="…"` block; the Run script writes the same content.
- [ ] `on=` / `as=` on every block when the page uses more than one terminal; `interactive` commands alone.
- [ ] What the reader must provide is a `required` variable, never an example in a block.

## 8. A worked request

> Write the page `host-a-nextjs-site` for a new series `vps` (title "A VPS from scratch"): from a freshly delivered Debian 12 VPS to a Next.js site served by nginx with a Let's Encrypt certificate, the app run by systemd, deployed by `git pull` + `npm run build`. Series variables: SERVER_IP (ip), USERNAME (user, the deploy user), DOMAIN (domain). Page variables: APP_DIR (path, default /srv/site), REPO_URL (url), NODE_MAJOR (text, default "22"), APP_PORT (port, default "3000"). Choices: TLS (select letsencrypt | selfsigned, default letsencrypt), PM (select systemd | pm2, default systemd). Deliver series.yaml, tuto.yaml, page-en.mdx, page-fr.mdx, diagram.yaml.

The answer is five fenced blocks, complete, following sections 3–6 and passing the checklist in 7.
