What a restore means
The problem. A backup nobody has restored is a hope, not a backup. The restore has one step that looks harmless and wipes the cluster if done wrong - this lesson is about getting it right.
What you need to know already: the etcd backup (18.20), static pod manifests (18.3), hostPath volumes and mountPath (16.37), sed (7.6).
Restoring etcd rewinds the entire cluster state to the moment of the snapshot: every object created since is gone, every object deleted since is back, including Nodes' last known state and your RBAC changes. It is the right tool when the state itself is lost or corrupted (an etcd disk died, someone deleted a namespace full of Secrets, a bad migration mangled CRDs). It is the wrong tool for "undo my Deployment change" - that is kubectl rollout undo (17.24) or git.
The procedure (single control-plane node)
1. Restore the snapshot into a new data directory. Never over the live one - etcdutl refuses anyway:
# an illustration: the restore mission runs these on cp-1
sudo etcdutl snapshot restore /opt/backup/etcd-snap.db --data-dir /var/lib/etcd
Error: data-dir "/var/lib/etcd" not empty or could not be read
sudo etcdutl snapshot restore /opt/backup/etcd-snap.db --data-dir /var/lib/etcd-restore
2026-09-22T20:00:03Z info snapshot/v3_snapshot.go:305 restoring snapshot {"path": "/opt/backup/etcd-snap.db", "wal-dir": "/var/lib/etcd-restore/member/wal", "data-dir": "/var/lib/etcd-restore", "snap-dir": "/var/lib/etcd-restore/member/snap", "initial-memory-map-size": 10737418240}
2026-09-22T20:00:03Z info membership/store.go:141 Trimming membership information from the backend...
2026-09-22T20:00:03Z info membership/cluster.go:421 added member {"cluster-id": "cdf818194e3a8c32", "local-member-id": "0", "added-peer-id": "8e9e05c52164694d", "added-peer-peer-urls": ["http://localhost:2380"], "added-peer-is-learner": false}
2026-09-22T20:00:03Z info snapshot/v3_snapshot.go:333 restored snapshot {"path": "/opt/backup/etcd-snap.db", ...}
Without --data-dir, etcdutl restores into ./default.etcd in whatever directory you happen to be in - a classic "I restored it, where did it go".
2. Point the etcd static pod at it. In etcd.yaml the container always reads --data-dir=/var/lib/etcd inside the container; what that path is on the host is the etcd-data hostPath volume. Change the hostPath:
volumes:
- hostPath:
path: /etc/kubernetes/pki/etcd
type: DirectoryOrCreate
name: etcd-certs
- hostPath:
path: /var/lib/etcd-restore # was /var/lib/etcd
type: DirectoryOrCreate
name: etcd-data
# an illustration: the restore mission runs these on cp-1
sudo cp /etc/kubernetes/manifests/etcd.yaml /root/etcd.yaml.bak
sudo sed -i 's#path: /var/lib/etcd$#path: /var/lib/etcd-restore#' /etc/kubernetes/manifests/etcd.yaml
The $ anchors the match so the /var/lib/etcd inside other paths is not touched. Always grep the file afterwards to see what actually changed.
3. Wait. The kubelet notices the manifest change, stops etcd, starts it on the restored data. The apiserver loses its backend for a moment (kubectl errors), then reconnects. crictl ps --name etcd shows a fresh container.
4. Verify. The objects deleted after the snapshot are back; objects created after it are gone. etcdctl endpoint health, kubectl get nodes, kubectl get ns.
(Also valid: restore into /var/lib/etcd after moving the old directory away, and leave the manifest untouched. Moving etcd.yaml out of the manifests dir first stops etcd so nothing writes while you swap directories.)
The trap: changing --data-dir instead
The flag --data-dir=/var/lib/etcd is the path inside the container. Change only the flag to /var/lib/etcd-restore and etcd writes to a directory that is not on any volume - inside the container's own filesystem, empty:
- etcd starts a brand-new, empty cluster there;
- the apiserver connects happily and finds nothing: no namespaces except the defaults, no Deployments, no CoreDNS, no RBAC beyond what it bootstraps;
- the Node objects re-register, so
kubectl get nodeslooks deceptively fine; - and the data vanishes again on the next container restart.
The symptom is "the restore wiped the cluster". The fix is the manifest: the hostPath must be the restored directory, and --data-dir must be inside the mountPath of that volume. If you change both, change them consistently.
With more than one control-plane node
Every member must be restored from the same snapshot, each with its own --name, --initial-cluster and --initial-advertise-peer-urls (the values from its etcd.yaml), and the old members must be stopped first. That is beyond this lab's single member, and exactly why "restore procedure" belongs in a practised runbook, not in your head at 3am.
What you can now do
- Restore a snapshot into a new directory and point etcd's hostPath at it.
- Explain the
--data-dir-only trap and its "empty cluster" symptom. - Say what a restore brings back and what it throws away.