OnCallReady

Lesson 18.22 · Kubernetes: Cluster Operations & Troubleshooting · 8 min read

Restoring etcd from a snapshot

In plain words

Imagine a video game save file. Loading an old save takes you back to that exact moment: the treasure you found since is gone, the monster you defeated is back. It's the right move when your current game is corrupted beyond repair, and the wrong move if you only wanted to undo one silly choice.

Restoring etcd is loading a save for the whole cluster: everything created since the snapshot disappears, everything deleted since returns. The procedure: etcdutl snapshot restore into a new data directory, then point the etcd static pod's hostPath volume at it. The famous trap is changing --data-dir instead: that path is inside the container, so etcd starts fresh and empty and the cluster looks wiped.

What a restore means

The problem. A backup nobody has restored is a hope, not a backup. The restore has one step that looks harmless and wipes the cluster if done wrong - this lesson is about getting it right.

What you need to know already: the etcd backup (18.20), static pod manifests (18.3), hostPath volumes and mountPath (16.37), sed (7.6).

Restoring etcd rewinds the entire cluster state to the moment of the snapshot: every object created since is gone, every object deleted since is back, including Nodes' last known state and your RBAC changes. It is the right tool when the state itself is lost or corrupted (an etcd disk died, someone deleted a namespace full of Secrets, a bad migration mangled CRDs). It is the wrong tool for "undo my Deployment change" - that is kubectl rollout undo (17.24) or git.

The procedure (single control-plane node)

1. Restore the snapshot into a new data directory. Never over the live one - etcdutl refuses anyway:

# an illustration: the restore mission runs these on cp-1
sudo etcdutl snapshot restore /opt/backup/etcd-snap.db --data-dir /var/lib/etcd
Error: data-dir "/var/lib/etcd" not empty or could not be read
sudo etcdutl snapshot restore /opt/backup/etcd-snap.db --data-dir /var/lib/etcd-restore
2026-09-22T20:00:03Z	info	snapshot/v3_snapshot.go:305	restoring snapshot	{"path": "/opt/backup/etcd-snap.db", "wal-dir": "/var/lib/etcd-restore/member/wal", "data-dir": "/var/lib/etcd-restore", "snap-dir": "/var/lib/etcd-restore/member/snap", "initial-memory-map-size": 10737418240}
2026-09-22T20:00:03Z	info	membership/store.go:141	Trimming membership information from the backend...
2026-09-22T20:00:03Z	info	membership/cluster.go:421	added member	{"cluster-id": "cdf818194e3a8c32", "local-member-id": "0", "added-peer-id": "8e9e05c52164694d", "added-peer-peer-urls": ["http://localhost:2380"], "added-peer-is-learner": false}
2026-09-22T20:00:03Z	info	snapshot/v3_snapshot.go:333	restored snapshot	{"path": "/opt/backup/etcd-snap.db", ...}

Without --data-dir, etcdutl restores into ./default.etcd in whatever directory you happen to be in - a classic "I restored it, where did it go".

2. Point the etcd static pod at it. In etcd.yaml the container always reads --data-dir=/var/lib/etcd inside the container; what that path is on the host is the etcd-data hostPath volume. Change the hostPath:

  volumes:
  - hostPath:
      path: /etc/kubernetes/pki/etcd
      type: DirectoryOrCreate
    name: etcd-certs
  - hostPath:
      path: /var/lib/etcd-restore        # was /var/lib/etcd
      type: DirectoryOrCreate
    name: etcd-data
# an illustration: the restore mission runs these on cp-1
sudo cp /etc/kubernetes/manifests/etcd.yaml /root/etcd.yaml.bak
sudo sed -i 's#path: /var/lib/etcd$#path: /var/lib/etcd-restore#' /etc/kubernetes/manifests/etcd.yaml

The $ anchors the match so the /var/lib/etcd inside other paths is not touched. Always grep the file afterwards to see what actually changed.

3. Wait. The kubelet notices the manifest change, stops etcd, starts it on the restored data. The apiserver loses its backend for a moment (kubectl errors), then reconnects. crictl ps --name etcd shows a fresh container.

4. Verify. The objects deleted after the snapshot are back; objects created after it are gone. etcdctl endpoint health, kubectl get nodes, kubectl get ns.

(Also valid: restore into /var/lib/etcd after moving the old directory away, and leave the manifest untouched. Moving etcd.yaml out of the manifests dir first stops etcd so nothing writes while you swap directories.)

The trap: changing --data-dir instead

The flag --data-dir=/var/lib/etcd is the path inside the container. Change only the flag to /var/lib/etcd-restore and etcd writes to a directory that is not on any volume - inside the container's own filesystem, empty:

The symptom is "the restore wiped the cluster". The fix is the manifest: the hostPath must be the restored directory, and --data-dir must be inside the mountPath of that volume. If you change both, change them consistently.

With more than one control-plane node

Every member must be restored from the same snapshot, each with its own --name, --initial-cluster and --initial-advertise-peer-urls (the values from its etcd.yaml), and the old members must be stopped first. That is beyond this lab's single member, and exactly why "restore procedure" belongs in a practised runbook, not in your head at 3am.

What you can now do

Why it helps

"Restore etcd from /opt/backup/etcd-snap.db" is one of the most feared exam tasks, because a single wrong edit either loses the data or makes the cluster look empty. Knowing the hostPath versus --data-dir distinction turns it into a calm five-minute procedure.

In real operations it's rare but critical: the etcd disk died, someone deleted a namespace full of Secrets and CRDs, a bad migration mangled resources. Knowing what a restore does (rewinds everything, including RBAC and Node state) lets you argue for the right tool: often kubectl rollout undo, Git, or a namespace backup tool (Velero) is better. And practising it once is what makes the backup job actually worth something.

FAQ

Why restore into a new directory instead of /var/lib/etcd?

Because the live directory is in use and must not be overwritten; etcdutl refuses a non-empty data dir anyway ("not empty or could not be read"). Restore into something like /var/lib/etcd-restore, then point etcd at it. Alternatively, stop etcd (move its manifest out), move the old directory aside, restore into /var/lib/etcd, and put the manifest back.

I changed --data-dir in etcd.yaml and the cluster looks empty. What happened?

--data-dir is a path inside the container. The host directory is mounted through the etcd-data hostPath volume at /var/lib/etcd. Changing only the flag points etcd at an unmounted, empty path in the container's own filesystem, so it starts a brand-new cluster. Fix the hostPath path instead, and keep --data-dir inside the volume's mountPath.

Where did my restore go when I didn't pass --data-dir?

Into ./default.etcd in the directory you ran the command from. etcdutl uses that default when no --data-dir is given, which is a classic "I restored it, but where is it?". Always pass --data-dir explicitly and check the path in the output.

Why does kubectl get nodes look fine even after a bad restore?

Because the kubelets re-register their Node objects with whatever API server and etcd they find. So nodes appear, while namespaces, Deployments, CoreDNS and your RBAC are missing. Check kubectl get ns and a few known objects, not just nodes, to verify a restore.

Is restoring etcd the right way to undo a bad deployment?

No. A restore rewinds the entire cluster, discarding every change since the snapshot by every team. For one Deployment use kubectl rollout undo or revert in Git; for one namespace's objects, re-apply them from Git or restore them with a backup tool such as Velero. An etcd restore is for lost or corrupted state: a dead etcd disk, mass deletions, a broken migration.

In an interview Mid

How do you restore etcd from a snapshot on a single control-plane kubeadm cluster?

A restore rewinds the entire cluster to the snapshot - everything created since is gone. Use it when the state is lost, not to undo a deploy.

  1. Restore into a new data directory: sudo etcdutl snapshot restore /opt/backup/etcd-snap.db --data-dir /var/lib/etcd-restore (without --data-dir it lands in ./default.etcd).
  2. Back up etcd.yaml outside the manifests directory, then change the hostPath of the etcd-data volume to /var/lib/etcd-restore. Check with grep.
  3. Wait: the kubelet restarts etcd on the restored data; the apiserver reconnects.
  4. Verify: etcdctl endpoint health, kubectl get ns, the deleted objects are back.

The trap: changing only the --data-dir flag. That path is inside the container; pointing it at a directory on no volume makes etcd start an empty cluster - "the restore wiped the cluster", and it vanishes again on the next restart. With several control-plane nodes, every member is restored from the same snapshot with its own member settings.

Also asked: What are the risks of an etcd restore? · When would you not use an etcd restore to fix a problem? · How do you verify a restore worked?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.