Everything is in etcd
The problem. Someone deletes a namespace full of config nobody kept, or the control-plane disk dies. Every object in the cluster lives in one database, etcd; without a tested backup, it is gone. This lesson is how to back it up.
What you need to know already: etcd's role (15.5), static pods and their manifests (18.3), TLS client certificates and the separate etcd CA (18.11), scp (18.5), systemd timers (2.16).
Every object you have created - Deployments, Secrets, ConfigMaps, RBAC, the Node objects, even Events - is a key in etcd. The apiserver is the only thing that talks to it. Lose etcd and you lose the cluster's memory: the containers keep running for a while, but nothing knows they should exist.
You can look (read-only curiosity; never write to it by hand). etcdctl is etcd's client; --endpoints = where etcd listens, --cacert/--cert/--key = the TLS files (next section), get PREFIX --prefix --keys-only = list the keys that start with PREFIX, without their values. On cp-1 (exit when you leave the lesson):
$ ssh cp-1
learner@cp-1:~$ sudo etcdctl --endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
get /registry/deployments/ --prefix --keys-only
/registry/deployments/kube-system/calico-kube-controllers
/registry/deployments/kube-system/coredns
/registry/deployments/kube-system/metrics-server
Keys are /registry/<resource>/<namespace>/<name> (Nodes are the historical /registry/minions/). Values are protobuf, not YAML - which is one more reason the apiserver, not you, reads them.
Quorum (the minimum number of members that must agree). etcd is a Raft cluster (Raft = the agreement algorithm etcd uses; each copy of etcd is a member): a write is committed when a majority of members agree. 3 members tolerate 1 failure, 5 tolerate 2; an even number buys nothing (4 still tolerates only 1). This lab has a single member on cp-1 - no tolerance at all, which is why the backup below is not optional.
Connecting: the TLS flags
kubeadm's etcd listens on https://127.0.0.1:2379 (and the node IP) and requires client certificates signed by the etcd CA. You never have to remember the paths - they are in the manifest:
learner@cp-1:~$ sudo grep -E 'advertise-client-urls|cert-file|key-file|trusted-ca-file' /etc/kubernetes/manifests/etcd.yaml
- --advertise-client-urls=https://10.64.0.10:2379
- --cert-file=/etc/kubernetes/pki/etcd/server.crt
- --key-file=/etc/kubernetes/pki/etcd/server.key
- --peer-cert-file=/etc/kubernetes/pki/etcd/peer.crt
- --peer-key-file=/etc/kubernetes/pki/etcd/peer.key
- --peer-trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt
- --trusted-ca-file=/etc/kubernetes/pki/etcd/ca.crt
trusted-ca-file -> your --cacert; cert-file/key-file -> your --cert/--key (any client cert signed by the etcd CA works: server, peer, healthcheck-client, or the apiserver's apiserver-etcd-client).
What each mistake looks like - learn to recognise them:
# no sudo: the keys are 0600 root
Error: open /etc/kubernetes/pki/etcd/server.key: permission denied
# --endpoints http://... (or no scheme and no cert flags): plaintext to a TLS port (waits 5s)
# (no scheme WITH --cacert/--cert uses TLS, so the default 127.0.0.1:2379 works; https:// is explicit)
{"level":"warn",...,"msg":"retrying of unary invoker failed",...,"error":"rpc error: code = DeadlineExceeded desc = latest balancer error: last connection error: connection error: desc = \"error reading server preface: http2: frame too large\""}
Error: context deadline exceeded
# wrong CA (e.g. the cluster ca.crt instead of etcd/ca.crt)
... "transport: authentication handshake failed: tls: failed to verify certificate: x509: certificate signed by unknown authority"
Error: context deadline exceeded
# no client cert
... "error reading server preface: remote error: tls: certificate required"
A quick health check before and after anything:
# an illustration: the etcd backup mission runs these on cp-1
sudo etcdctl --endpoints=https://127.0.0.1:2379 --cacert=... --cert=... --key=... endpoint health
https://127.0.0.1:2379 is healthy: successfully committed proposal: took = 8.911236ms
sudo etcdctl ... member list -w table
+------------------+---------+------+----------------------------+----------------------------+------------+
| ID | STATUS | NAME | PEER ADDRS | CLIENT ADDRS | IS LEARNER |
+------------------+---------+------+----------------------------+----------------------------+------------+
| 8e9e05c52164694d | started | cp-1 | https://10.64.0.10:2380 | https://10.64.0.10:2379 | false |
+------------------+---------+------+----------------------------+----------------------------+------------+
Taking the snapshot
learner@cp-1:~$ sudo mkdir -p /opt/backup
learner@cp-1:~$ sudo ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
snapshot save /opt/backup/etcd-$(date +%F).db
{"level":"info","ts":"2026-09-22T20:00:03.300Z","caller":"snapshot/v3_snapshot.go:83","msg":"created temporary db file","path":"/opt/backup/etcd-2026-09-22.db.part"}
{"level":"info",...,"msg":"fetching snapshot","endpoint":"https://127.0.0.1:2379"}
{"level":"info",...,"msg":"fetched snapshot","endpoint":"https://127.0.0.1:2379","size":"5.83 MB","took":"171.418802ms"}
{"level":"info",...,"msg":"saved","path":"/opt/backup/etcd-2026-09-22.db"}
Snapshot saved at /opt/backup/etcd-2026-09-22.db
Server version 3.6.0
ETCDCTL_API=3has been the default since etcd 3.4; it is harmless, and many guides and exam tasks still include it. With the old default (v2) you would get a completely different command set.- The file is written as
.partfirst and renamed when complete - a half snapshot never looks like a real one. - The snapshot is a point-in-time copy of the whole keyspace. It contains every Secret in plain protobuf (unless you configured encryption at rest). Treat the file like the cluster-admin credential it effectively is:
0600, encrypted storage, off the node.
Inspecting it: etcdutl (etcd 3.6)
etcd 3.5 deprecated etcdctl snapshot status and etcdctl snapshot restore; etcd 3.6 removed them. etcdctl talks to a running member (save); the offline tool etcdutl works on files (status, restore). On 3.6:
$ etcdctl snapshot status /opt/backup/etcd-2026-09-22.db
NAME:
snapshot - Manages etcd node snapshots
...
COMMANDS:
save Stores an etcd node backend snapshot to a given file
# an illustration: the etcd backup mission takes this snapshot on cp-1
$ sudo etcdutl snapshot status /opt/backup/etcd-2026-09-22.db -w table
+----------+----------+------------+------------+---------+
| HASH | REVISION | TOTAL KEYS | TOTAL SIZE | VERSION |
+----------+----------+------------+------------+---------+
| ccc2c4cc | 10194 | 180 | 5.83 MB | 3.6.0 |
+----------+----------+------------+------------+---------+
Older guides (and exam material written for etcd 3.5) use etcdctl snapshot restore. Know both names; use the one your etcd version has - etcdctl version tells you.
(simulator) On a real kubeadm node neither binary is installed by default - the etcd image contains them, and you install the matching release tarball from github.com/etcd-io/etcd (or apt install etcd-client, which is usually older). The lab put etcd v3.6.4's etcdctl and etcdutl in /usr/local/bin on cp-1, the way the exam environment does.
Off the node
A backup that lives only on cp-1 dies with cp-1. Copy it off - scp runs as you, so go through a file you own:
# an illustration: the etcd backup mission runs these on cp-1
sudo cp /opt/backup/etcd-2026-09-22.db ~ && sudo chown learner: ~/etcd-2026-09-22.db
scp cp-1:etcd-2026-09-22.db ~/backups/
In production that is a cron job (or a systemd timer, 2.16) that saves, checks snapshot status, uploads to object storage (a cloud file store reached over HTTP, like a network drive for backups), and alerts when the newest backup is older than it should be. The part people skip: practising the restore. The next lesson.
What you can now do
- Connect etcdctl to kubeadm's etcd with the TLS flags read from its manifest.
- Take a snapshot, check it with etcdutl, and copy it off the node.
- Recognise the three etcdctl connection errors (no sudo, wrong scheme/CA, no cert).