Why a cluster has so many certificates
The problem. Exactly one year after a kubeadm cluster is built, kubectl stops working for everyone with "x509: certificate has expired". Nothing was deployed. This lesson shows which certificates exist, how to read their expiry, and the renewal steps people forget.
What you need to know already: TLS, CAs, certificate chains and openssl x509 (9.15, 9.17), client certificates and CN/O (17.43), the kubeadm PKI list from init (18.6), static pods and the move-out/move-back restart (18.3).
Every connection inside a kubeadm cluster is mutual TLS (mTLS - both sides show a certificate, not just the server): both sides present a certificate, both sides check the other's against a CA. That is how the apiserver knows a request comes from the scheduler, and how the kubelet knows it is talking to the real apiserver. kubeadm creates it all at init:
On cp-1 (exit when you leave the lesson):
$ ssh cp-1
learner@cp-1:~$ sudo ls /etc/kubernetes/pki /etc/kubernetes/pki/etcd
/etc/kubernetes/pki:
apiserver-etcd-client.crt apiserver-kubelet-client.key ca.crt front-proxy-ca.crt front-proxy-client.key
apiserver-etcd-client.key apiserver.crt ca.key front-proxy-ca.key sa.key
apiserver-kubelet-client.crt apiserver.key etcd front-proxy-client.crt sa.pub
/etc/kubernetes/pki/etcd:
ca.crt ca.key healthcheck-client.crt healthcheck-client.key peer.crt peer.key server.crt server.key
Three CAs, each signing a family:
| CA | signs | used for |
|---|---|---|
ca (kubernetes) | apiserver | the apiserver's serving cert - what kubectl verifies |
apiserver-kubelet-client | apiserver -> kubelet (logs, exec) | |
| admin/super-admin/controller-manager/scheduler .conf | client certs embedded in those kubeconfigs | |
| every kubelet's client cert | node -> apiserver | |
etcd/ca | etcd/server, etcd/peer, etcd/healthcheck-client, apiserver-etcd-client | everything that talks to etcd - a separate trust domain |
front-proxy-ca | front-proxy-client | the apiserver talking to aggregated APIs (extra APIs served by add-ons such as metrics-server, 17.1) |
sa.key/sa.pub is not a certificate: it is the key pair that signs and verifies ServiceAccount tokens.
Lifetimes: CAs 10 years, every leaf 1 year. That second number is the one that causes outages.
Reading expiry
The kubeadm way, on the control plane:
learner@cp-1:~$ sudo kubeadm certs check-expiration
[check-expiration] Reading configuration from the "kubeadm-config" ConfigMap in namespace "kube-system"...
[check-expiration] Use 'kubeadm init phase upload-config kubeadm --config your-config-file' to re-upload it.
CERTIFICATE EXPIRES RESIDUAL TIME CERTIFICATE AUTHORITY EXTERNALLY MANAGED
admin.conf Sep 10, 2027 16:43 UTC 352d ca no
apiserver Sep 10, 2027 16:43 UTC 352d ca no
apiserver-etcd-client Sep 10, 2027 16:43 UTC 352d etcd-ca no
apiserver-kubelet-client Sep 10, 2027 16:43 UTC 352d ca no
controller-manager.conf Sep 10, 2027 16:43 UTC 352d ca no
etcd-healthcheck-client Sep 10, 2027 16:43 UTC 352d etcd-ca no
etcd-peer Sep 10, 2027 16:43 UTC 352d etcd-ca no
etcd-server Sep 10, 2027 16:43 UTC 352d etcd-ca no
front-proxy-client Sep 10, 2027 16:43 UTC 352d front-proxy-ca no
scheduler.conf Sep 10, 2027 16:43 UTC 352d ca no
super-admin.conf Sep 10, 2027 16:43 UTC 352d ca no
CERTIFICATE AUTHORITY EXPIRES RESIDUAL TIME EXTERNALLY MANAGED
ca Sep 07, 2036 16:43 UTC 9y no
etcd-ca Sep 07, 2036 16:43 UTC 9y no
front-proxy-ca Sep 07, 2036 16:43 UTC 9y no
An expired one shows <invalid> in RESIDUAL TIME. It needs sudo (it reads the keys) and it lists the kubeconfig certs by the file name (admin.conf) and the pki ones by their short name (apiserver).
The openssl way works for any certificate, on any node, and is what you use when kubeadm is not the tool that made it:
-in = the file; -noout = do not print the certificate itself; -subject -issuer -dates = print those fields; -ext subjectAltName = print the SAN extension:
$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -subject -issuer -dates
subject=CN=kube-apiserver
issuer=CN=kubernetes
notBefore=Sep 10 16:43:03 2026 GMT
notAfter=Sep 10 16:43:03 2027 GMT
$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -ext subjectAltName
X509v3 Subject Alternative Name:
DNS:cp-1, DNS:kubernetes, DNS:kubernetes.default, DNS:kubernetes.default.svc, DNS:kubernetes.default.svc.cluster.local, IP Address:10.96.0.1, IP Address:10.64.0.10
$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -checkend $((30*86400))
Certificate will not expire
-checkend N exits 1 if the cert expires within N seconds - the one-liner for a monitoring check. The certificate inside a kubeconfig is base64 in client-certificate-data; the pipe below finds that line (grep), keeps the value (awk '{print $2}' = the 2nd field, 7.8), decodes it and hands it to openssl (reading from stdin when there is no -in):
$ sudo grep client-certificate-data /etc/kubernetes/admin.conf | awk '{print $2}' | base64 -d | openssl x509 -noout -subject -enddate
subject=O=kubeadm:cluster-admins, CN=kubernetes-admin
notAfter=Sep 10 16:43:03 2027 GMT
O= is the group (RBAC binds kubeadm:cluster-admins to cluster-admin) and CN= the user. That is all "logging in" means with a client certificate.
The kubelet's certificates renew themselves
$ sudo openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -noout -subject -issuer -enddate
subject=O=system:nodes, CN=system:node:cp-1
issuer=CN=kubernetes
notAfter=Sep 10 16:47:03 2027 GMT
(system:node:cp-1 in group system:nodes is how every kubelet identifies itself.) rotateCertificates: true in /var/lib/kubelet/config.yaml makes the kubelet request a new client certificate (a CSR, auto-approved) as it approaches expiry and swap the kubelet-client-current.pem symlink. kubeadm's renew commands do not touch these - and do not need to.
What expiry looks like
The apiserver's serving cert expired - every client refuses the connection:
# an illustration (no ▶): the certificates mission makes this happen
$ kubectl get nodes
E0922 20:00:06.500514 4012 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://10.64.0.10:6443/api?timeout=32s\": tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:06Z is after 2026-09-07T20:00:03Z"
Unable to connect to the server: tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:06Z is after 2026-09-07T20:00:03Z
Every kubelet logs the same from the node side ("Error updating node status, will retry" err="... x509: certificate has expired ..."), so after the grace period all nodes go NotReady. Workloads keep running - the containers do not care - but nothing can be changed and nothing reschedules.
Your client cert (in ~/.kube/config) expired but the server's is fine - the apiserver rejects you instead:
# an illustration (no ▶)
$ kubectl get nodes
error: You must be logged in to the server (Unauthorized)
Two different messages, two different certificates. "x509 ... expired" = the server's. "Unauthorized" = yours (or a bad token).
Renewing
# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo kubeadm certs renew all
[renew] Reading configuration from the "kubeadm-config" ConfigMap in namespace "kube-system"...
[renew] FYI: You can look at this config file with 'kubectl -n kube-system get cm kubeadm-config -o yaml'
certificate embedded in the kubeconfig file for the admin to use and for kubeadm itself renewed
certificate for serving the Kubernetes API renewed
certificate the apiserver uses to access etcd renewed
certificate for the API server to connect to kubelet renewed
certificate embedded in the kubeconfig file for the controller manager to use renewed
certificate for liveness probes to healthcheck etcd renewed
certificate for etcd nodes to communicate with each other renewed
certificate for serving etcd renewed
certificate for the front proxy client renewed
certificate embedded in the kubeconfig file for the scheduler manager to use renewed
certificate embedded in the kubeconfig file for the super-admin renewed
Done renewing certificates. You must restart the kube-apiserver, kube-controller-manager, kube-scheduler and etcd, so that they can use the new certificates.
(When the apiserver cert has already expired, the first line instead says Error reading configuration from the Cluster. Falling back to default configuration - it cannot read the ConfigMap. Renewing works anyway.)
Renew re-issues the leaf certs with the same CA and the same SANs. It does not renew the CAs. Then the two steps everybody forgets:
1. Restart the components. They read their certificates at start. The files are new, the processes still hold the old ones - kubectl keeps failing with exactly the same x509 error until the apiserver restarts. Static pods do not have a restart command; move the manifests out and back, or stop the containers:
# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo mkdir -p /root/mf && sudo mv /etc/kubernetes/manifests/*.yaml /root/mf/
sudo crictl ps | grep -E 'kube-|etcd' # wait until they are gone
sudo sh -c 'mv /root/mf/*.yaml /etc/kubernetes/manifests/'
Why sh -c on the way back: your shell expands the * before sudo runs. You can list /etc/kubernetes/manifests, so the first glob works; you cannot list /root (0700), so sudo mv /root/mf/*.yaml ... passes the literal string /root/mf/*.yaml and fails with mv: cannot stat '/root/mf/*.yaml': No such file or directory. sudo sh -c '...' makes root's shell do the expansion.
or, container by container: sudo crictl stop $(sudo crictl ps --name kube-apiserver -q)
- the kubelet starts a fresh one (ATTEMPT goes up).
2. Refresh your kubeconfig. admin.conf got a new client cert; your ~/.kube/config is a copy from last year and did not:
# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo cp /etc/kubernetes/admin.conf ~/.kube/config
$ sudo chown $(id -u):$(id -g) ~/.kube/config
Every other copy (your laptop, oncall-lab, the CI system) needs the new one too. scp runs as you on the remote side, so copy it to your home on cp-1 first - scp cp-1:/etc/kubernetes/admin.conf . fails with Permission denied.
The real fix: upgrade
kubeadm upgrade apply renews every leaf certificate as part of the upgrade. Kubernetes ships a minor release every ~4 months and each minor is supported for about a year, so a cluster that is upgraded on schedule never gets near its certificate expiry. "Everything broke after exactly one year" means "nobody upgraded this cluster for a year" - which is a second problem worth raising.
Monitor it anyway, alerting at 30 days: kubeadm certs check-expiration (or openssl x509 -checkend over /etc/kubernetes/pki) from a systemd timer (2.16), plus a check from outside that connects to :6443 and reads the serving cert's expiry (a "blackbox" probe - it tests like a client would, without looking inside).
Later (Ch 27): your monitoring system can run that TLS probe and alert on the days left.
What you can now do
- List the kubeadm PKI and say which CA signs what.
- Read any certificate's expiry and SANs with openssl, including the one inside a kubeconfig.
- Renew, restart the four static pods, and refresh every kubeconfig copy - and tell "x509 expired" from "Unauthorized".