OnCallReady

Lesson 18.11 · Kubernetes: Cluster Operations & Troubleshooting · 22 min read

kubeadm's certificates: who trusts whom, and for how long

In plain words

Imagine every door in a building has a guard, and everyone who walks through must show an ID card, while the guard shows theirs too. The cards are issued by the building's official card office, and each card expires after one year. The card office's own seal lasts ten years. When the cards expire, nobody gets through any door, even though nothing else is wrong.

A kubeadm cluster is that building. Every connection is mutual TLS, and kubeadm made three CAs (ca, etcd/ca, front-proxy-ca, 10 years) and every leaf certificate (1 year) in /etc/kubernetes/pki, plus client certificates inside kubeconfigs like admin.conf. kubeadm certs check-expiration shows what's left; kubeadm certs renew all renews them. The kubelet renews its own.

Why a cluster has so many certificates

The problem. Exactly one year after a kubeadm cluster is built, kubectl stops working for everyone with "x509: certificate has expired". Nothing was deployed. This lesson shows which certificates exist, how to read their expiry, and the renewal steps people forget.

What you need to know already: TLS, CAs, certificate chains and openssl x509 (9.15, 9.17), client certificates and CN/O (17.43), the kubeadm PKI list from init (18.6), static pods and the move-out/move-back restart (18.3).

Every connection inside a kubeadm cluster is mutual TLS (mTLS - both sides show a certificate, not just the server): both sides present a certificate, both sides check the other's against a CA. That is how the apiserver knows a request comes from the scheduler, and how the kubelet knows it is talking to the real apiserver. kubeadm creates it all at init:

On cp-1 (exit when you leave the lesson):

$ ssh cp-1
learner@cp-1:~$ sudo ls /etc/kubernetes/pki /etc/kubernetes/pki/etcd
/etc/kubernetes/pki:
apiserver-etcd-client.crt     apiserver-kubelet-client.key  ca.crt  front-proxy-ca.crt      front-proxy-client.key
apiserver-etcd-client.key     apiserver.crt                 ca.key  front-proxy-ca.key      sa.key
apiserver-kubelet-client.crt  apiserver.key                 etcd    front-proxy-client.crt  sa.pub

/etc/kubernetes/pki/etcd:
ca.crt  ca.key  healthcheck-client.crt  healthcheck-client.key  peer.crt  peer.key  server.crt  server.key

Three CAs, each signing a family:

CAsignsused for
ca (kubernetes)apiserverthe apiserver's serving cert - what kubectl verifies
apiserver-kubelet-clientapiserver -> kubelet (logs, exec)
admin/super-admin/controller-manager/scheduler .confclient certs embedded in those kubeconfigs
every kubelet's client certnode -> apiserver
etcd/caetcd/server, etcd/peer, etcd/healthcheck-client, apiserver-etcd-clienteverything that talks to etcd - a separate trust domain
front-proxy-cafront-proxy-clientthe apiserver talking to aggregated APIs (extra APIs served by add-ons such as metrics-server, 17.1)

sa.key/sa.pub is not a certificate: it is the key pair that signs and verifies ServiceAccount tokens.

Lifetimes: CAs 10 years, every leaf 1 year. That second number is the one that causes outages.

Reading expiry

The kubeadm way, on the control plane:

learner@cp-1:~$ sudo kubeadm certs check-expiration
[check-expiration] Reading configuration from the "kubeadm-config" ConfigMap in namespace "kube-system"...
[check-expiration] Use 'kubeadm init phase upload-config kubeadm --config your-config-file' to re-upload it.

CERTIFICATE                EXPIRES                  RESIDUAL TIME   CERTIFICATE AUTHORITY   EXTERNALLY MANAGED
admin.conf                 Sep 10, 2027 16:43 UTC   352d            ca                      no
apiserver                  Sep 10, 2027 16:43 UTC   352d            ca                      no
apiserver-etcd-client      Sep 10, 2027 16:43 UTC   352d            etcd-ca                 no
apiserver-kubelet-client   Sep 10, 2027 16:43 UTC   352d            ca                      no
controller-manager.conf    Sep 10, 2027 16:43 UTC   352d            ca                      no
etcd-healthcheck-client    Sep 10, 2027 16:43 UTC   352d            etcd-ca                 no
etcd-peer                  Sep 10, 2027 16:43 UTC   352d            etcd-ca                 no
etcd-server                Sep 10, 2027 16:43 UTC   352d            etcd-ca                 no
front-proxy-client         Sep 10, 2027 16:43 UTC   352d            front-proxy-ca          no
scheduler.conf             Sep 10, 2027 16:43 UTC   352d            ca                      no
super-admin.conf           Sep 10, 2027 16:43 UTC   352d            ca                      no

CERTIFICATE AUTHORITY   EXPIRES                  RESIDUAL TIME   EXTERNALLY MANAGED
ca                      Sep 07, 2036 16:43 UTC   9y              no
etcd-ca                 Sep 07, 2036 16:43 UTC   9y              no
front-proxy-ca          Sep 07, 2036 16:43 UTC   9y              no

An expired one shows <invalid> in RESIDUAL TIME. It needs sudo (it reads the keys) and it lists the kubeconfig certs by the file name (admin.conf) and the pki ones by their short name (apiserver).

The openssl way works for any certificate, on any node, and is what you use when kubeadm is not the tool that made it:

-in = the file; -noout = do not print the certificate itself; -subject -issuer -dates = print those fields; -ext subjectAltName = print the SAN extension:

$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -subject -issuer -dates
subject=CN=kube-apiserver
issuer=CN=kubernetes
notBefore=Sep 10 16:43:03 2026 GMT
notAfter=Sep 10 16:43:03 2027 GMT
$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -ext subjectAltName
X509v3 Subject Alternative Name:
    DNS:cp-1, DNS:kubernetes, DNS:kubernetes.default, DNS:kubernetes.default.svc, DNS:kubernetes.default.svc.cluster.local, IP Address:10.96.0.1, IP Address:10.64.0.10
$ sudo openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -checkend $((30*86400))
Certificate will not expire

-checkend N exits 1 if the cert expires within N seconds - the one-liner for a monitoring check. The certificate inside a kubeconfig is base64 in client-certificate-data; the pipe below finds that line (grep), keeps the value (awk '{print $2}' = the 2nd field, 7.8), decodes it and hands it to openssl (reading from stdin when there is no -in):

$ sudo grep client-certificate-data /etc/kubernetes/admin.conf | awk '{print $2}' | base64 -d | openssl x509 -noout -subject -enddate
subject=O=kubeadm:cluster-admins, CN=kubernetes-admin
notAfter=Sep 10 16:43:03 2027 GMT

O= is the group (RBAC binds kubeadm:cluster-admins to cluster-admin) and CN= the user. That is all "logging in" means with a client certificate.

The kubelet's certificates renew themselves

$ sudo openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -noout -subject -issuer -enddate
subject=O=system:nodes, CN=system:node:cp-1
issuer=CN=kubernetes
notAfter=Sep 10 16:47:03 2027 GMT

(system:node:cp-1 in group system:nodes is how every kubelet identifies itself.) rotateCertificates: true in /var/lib/kubelet/config.yaml makes the kubelet request a new client certificate (a CSR, auto-approved) as it approaches expiry and swap the kubelet-client-current.pem symlink. kubeadm's renew commands do not touch these - and do not need to.

What expiry looks like

The apiserver's serving cert expired - every client refuses the connection:

# an illustration (no ▶): the certificates mission makes this happen
$ kubectl get nodes
E0922 20:00:06.500514    4012 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://10.64.0.10:6443/api?timeout=32s\": tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:06Z is after 2026-09-07T20:00:03Z"
Unable to connect to the server: tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:06Z is after 2026-09-07T20:00:03Z

Every kubelet logs the same from the node side ("Error updating node status, will retry" err="... x509: certificate has expired ..."), so after the grace period all nodes go NotReady. Workloads keep running - the containers do not care - but nothing can be changed and nothing reschedules.

Your client cert (in ~/.kube/config) expired but the server's is fine - the apiserver rejects you instead:

# an illustration (no ▶)
$ kubectl get nodes
error: You must be logged in to the server (Unauthorized)

Two different messages, two different certificates. "x509 ... expired" = the server's. "Unauthorized" = yours (or a bad token).

Renewing

# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo kubeadm certs renew all
[renew] Reading configuration from the "kubeadm-config" ConfigMap in namespace "kube-system"...
[renew] FYI: You can look at this config file with 'kubectl -n kube-system get cm kubeadm-config -o yaml'

certificate embedded in the kubeconfig file for the admin to use and for kubeadm itself renewed
certificate for serving the Kubernetes API renewed
certificate the apiserver uses to access etcd renewed
certificate for the API server to connect to kubelet renewed
certificate embedded in the kubeconfig file for the controller manager to use renewed
certificate for liveness probes to healthcheck etcd renewed
certificate for etcd nodes to communicate with each other renewed
certificate for serving etcd renewed
certificate for the front proxy client renewed
certificate embedded in the kubeconfig file for the scheduler manager to use renewed
certificate embedded in the kubeconfig file for the super-admin renewed

Done renewing certificates. You must restart the kube-apiserver, kube-controller-manager, kube-scheduler and etcd, so that they can use the new certificates.

(When the apiserver cert has already expired, the first line instead says Error reading configuration from the Cluster. Falling back to default configuration - it cannot read the ConfigMap. Renewing works anyway.)

Renew re-issues the leaf certs with the same CA and the same SANs. It does not renew the CAs. Then the two steps everybody forgets:

1. Restart the components. They read their certificates at start. The files are new, the processes still hold the old ones - kubectl keeps failing with exactly the same x509 error until the apiserver restarts. Static pods do not have a restart command; move the manifests out and back, or stop the containers:

# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo mkdir -p /root/mf && sudo mv /etc/kubernetes/manifests/*.yaml /root/mf/
sudo crictl ps | grep -E 'kube-|etcd'          # wait until they are gone
sudo sh -c 'mv /root/mf/*.yaml /etc/kubernetes/manifests/'

Why sh -c on the way back: your shell expands the * before sudo runs. You can list /etc/kubernetes/manifests, so the first glob works; you cannot list /root (0700), so sudo mv /root/mf/*.yaml ... passes the literal string /root/mf/*.yaml and fails with mv: cannot stat '/root/mf/*.yaml': No such file or directory. sudo sh -c '...' makes root's shell do the expansion.

or, container by container: sudo crictl stop $(sudo crictl ps --name kube-apiserver -q)

2. Refresh your kubeconfig. admin.conf got a new client cert; your ~/.kube/config is a copy from last year and did not:

# an illustration: the renewal dance (the certificates mission does it on cp-1)
sudo cp /etc/kubernetes/admin.conf ~/.kube/config
$ sudo chown $(id -u):$(id -g) ~/.kube/config

Every other copy (your laptop, oncall-lab, the CI system) needs the new one too. scp runs as you on the remote side, so copy it to your home on cp-1 first - scp cp-1:/etc/kubernetes/admin.conf . fails with Permission denied.

The real fix: upgrade

kubeadm upgrade apply renews every leaf certificate as part of the upgrade. Kubernetes ships a minor release every ~4 months and each minor is supported for about a year, so a cluster that is upgraded on schedule never gets near its certificate expiry. "Everything broke after exactly one year" means "nobody upgraded this cluster for a year" - which is a second problem worth raising.

Monitor it anyway, alerting at 30 days: kubeadm certs check-expiration (or openssl x509 -checkend over /etc/kubernetes/pki) from a systemd timer (2.16), plus a check from outside that connects to :6443 and reads the serving cert's expiry (a "blackbox" probe - it tests like a client would, without looking inside).

Later (Ch 27): your monitoring system can run that TLS probe and alert on the days left.

What you can now do

Why it helps

"Everything broke after exactly one year" is one of the most famous kubeadm incidents: every kubectl command fails with x509: certificate has expired, all nodes go NotReady, workloads keep running but nothing can change. Knowing the fix (renew, then restart the static pods, then refresh every copy of admin.conf) turns a panic into a 15-minute runbook, and knowing the prevention (upgrade on schedule, alert at 30 days) means it never happens on your watch.

It also teaches you to read TLS errors precisely: x509 ... expired is the server's certificate, Unauthorized is your client certificate. And decoding a certificate with openssl x509 to see subject, SANs and dates is a skill you'll use for ingress TLS, service meshes and the admin exam.

Commands in this lesson

ssh openssl grep kubectl chown

FAQ

I renewed the certificates but kubectl still says x509 expired. Why?

The files are new, but the running components still hold the old certificates in memory; they only read them at start. Restart the static pods (move the manifests out of /etc/kubernetes/manifests and back, or crictl stop each container). Then refresh your ~/.kube/config from the renewed /etc/kubernetes/admin.conf, since your copy has last year's client certificate.

x509 expired or Unauthorized: which certificate is it?

x509: certificate has expired or is not yet valid from kubectl means the server's certificate (the API server's serving cert) failed your client's verification. error: You must be logged in to the server (Unauthorized) means the server rejected your credential: your client certificate in the kubeconfig is expired or wrong, or a token is invalid. Different certificates, different fixes.

Do I need to renew the kubelet certificates too?

Usually not. With rotateCertificates: true in the kubelet config (the kubeadm default), each kubelet requests a new client certificate through a CSR as expiry approaches and swaps the kubelet-client-current.pem symlink. kubeadm certs renew doesn't touch them and doesn't need to. Check with openssl x509 -in /var/lib/kubelet/pki/kubelet-client-current.pem -noout -enddate.

Does renewing change the CA?

No. kubeadm certs renew re-issues the leaf certificates with the same CA and the same SANs. The CAs last 10 years and rotating them is a separate, more involved procedure, because every kubeconfig and kubelet trusts the old CA. If you need new SANs (say, a load balancer name), you regenerate the API server certificate with an updated config instead.

How do I check a certificate's expiry without kubeadm?

With openssl: openssl x509 -in <file> -noout -subject -issuer -dates, -ext subjectAltName for the names, and -checkend 2592000 to exit non-zero if it expires within 30 days, which is handy for monitoring. For a kubeconfig, base64-decode client-certificate-data first and pipe it into openssl.

In an interview Mid

All kubectl commands fail with "x509: certificate has expired". How do you fix it?

kubeadm's leaf certificates are valid one year; a cluster nobody upgraded for a year hits this. "x509 expired" means the apiserver's serving certificate (your own client certificate expiring gives Unauthorized instead). On the control plane:

  1. sudo kubeadm certs check-expiration - which ones, how long left (<invalid> = expired). Or openssl x509 -in /etc/kubernetes/pki/apiserver.crt -noout -dates.
  2. sudo kubeadm certs renew all - new leaves, same CAs and SANs.
  3. Restart the static pods - they hold the old certificates in memory: move the manifests out of /etc/kubernetes/manifests and back, or crictl stop the containers.
  4. Refresh every kubeconfig copy: sudo cp /etc/kubernetes/admin.conf ~/.kube/config, plus laptops and CI.

The kubelet's own client certificate rotates itself. The real prevention: upgrade on schedule (kubeadm upgrade apply renews everything) and alert at 30 days with openssl x509 -checkend from a timer.

Also asked: Why does a Kubernetes cluster use so many certificates? · What is the difference between "x509: certificate has expired" and "Unauthorized"? · Which certificates does the kubelet renew by itself?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.