OnCallReady

Kubernetes · 5 min read

kubectl: x509: certificate has expired or is not yet valid (renewing kubeadm certificates)

A kubeadm cluster a year old and never upgraded stops answering kubectl. Check expiry, renew, restart the static pods, refresh your kubeconfig.

terminal
$ kubectl get nodes
E0922 20:00:04.500676   4012 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://10.64.0.10:6443/api?timeout=32s\": tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:04Z is after 2026-09-20T19:58:53Z"
Unable to connect to the server: tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:04Z is after 2026-09-20T19:58:53Z

Monday morning, and kubectl is dead for everyone. Nothing was deployed over the weekend. Oddly, the pods still serve traffic.

What is actually happening

A kubeadm cluster runs on certificates. A CA (certificate authority) signs leaf certificates for every component: the API server's serving certificate, the client certificates that the controller manager, the scheduler and the API server itself use to talk to etcd and the kubelets, and the client certificate inside admin.conf. kubeadm makes the CAs valid for 10 years and the leaf certificates for 1 year.

The design assumes you upgrade. kubeadm upgrade apply renews every leaf certificate by default, so a cluster upgraded at least once a year never sees this. Leave a cluster alone for 12 months and the leaves all expire together, to the minute: they were all issued by kubeadm init.

Read the error closely. tls: failed to verify certificate comes from the client, which is checking the server's certificate. The certificate the API server presents expired on 20 September, so every client refuses it. That includes kubectl, the controller manager, the scheduler and the kubelets.

Pods keep serving because they don't need the control plane. Containers run on the nodes under the kubelet and the container runtime, and Service routing rules already sit in each node's kernel. Even a container that crashes now is restarted locally by the kubelet: the CrashLoopBackOff restart loop runs on the node. What stops is change: no deploys, no rescheduling, no self-healing, until the API server is trusted again.

Diagnosis

1. Go to the control plane node

kubectl is the thing that's failing, so it can't help. SSH to the control plane node and ask kubeadm:

terminal
$ sudo kubeadm certs check-expiration
[check-expiration] Error reading configuration from the Cluster. Falling back to default configuration

CERTIFICATE                EXPIRES                  RESIDUAL TIME   CERTIFICATE AUTHORITY   EXTERNALLY MANAGED
admin.conf                 Sep 20, 2026 19:58 UTC   <invalid>       ca                      no
apiserver                  Sep 20, 2026 19:58 UTC   <invalid>       ca                      no
apiserver-etcd-client      Sep 20, 2026 19:58 UTC   <invalid>       etcd-ca                 no
apiserver-kubelet-client   Sep 20, 2026 19:58 UTC   <invalid>       ca                      no
controller-manager.conf    Sep 20, 2026 19:58 UTC   <invalid>       ca                      no
etcd-healthcheck-client    Sep 20, 2026 19:58 UTC   <invalid>       etcd-ca                 no
etcd-peer                  Sep 20, 2026 19:58 UTC   <invalid>       etcd-ca                 no
etcd-server                Sep 20, 2026 19:58 UTC   <invalid>       etcd-ca                 no
front-proxy-client         Sep 20, 2026 19:58 UTC   <invalid>       front-proxy-ca          no
scheduler.conf             Sep 20, 2026 19:58 UTC   <invalid>       ca                      no
super-admin.conf           Sep 20, 2026 19:58 UTC   <invalid>       ca                      no

CERTIFICATE AUTHORITY   EXPIRES                  RESIDUAL TIME   EXTERNALLY MANAGED
ca                      Sep 17, 2035 19:58 UTC   8y              no
etcd-ca                 Sep 17, 2035 19:58 UTC   8y              no
front-proxy-ca          Sep 17, 2035 19:58 UTC   8y              no

Every leaf shows <invalid>, and the CAs are fine. The first line is expected: kubeadm can't read its config from a cluster it can't reach.

2. Confirm on the file, if you want proof

terminal
$ sudo openssl x509 -noout -subject -enddate -in /etc/kubernetes/pki/apiserver.crt
subject=CN=kube-apiserver
notAfter=Sep 20 19:58:53 2026 GMT

The kubelet is a systemd service, so its journal tells the same story from the node's side (the same two commands as when the kubelet refuses to start because of swap):

terminal
$ sudo journalctl -u kubelet -n 1 --no-pager
Sep 22 20:00:04 cp-1 kubelet[852]: E0922 20:00:04.000476     852 kubelet_node_status.go:548] "Error updating node status, will retry" err="error getting node \"cp-1\": Get \"https://10.64.0.10:6443/api/v1/nodes/cp-1?resourceVersion=0&timeout=10s\": tls: failed to verify certificate: x509: certificate has expired or is not yet valid: current time 2026-09-22T20:00:04Z is after 2026-09-20T19:58:53Z"

The fix

1. Renew

terminal
$ sudo kubeadm certs renew all
[renew] Error reading configuration from the Cluster. Falling back to default configuration

certificate embedded in the kubeconfig file for the admin to use and for kubeadm itself renewed
certificate for serving the Kubernetes API renewed
...
certificate embedded in the kubeconfig file for the super-admin renewed

Done renewing certificates. You must restart the kube-apiserver, kube-controller-manager, kube-scheduler and etcd, so that they can use the new certificates.

The last line matters. Right after the renewal, kubectl still fails with the same x509 error. The API server loaded its certificate when it started, 12 days ago, and still serves the old one from memory.

2. Restart the four static pods

The control plane runs as static pods: the kubelet starts them from files in /etc/kubernetes/manifests. Restarting the kubelet won't help, because running static pods keep running. Move the manifests out, wait for the kubelet to stop the pods (about 20 seconds), then move them back:

output
sudo mkdir -p /root/mf
sudo mv /etc/kubernetes/manifests/{etcd,kube-apiserver,kube-controller-manager,kube-scheduler}.yaml /root/mf/
# wait until crictl ps no longer shows them
sudo mv /root/mf/*.yaml /etc/kubernetes/manifests/

3. Expect a second error, and refresh your kubeconfig

terminal
$ kubectl get nodes
error: You must be logged in to the server (Unauthorized)

That's progress. The server is trusted now, but your ~/.kube/config still carries the old admin client certificate, which has also expired. Copy the renewed /etc/kubernetes/admin.conf from the control plane to ~/.kube/config:

terminal
$ kubectl get nodes
NAME       STATUS   ROLES           AGE    VERSION
cp-1       Ready    control-plane   367d   v1.34.1
worker-1   Ready    <none>          367d   v1.34.1
worker-2   Ready    <none>          367d   v1.34.1

The kubelets reconnect by themselves. Their own client certificates rotate automatically, so they were never the problem. Every other copy of admin.conf is, though: CI jobs, monitoring, deployment tools. Each one fails with Unauthorized until it gets a fresh copy. A copied admin.conf is also a credential to the whole cluster sitting in a CI system, with all the risks of a secret leaked in a CI log. That's the best argument for giving those systems ServiceAccount tokens or OIDC instead of a copied admin certificate.

How to prevent it

  • Upgrade at least once a year. kubeadm upgrade apply renews the leaf certificates as a side effect, and you get the security fixes too.
  • Alert at 30 days. Run kubeadm certs check-expiration on a timer, and probe the API server's certificate from outside (for example the blackbox exporter's probe_ssl_earliest_cert_expiry).
  • Know the CA date. If the 10-year CA expires, renewing leaves isn't enough: every kubeconfig and every kubelet needs new credentials. That's a planned migration, not an incident fix.

Practise it

Chapter 18's incident "Monday, everything says x509" expires a kubeadm cluster's certificates for real. You renew them, restart the control plane and meet the second Unauthorized error yourself.

OnCallReady is free, with no ads and no tracking. RSS · All posts