OnCallReady

Lesson 34.10 · Kubernetes: Ingress, Gateway API & Service Mesh · 16 min read

Certificates that renew themselves: cert-manager and ACME

In plain words

Imagine a passport office that only gives out passports valid for three months. Instead of everyone queueing every quarter, you hire an assistant who keeps a list of your family's passports, books the appointment a month before each one expires, brings the proof of address the office asks for, and puts the new passport in the drawer.

cert-manager is that assistant for TLS certificates. You describe the certificate you want (names, which issuer); it creates a private key, asks an ACME server for a certificate, proves you control the name by answering an HTTP-01 challenge (a small token file served on your domain), stores the result in a Secret and renews it before it expires. The lab uses Pebble, a test ACME server, so you can watch every object in the chain: Certificate, CertificateRequest, Order and Challenge.

Nobody should copy certificates by hand

In 16.22 you made a TLS Secret from files. In production certificates expire every 90 days (Let's Encrypt; 47 days by 2029 under the CA/Browser Forum's plan), and an expired certificate is one of the most common self-inflicted outages. cert-manager is the controller that requests certificates, stores them in Secrets, and renews them long before they expire.

What you need to know already: certificates, chains, SANs and expiry (9.15-9.17), TLS Secrets on an Ingress (16.22), the ingress-nginx controller (16.19), CRDs and controllers (16.26), DNS names in the lab zone (8.16).

The objects

Issuer / ClusterIssuer   WHERE certificates come from: an ACME server (Let's Encrypt), your own CA,
                         Vault, or self-signed. Issuer = one namespace, ClusterIssuer = all of them.
Certificate              WHAT you want: dnsNames, the issuerRef, the Secret to write (secretName),
                         duration, renewBefore. cert-manager keeps that Secret valid forever.
CertificateRequest       one attempt to get it signed (named <certificate>-<revision>)
Order                    (ACME only) the order placed with the ACME server for those names
Challenge                (ACME only) one proof per name that you control it: HTTP-01 or DNS-01

You write the first two; cert-manager writes the other three while it works, and deletes the Order's Challenges when they pass. When something is stuck, the reason is on the lowest object that exists - usually the Challenge.

ACME and HTTP-01 in one picture

ACME (RFC 8555) is the protocol Let's Encrypt speaks. To prove you control shop.lab, the CA gives cert-manager a token, and expects to fetch it from http://shop.lab/.well-known/acme-challenge/<token> on port 80:

1. cert-manager creates a solver pod (answers the token), a Service, and a
   temporary Ingress for /.well-known/acme-challenge/<token> on shop.lab
2. cert-manager checks it itself first (the "self check"), from its own pod,
   through real DNS and the real ingress controller
3. it tells the ACME server "ready"; the CA fetches the same URL
4. valid -> the CA signs -> the certificate lands in the Secret

That is why HTTP-01 needs three things that are not cert-manager's: public DNS pointing at your ingress, the ingress controller answering on port 80, and the solver Ingress being picked up by that controller (its class).

DNS-01 proves control by creating a TXT record _acme-challenge.shop.lab through your DNS provider's API instead. It works for private hosts and is the only way to get a wildcard (*.shop.lab). The price: cert-manager needs credentials for your DNS zone.

The lab's ACME server: Pebble

The lab cannot reach Let's Encrypt, so it runs Pebble, the small ACME test server Let's Encrypt publishes, in namespace pebble. Its directory URL:

https://pebble.pebble.svc.cluster.local:14000/dir

Pebble's own HTTPS certificate is signed by a throwaway CA nobody trusts, so the issuer must either carry that CA (caBundle) or skip verification (skipTLSVerify: true, test servers only). The certificates it issues come from "Pebble Intermediate CA" - your browser would not trust them either; for the lab that is fine, the mechanics are identical to Let's Encrypt's.

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: lab-acme
spec:
  acme:
    server: https://pebble.pebble.svc.cluster.local:14000/dir
    email: [email protected]
    skipTLSVerify: true                 # Pebble only. Never for a real CA.
    privateKeySecretRef:
      name: lab-acme-account            # cert-manager creates the ACME account key here
    solvers:
    - http01:
        ingress:
          ingressClassName: nginx       # the solver Ingress must reach THIS controller

A Ready issuer says so:

$ kubectl get clusterissuer
NAME       READY   AGE
lab-acme   True    40s

READY False with reason ErrRegisterACMEAccount means cert-manager could not even talk to the ACME server: wrong URL, DNS, or the TLS trust above.

The easy way: annotate the Ingress

The ingress-shim (part of cert-manager) watches Ingresses. One annotation and a tls section is enough:

metadata:
  annotations:
    cert-manager.io/cluster-issuer: lab-acme
spec:
  tls:
  - hosts: [shop.lab]
    secretName: shop-tls          # cert-manager creates and owns this Secret

It creates a Certificate named after the Secret (shop-tls) with the tls hosts as dnsNames. The same works on a Gateway (cert-manager.io/cluster-issuer on the Gateway, with --enable-gateway-api on cert-manager).

Watching an issuance

$ kubectl get certificate,certificaterequest,order,challenge -n shop
NAME                                   READY   SECRET     AGE
certificate.cert-manager.io/shop-tls   False   shop-tls   19s
NAME                                            APPROVED   DENIED   READY   ISSUER     REQUESTER                                         AGE
certificaterequest.cert-manager.io/shop-tls-1   True                False   lab-acme   system:serviceaccount:cert-manager:cert-manager   19s
NAME                                               STATE     AGE
order.acme.cert-manager.io/shop-tls-1-2574616885   pending   19s
NAME                                                              STATE     DOMAIN     AGE
challenge.acme.cert-manager.io/shop-tls-1-2574616885-2848982576   pending   shop.lab   19s

The conditions and messages you will read, top to bottom:

Certificate Ready False   DoesNotExist   Issuing certificate as Secret does not exist
CertificateRequest Ready False Pending   Waiting on certificate issuance from order shop/shop-tls-1-...: "pending"
Challenge reason          Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200'
Certificate Ready True    Ready          Certificate is up to date and has not expired

cmctl status certificate shop-tls -n shop prints all of it in one view: the Certificate's conditions, the Issuer, the Secret's certificate, and while issuing, the CertificateRequest, the Order and each Challenge with its reason. cmctl is a separate binary (v2.6, github.com/cert-manager/cmctl).

Renewal

A Certificate defaults to duration: 2160h (90 days) and renews at 2/3 of its life (renewBefore = 1/3 of the duration, or renewBeforePercentage). Each renewal is a new CertificateRequest revision (shop-tls-2), and since v1.18 a new private key every time (privateKey.rotationPolicy: Always). cmctl renew shop-tls forces one now - the usual reaction to "we rotated the CA" or "the key may have leaked".

What breaks renewals in real life is always outside cert-manager: someone changed the DNS, a NetworkPolicy now blocks the solver pod, the Ingress class was renamed in a migration. Alert on the certificate's expiry (cert-manager exports certmanager_certificate_expiration_timestamp_seconds), not on the renewal job.

In an interview: "Certificate -> CertificateRequest -> Order -> Challenge. When a certificate is stuck I go down that chain to the lowest object; for HTTP-01 the Challenge's reason says whether DNS, the ingress class or the self check failed."

What you can now do:

Why it helps

Expired certificates cause real, embarrassing outages, and short-lived certificates (Let's Encrypt's are 90 days, and getting shorter) make manual renewal impossible. cert-manager is the standard answer in Kubernetes, so you will meet it in nearly every cluster.

When issuance gets stuck, the error is never on the Certificate itself: it is in the lowest object of the chain, usually the Challenge, with a reason such as a 404 from the self check or a DNS name that does not resolve. Knowing to walk down the chain, and what each reason means, is what turns a stuck certificate from an afternoon into five minutes. cmctl status certificate shows the whole chain at once.

Commands in this lesson

kubectl cmctl

FAQ

Issuer or ClusterIssuer?

An Issuer lives in one namespace and can only issue certificates there; a ClusterIssuer is cluster-wide and any namespace can use it. Platform teams usually provide a ClusterIssuer for the company's ACME account, and an annotation such as cert-manager.io/cluster-issuer on an Ingress picks it.

When do I need DNS-01 instead of HTTP-01?

For wildcard certificates (*.shop.example), which HTTP-01 cannot issue, and for names that are not reachable from the internet on port 80. DNS-01 proves control by creating a TXT record, so cert-manager needs credentials for your DNS provider.

What is the self check?

Before telling the ACME server "ready", cert-manager fetches the challenge URL itself through the same path the server will use. If it gets a 404, a redirect to HTTPS or a DNS error, it waits and retries, and the Challenge's reason says what it saw, so a misrouted solver shows up there.

Why did cert-manager create an extra Ingress and pod?

For HTTP-01 it starts a small solver pod that serves the token and a temporary Ingress (or HTTPRoute) that sends /.well-known/acme-challenge/<token> for your host to it. They are removed when the challenge finishes, so seeing them hang around means a challenge is still pending.

How do I force a renewal to test it?

cmctl renew <certificate> -n <namespace> marks the Certificate for re-issuance now. A new CertificateRequest with the next revision number appears, and when it is done the Secret holds the new certificate. Testing renewal once, long before expiry, is much better than finding out it fails on the last day.

In an interview Mid

A cert-manager Certificate has been not Ready for an hour. How do you debug it?

I walk down the chain: Certificate, then the CertificateRequest it created, then the Order, then the Challenge, because the real error sits in the lowest object. cmctl status certificate <name> -n <ns> shows the whole chain. For an HTTP-01 Challenge the reason tells me what the self check saw: a DNS error means the name does not resolve to the ingress yet, a 404 means the solver Ingress is not picked up (often a wrong ingress class), a redirect means ssl-redirect is catching the challenge path. If the Issuer itself is not Ready, the ACME account could not register, for example because of the server URL or trust.

Also asked: What is the difference between HTTP-01 and DNS-01 challenges? · How does cert-manager decide when to renew a certificate? · What does the cert-manager.io/cluster-issuer annotation on an Ingress do?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.