OnCallReady

KubernetesNetworking · 5 min read

cert-manager certificate stuck: Ready False, a pending Order and an HTTP-01 challenge that gets 404

The Certificate never becomes Ready and the issuer says it is fine. Follow CertificateRequest, Order and Challenge down to the HTTP-01 self-check that fails.

pay.lab has served a self-signed certificate since the cluster migration last week. cert-manager is running, the issuer is Ready, DNS is right, and the team has already deleted and recreated the Certificate twice:

terminal
$ kubectl get certificate -n pay-lab
NAME      READY   SECRET    AGE
pay-tls   False   pay-tls   15m
$ kubectl get clusterissuer
NAME            READY   AGE
migrated-acme   True    15m

Recreating the Certificate does not help, because the Certificate is not where it is stuck.

What is happening: four objects, and the last one is waiting

For an ACME issuer (Let's Encrypt, or the lab's Pebble test server), cert-manager turns one Certificate into a chain of objects, each created by the one before:

text
Certificate          "I want a cert for pay.lab in Secret pay-tls"
 └ CertificateRequest   the CSR, approved, sent to the issuer
    └ Order               the ACME order at the CA
       └ Challenge          one per domain: prove you control pay.lab

With HTTP-01, proving control means serving a token at http://pay.lab/.well-known/acme-challenge/<token>. cert-manager starts a small solver pod, plus a Service and an Ingress (or an HTTPRoute) that routes that path to it. Before it tells the CA to check, it runs a self-check: it requests the URL itself, the way the CA will. Until that returns the token with a 200, the Challenge stays pending, and so does everything above it.

An issuer that is Ready only means the ACME account is registered. It says nothing about whether challenges can be solved.

Diagnosis: walk down the chain

1. See all the objects at once

terminal
$ kubectl get certificaterequest,order,challenge -n pay-lab
NAME                                           APPROVED   DENIED   READY   ISSUER          REQUESTER                                         AGE
certificaterequest.cert-manager.io/pay-tls-1   True                False   migrated-acme   system:serviceaccount:cert-manager:cert-manager   15m

NAME                                              STATE     AGE
order.acme.cert-manager.io/pay-tls-1-2471426364   pending   15m

NAME                                                             STATE     DOMAIN    AGE
challenge.acme.cert-manager.io/pay-tls-1-2471426364-1765097894   pending   pay.lab   15m

Approved, ordered, and a Challenge pending for 15 minutes. cmctl status certificate prints the whole chain in one go, and its last line is the one that matters:

terminal
$ cmctl status certificate pay-tls -n pay-lab | tail -3
    URL: https://pebble.pebble.svc.cluster.local:14000/authZ/NjkzNTQxYTYzZDhmNWNlZD, Identifier: pay.lab, Initial State: pending, Wildcard: false
Challenges:
- Name: pay-tls-1-2471426364-1765097894, Type: HTTP-01, Token: NjkzNTQxYTYzZDhmNWNlZDY0NTQ1MzFmMzlmYWJjNjA, Key: NjkzNTQxYTYzZDhmNWNlZDY0NTQ1MzFmMzlmYWJjNjA.ZGIyYjBiNjcyNTUwMTNkZmZjMTc1NzVhZWE3ODhjZjN, State: pending, Reason: Waiting for HTTP-01 challenge propagation: wrong status code '404', expected '200', Processing: true, Presented: true

Presented: true: cert-manager created the solver. wrong status code '404': the self-check cannot reach it.

2. Repeat the self-check yourself

terminal
$ TOKEN=$(kubectl get challenge -n pay-lab -o jsonpath='{.items[0].spec.token}'); curl -s -o /dev/null -w '%{http_code}\n' -H 'Host: pay.lab' http://10.64.0.240/.well-known/acme-challenge/$TOKEN
404

The ingress controller answers, but nothing routes the challenge path. Which Ingress should?

terminal
$ kubectl get ingress -n pay-lab
NAME                        CLASS          HOSTS     ADDRESS       PORTS     AGE
cm-acme-http-solver-dcggg   nginx-public   pay.lab                 80        15m
pay                         nginx          pay.lab   10.64.0.240   80, 443   15m

The solver's Ingress has class nginx-public and no address. Nobody picked it up.

terminal
$ kubectl get ingressclass
NAME    CONTROLLER             PARAMETERS   AGE
nginx   k8s.io/ingress-nginx   <none>       2m4s

nginx-public was renamed away in the migration. The controller's log says so too:

terminal
$ kubectl logs -n ingress-nginx deploy/ingress-nginx-controller | grep cm-acme-http-solver | tail -1
I0922 19:45:05.000000       7 store.go:369] "Ignoring ingress because of error while validating ingress class" ingress="pay-lab/cm-acme-http-solver-dcggg" error="error getting IngressClass: ingressclass.networking.k8s.io "nginx-public" not found"

3. Find where the class comes from

The solver Ingress is built from the issuer's solver config, not from the app's Ingress. The Challenge records the one it used:

terminal
$ kubectl describe challenge -n pay-lab | sed -n '/^Spec:/,$p'
Spec:
  ...
  Solver:
    Http01:
      Ingress:
        IngressClassName: nginx-public
  ...
Status:
  Presented: true
  Processing: true
  Reason: 'Waiting for HTTP-01 challenge propagation: wrong status code ''404'', expected ''200'''
  State: pending

And the ClusterIssuer still says it:

terminal
$ kubectl get clusterissuer migrated-acme -o yaml | sed -n '/solvers:/,/^status/p'
    solvers:
    - http01:
        ingress:
          ingressClassName: nginx-public
status:

The fix

Point the solver at the class that exists. Leave the app's Ingress alone, because its class was never the problem:

terminal
$ kubectl patch clusterissuer migrated-acme --type=json -p '[{"op":"replace","path":"/spec/acme/solvers/0/http01/ingress/ingressClassName","value":"nginx"}]'
clusterissuer.cert-manager.io/migrated-acme patched

The existing Challenge keeps the solver it was created with, so delete it. The Order creates a new one from the fixed issuer:

terminal
$ kubectl delete challenge --all -n pay-lab
challenge.acme.cert-manager.io "pay-tls-1-2471426364-1765097894" deleted from pay-lab namespace
$ kubectl wait --for=condition=Ready certificate/pay-tls -n pay-lab --timeout=120s
certificate.cert-manager.io/pay-tls condition met
$ kubectl get secret pay-tls -n pay-lab -o jsonpath='{.data.tls\.crt}' | base64 -d | openssl x509 -noout -issuer -enddate
issuer=CN=Pebble Intermediate CA 645fc5
notAfter=Dec 21 20:00:19 2026 GMT

Issued by the CA, valid for 90 days, and the temporary solver Ingress is cleaned up.

Other reasons a challenge stays pending

The self-check message names the failure. The usual ones:

  • wrong status code '404': the solver is not routed. A wrong or missing ingress class (as here), a second ingress controller grabbing the host, or an app Ingress whose rules win over the solver's path.
  • connection refused / timeout: port 80 is not reachable from inside the cluster at the domain's address. Hairpin NAT, a firewall, or a load balancer that only listens on 443. The CA must reach port 80 from the internet as well.
  • no such host: DNS does not point at the ingress yet. external-dns not running, or a TTL still serving the old address.
  • The Order says invalid: the CA checked and failed (look at kubectl describe order). Deleting the Certificate's failed request (or cmctl renew pay-tls) starts a new order once the cause is fixed. Mind Let's Encrypt's rate limits while you experiment, and test against the staging endpoint.
  • On Gateway API, use the gatewayHTTPRoute solver with parentRefs to your Gateway. Then the route must be accepted by the Gateway like any other route, and a solver route that is not accepted gives the same 404.

ingress-nginx itself was retired in March 2026 (no more releases or security fixes). Moving to a Gateway API implementation is the long-term fix, and the solver config moves with it.

Keeping it from coming back

  • Alert on the Certificate, not on the browser. cert-manager exports certmanager_certificate_ready_status and certmanager_certificate_expiration_timestamp_seconds. Page when a certificate is not Ready for an hour, or expires in less than 14 days. An expired certificate is a full outage, as with kubeadm's cluster certificates.
  • Treat issuer config as part of every rename. An ingress class, a Gateway name or a namespace change touches the solvers too: grep the issuers in the same pull request.
  • Test the chain after migrations with a throwaway Certificate on a test host, before the real ones come up for renewal 30 days before expiry. When TLS works and the page still fails, it is usually the routing behind it: OpenShift's 503 and a Service with no endpoints are the usual suspects.

Practise it

In the chapter Kubernetes: Ingress, Gateway API & Service Mesh, Incident: "the certificate never issues" is this migration: Pebble as the ACME server, the renamed class, and a certificate that must end up issued without touching the app's Ingress class. The mission A challenge stuck on DNS and the drill Why is this certificate not Ready? cover the other self-check failures.

OnCallReady is free, with no ads and no tracking. RSS · All posts