OnCallReady

Kubernetes · 4 min read

CreateContainerConfigError: couldn't find key in ConfigMap (and other missing config)

New pods stuck in CreateContainerConfigError or ContainerCreating after a config cleanup? Read the Events, find the missing ConfigMap or key, fix the config.

terminal
$ kubectl get pods -n orders
NAME                          READY   STATUS                       RESTARTS   AGE
orders-api-bqjjjczwtp-4rgql   0/1     CreateContainerConfigError   0          10m
orders-api-bqjjjczwtp-rlfg7   1/1     Running                      0          120m
orders-api-bqjjjczwtp-vt8m4   1/1     Running                      0          120m

Someone "tidied up" the config in namespace orders this morning. Two old pods still serve traffic. The third was moved by a node drain, and its replacement won't start. kubectl logs has nothing for you: no container was ever created, so nothing has logged anything. That's the difference from CrashLoopBackOff, where the container starts and then dies.

What is actually happening

Before the kubelet starts a container, it builds that container's configuration: environment variables from env and envFrom, files from the pod's volumes. A pod spec that says

output
env:
- name: DB_URL
  valueFrom:
    configMapKeyRef:
      name: orders-env
      key: DB_URL

is a promise that ConfigMap orders-env exists and has a key DB_URL. If it doesn't at the moment the container starts, the kubelet can't build the environment, so it doesn't create the container. The container stays Waiting with reason CreateContainerConfigError, and the kubelet keeps retrying on its own.

Why are the old pods fine? Environment variables are resolved once, when the container starts. Renaming a key or deleting a ConfigMap changes nothing inside a running container. That makes config mistakes latent: the cleanup breaks nothing until the next pod start, which might be a drain, a scale-up, a crash or the next deploy. That's also why they get through review.

A missing volume source fails at an earlier stage, and it looks different. The kubelet can't mount the volume, so the pod sits in ContainerCreating and the events say FailedMount:

What is missingPod STATUSEvent message
a ConfigMap mounted as a volumeContainerCreatingMountVolume.SetUp failed for volume "files" : configmap "orders-files" not found
a ConfigMap used for envCreateContainerConfigErrorError: configmap "orders-env" not found
one key in itCreateContainerConfigErrorError: couldn't find key DB_URL in ConfigMap orders/orders-env
a Secret or a Secret keythe same two shapessecret "x" not found, couldn't find key K in Secret ns/x

Diagnosis

1. Look at the pod that is trying to start

The Running pods tell you nothing about config. They resolved theirs hours ago. Pick the one that isn't Running. If it's Pending with no node at all, the problem comes before the kubelet: the scheduler couldn't place it (see how to read FailedScheduling).

2. Read its Events, from the bottom

terminal
$ kubectl describe pod -n orders orders-api-bqjjjczwtp-4rgql
...
    State:          Waiting
      Reason:       ContainerCreating
...
    Environment:
      DB_URL:  <set to the key 'DB_URL' of config map 'orders-env'>  Optional: false
      POOL_SIZE:  <set to the key 'POOL_SIZE' of config map 'orders-env'>  Optional: false
...
Events:
  Type     Reason       Age                From               Message
  ----     ------       ----               ----               -------
  Normal   Scheduled    10m                default-scheduler  Successfully assigned orders/orders-api-bqjjjczwtp-4rgql to worker-1
  Warning  FailedMount  1s (x76 over 10m)  kubelet            MountVolume.SetUp failed for volume "files" : configmap "orders-files" not found

The Environment block shows what the pod expects. Optional: false means "don't start without it". The last event shows what is wrong right now: the volume's ConfigMap is gone.

3. Compare with what exists

terminal
$ kubectl get cm -n orders
NAME               DATA   AGE
kube-root-ca.crt   1      1s
orders-env         2      120m

orders-files was deleted. Recreate it and look again. Errors arrive one at a time: volumes must mount before the kubelet builds the environment, so fixing the first problem shows you the second:

output
    State:          Waiting
      Reason:       CreateContainerConfigError
      Message:      couldn't find key DB_URL in ConfigMap orders/orders-env
...
  Warning  Failed       6s (x2 over 14s)    kubelet            Error: couldn't find key DB_URL in ConfigMap orders/orders-env
terminal
$ kubectl get cm orders-env -n orders -o yaml
apiVersion: v1
data:
  DATABASE_URL: jdbc:postgresql://pg.orders.svc:5432/orders
  POOL_SIZE: "20"

The tidy-up renamed DB_URL to DATABASE_URL.

The fix

Fix the config, not the Deployment. Editing the Deployment would roll every pod, the healthy ones too, and here the Deployment belongs to a pipeline anyway:

terminal
$ kubectl create configmap orders-files -n orders --from-literal=application.yaml='server: {port: 8080}'
configmap/orders-files created
$ kubectl patch cm orders-env -n orders -p '{"data":{"DB_URL":"jdbc:postgresql://pg.orders.svc:5432/orders"}}'
configmap/orders-env patched

A merge patch adds DB_URL and keeps DATABASE_URL, in case whoever renamed it needs the new name. Then wait. The kubelet retries by itself:

terminal
$ kubectl get pods -n orders -w
NAME                          READY   STATUS                       RESTARTS   AGE
orders-api-bqjjjczwtp-4rgql   0/1     CreateContainerConfigError   0          10m
orders-api-bqjjjczwtp-rlfg7   1/1     Running                      0          120m
orders-api-bqjjjczwtp-vt8m4   1/1     Running                      0          120m
orders-api-bqjjjczwtp-4rgql   1/1     Running   0          10m

Nothing needed a restart. The fix worked because the pod was still trying.

How to prevent it

  • Versioned, immutable ConfigMaps. Give each version a new name (kustomize's configMapGenerator appends a content hash) and set immutable: true. Then nobody can "tidy" a ConfigMap in place. A rename becomes a new object plus a Deployment change, reviewed together. A rolling update starts new pods first, so a broken reference stalls the rollout while the old pods keep serving (as long as the strategy really keeps them, see downtime on every deploy).
  • Ship config and workload from one pipeline, so a key rename and the code that reads it change in the same commit.
  • optional: true only when the app really copes without the value. Otherwise it just moves the failure from the kubelet into your app at runtime.
  • After any config change, start a pod (kubectl rollout restart in a test environment). Running pods prove nothing about config.

Practise it

Chapter 15 has this incident, "CreateContainerConfigError after a harmless cleanup", on a live cluster in your browser. Its lessons explain env vs volume ConfigMaps and what happens when you change one.

OnCallReady is free, with no ads and no tracking. RSS · All posts