OnCallReady

Lesson 15.9 · Kubernetes: Architecture & Workloads · 18 min read

The reconciliation loop

In plain words

Think of a thermostat. You don't tell it "heat for 20 minutes"; you tell it "I want 21 degrees". It keeps checking: too cold, turn the heating on; warm enough, turn it off. If someone opens a window, it doesn't need to be told; the next check notices and reacts. And if you fight it by opening the window further, it just keeps heating.

Every Kubernetes controller works like that: read the desired state, read the actual state, take one step to close the gap, repeat forever. It compares levels ("want 3, see 2") rather than reacting to events, so missed events don't matter. Objects are chained with ownerReferences (Deployment to ReplicaSet to Pod), the garbage collector deletes dependents when owners go, and controllers adopt any matching pod without an owner.

Why this is the most important idea

Sooner or later you will change something in a cluster - delete a pod, edit a label, scale something by hand - and watch it quietly "undo itself". That is not a bug. Kubernetes is built out of loops that keep putting things back the way they were declared. Once you see the loops, you stop fighting them and start using them.

What you need to know already: controllers and the controller-manager (15.5), the kubelet (15.7), the kind table in 15.3 (Deployment, ReplicaSet), Terraform's plan/apply against drift (12.24, 13.21), jsonpath as seen in 15.5.

Spec and status

Almost every object has two halves:

You met this pair in Terraform (Ch 12): your .tf files are the desired state, the real cloud is the observed state, and terraform plan shows the difference. The big change: in Terraform you run plan/apply when you choose; in Kubernetes it runs for ever, on its own.

The loop

Every controller runs the same loop:

loop:
    desired  = read the spec from the apiserver       (what you asked for)
    observed = read the actual state from the apiserver (what exists)
    if they differ:
        take ONE step to make observed closer to desired
    repeat

This is the reconciliation loop ("reconcile" = make two things agree). Kubernetes is not a deployment tool that runs a script once. It is a set of these loops, each owning one small piece:

Three consequences you have to take to heart:

1. You declare, you do not command. kubectl scale --replicas=5 deploy/web does not start two pods. It changes one number (spec.replicas) in etcd. The ReplicaSet controller notices the difference and creates two Pod objects. The scheduler notices two pods with no node. Two kubelets notice pods assigned to them. Each step is a different loop reacting to the previous one's output.

2. It is level-triggered, not edge-triggered. An edge is an event ("a pod was deleted"); a level is a state ("I want 3, I see 2"). Controllers compare levels. So a missed event, a controller restart, a network blip - none of it matters: the next pass sees the difference and fixes it. That is where the "self-healing" comes from. (systemd's Restart= from 2.10 reacts to one event, a process exiting; this checks the whole picture again and again.)

3. Fighting the loop always loses. Delete a pod of a Deployment and it comes back. Scale a ReplicaSet owned by a Deployment by hand and the Deployment scales it back. If something keeps "undoing" your change, look for the controller that owns the thing you changed.

Watching it happen

First create something to watch. k create deployment web --image=nginx:1.27 --replicas=3 makes a Deployment of 3 nginx pods; k get pods -l app=web lists only pods carrying the label app=web (the generator adds that label, 15.26 explains labels):

$ k create deployment web --image=nginx:1.27 --replicas=3     # AlreadyExists is fine
deployment.apps/web created
$ k get pods -l app=web
NAME                   READY   STATUS    RESTARTS   AGE
web-7d9f5cbb6d-8xk2p   1/1     Running   0          3m
web-7d9f5cbb6d-n4v7q   1/1     Running   0          3m
web-7d9f5cbb6d-zq9hm   1/1     Running   0          3m

Now delete one pod and watch what happens. Normally the watch runs in a second terminal; with one terminal, delete without waiting and start the watch straight after (Ctrl+C to stop it):

$ k delete $(k get pods -l app=web -o name | head -1) --wait=false; k get pods -l app=web -w
pod "web-7d9f5cbb6d-8xk2p" deleted from default namespace
NAME                   READY   STATUS    RESTARTS   AGE
web-7d9f5cbb6d-8xk2p   1/1     Running   0          3m
web-7d9f5cbb6d-n4v7q   1/1     Running   0          3m
web-7d9f5cbb6d-zq9hm   1/1     Running   0          3m
web-7d9f5cbb6d-8xk2p   1/1     Terminating   0          3m
web-7d9f5cbb6d-q2l8c   0/1     Pending       0          0s
web-7d9f5cbb6d-q2l8c   0/1     ContainerCreating   0          0s
web-7d9f5cbb6d-8xk2p   0/1     Terminating         0          3m
web-7d9f5cbb6d-q2l8c   1/1     Running             0          2s

Piece by piece:

Now read the lines. The old pod goes Terminating (its graceful shutdown started - SIGTERM, then a wait, 3.18). The ReplicaSet controller immediately creates a replacement, because it only counts pods that are not terminating. The new pod is Pending (no node yet), then ContainerCreating (the scheduler picked a node, the kubelet is starting it), then Running.

The replacement has a new name. Pods are cattle, not pets: nobody repairs a pod, the loop replaces it.

Who owns what: ownerReferences

Every object a controller creates carries a pointer back to its creator, in metadata.ownerReferences. The jsonpath below prints kind/name of the first owner of the first pod (.items[0] = the first object in the list):

$ k get pods -l app=web -o jsonpath='{.items[0].metadata.ownerReferences[0].kind}/{.items[0].metadata.ownerReferences[0].name}{"\n"}'
ReplicaSet/web-7d9f5cbb6d
$ k get rs -l app=web -o jsonpath='{.items[0].metadata.ownerReferences[0].kind}/{.items[0].metadata.ownerReferences[0].name}{"\n"}'
Deployment/web

The pod is owned by a ReplicaSet; the ReplicaSet is owned by the Deployment. In the YAML it looks like:

metadata:
  ownerReferences:
  - apiVersion: apps/v1
    blockOwnerDeletion: true
    controller: true
    kind: ReplicaSet
    name: web-7d9f5cbb6d
    uid: 1f5a0c3e-7b0b-4f1a-9f5e-2b7c0d8e4a11

controller: true means "this owner manages me". The uid is the owner's unique id (every object gets one when it is created). k describe shows the same as Controlled By: ReplicaSet/web-7d9f5cbb6d. The chain for a Deployment is always:

Deployment/web  ->  ReplicaSet/web-<template-hash>  ->  Pod/web-<template-hash>-<random>

The <template-hash> is a hash (a short fingerprint) of the pod description inside the Deployment. Change that description - a new image, say - and you get a new hash, so a new ReplicaSet. That is how updates work (15.16).

Garbage collection

The garbage collector (one of the controllers, 15.5) deletes objects whose owners are all gone. So deleting a Deployment deletes its ReplicaSets, which deletes their pods:

$ k delete deploy web
deployment.apps "web" deleted from default namespace
$ k get rs,pods -l app=web
No resources found in default namespace.

That is background cascading deletion (the default): the owner goes first, the garbage collector removes the dependents shortly after. --cascade picks another way:

k delete deploy web --cascade=foreground   # dependents first, then the owner
k delete deploy web --cascade=orphan       # delete ONLY the owner, leave the rest

--cascade=orphan is useful and dangerous: the ReplicaSet and its pods keep running with no owner - orphans. Recreate the Deployment with the same label rules and it adopts them (sets itself as their owner again) instead of creating new pods. That is the standard way to change a field that cannot be edited in place without restarting anything.

Adoption works both ways

A ReplicaSet owns every pod carrying its label that has no other controller. Create a bare pod with the label app=web next to a 3-replica ReplicaSet and the controller adopts it - and now it sees 4, so it deletes one. Usually yours, because the newest pods are deleted first. Labels are not decoration; they are the wiring (15.26).

What you can now do:

Why it helps

This is the idea that explains "Kubernetes keeps undoing my change". Situations: you delete a pod to "stop" an app and it comes straight back; you scale a ReplicaSet and the Deployment scales it back; you run a test pod with app=web next to a ReplicaSet and one of the real pods disappears because the controller adopted yours and deleted the surplus. Knowing to look for the owning controller, via ownerReferences or Controlled By in describe, turns these into one-minute answers. --cascade=orphan is also the trick for changing an immutable selector without downtime, and "explain level-triggered reconciliation" is a common senior interview question.

FAQ

Why does my deleted pod keep coming back?

It's owned by a ReplicaSet (usually through a Deployment), whose loop sees fewer pods than desired and creates a replacement, with a new name. To really stop it, change the owner: scale the Deployment to 0 or delete it. kubectl describe pod shows Controlled By:, and ownerReferences in the YAML shows the chain.

What does level-triggered mean?

Controllers compare the current level of state with the desired level ("I want 3, I see 2") on every pass, instead of reacting to individual events ("a pod was deleted"). So a missed event, a controller restart or a network blip is harmless: the next pass sees the difference. That's the source of Kubernetes' self-healing.

What happens to the pods when I delete a Deployment?

By default, background cascading deletion: the Deployment goes, then the garbage collector deletes its ReplicaSets, which deletes their pods. --cascade=foreground deletes dependents first; --cascade=orphan deletes only the Deployment and leaves ReplicaSets and pods running without an owner, which a new Deployment with the same selector can adopt.

What is adoption and why can it delete my pod?

A ReplicaSet owns every pod that matches its selector and has no other controller. If you create a bare pod with matching labels, the ReplicaSet adopts it, now sees one too many, and deletes one, often the newest, possibly yours. Deployments protect their ReplicaSets with a pod-template-hash label in the selector, but bare ReplicaSets and DaemonSets don't.

What is the pod-template-hash?

A hash of the Deployment's pod template, used in the ReplicaSet's name and as a label on its pods and in its selector. Changing the template produces a new hash, hence a new ReplicaSet, which is how a rollout creates a new version alongside the old one. It also stops ReplicaSets of different versions from adopting each other's pods.

In an interview Junior

What is a controller in Kubernetes, and what is the reconciliation loop?

Every object has a spec (what you want, which you write) and a status (what is, which the cluster writes). A controller is a loop that watches one kind of object and keeps acting until status matches spec:

loop: desired = spec; actual = observe; if different, act; repeat

The ReplicaSet controller keeps N pods with its labels, the Deployment controller keeps the ReplicaSets matching the template, the scheduler gives pods a node, the kubelet keeps containers running.

Consequences:

Also asked: What are ownerReferences used for? · What does kubectl delete --cascade=orphan do? · Why does a pod you delete come back with a different name?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.