Why this is the most important idea
Sooner or later you will change something in a cluster - delete a pod, edit a label, scale something by hand - and watch it quietly "undo itself". That is not a bug. Kubernetes is built out of loops that keep putting things back the way they were declared. Once you see the loops, you stop fighting them and start using them.
What you need to know already: controllers and the controller-manager (15.5), the kubelet (15.7), the kind table in 15.3 (Deployment, ReplicaSet), Terraform's plan/apply against drift (12.24, 13.21), jsonpath as seen in 15.5.
Spec and status
Almost every object has two halves:
- spec - what you want (the desired state): "3 copies of nginx:1.27". You write it.
- status - what is (the observed state): "3 exist, 3 are ready". The cluster writes it.
You met this pair in Terraform (Ch 12): your .tf files are the desired state, the real cloud is the observed state, and terraform plan shows the difference. The big change: in Terraform you run plan/apply when you choose; in Kubernetes it runs for ever, on its own.
The loop
Every controller runs the same loop:
loop:
desired = read the spec from the apiserver (what you asked for)
observed = read the actual state from the apiserver (what exists)
if they differ:
take ONE step to make observed closer to desired
repeat
This is the reconciliation loop ("reconcile" = make two things agree). Kubernetes is not a deployment tool that runs a script once. It is a set of these loops, each owning one small piece:
- the ReplicaSet controller owns "there are N pods with this label"
- the Deployment controller owns "the ReplicaSets match the Deployment"
- the scheduler owns "every pod has a node"
- the kubelet owns "the containers of pods on my node are running"
Three consequences you have to take to heart:
1. You declare, you do not command. kubectl scale --replicas=5 deploy/web does not start two pods. It changes one number (spec.replicas) in etcd. The ReplicaSet controller notices the difference and creates two Pod objects. The scheduler notices two pods with no node. Two kubelets notice pods assigned to them. Each step is a different loop reacting to the previous one's output.
2. It is level-triggered, not edge-triggered. An edge is an event ("a pod was deleted"); a level is a state ("I want 3, I see 2"). Controllers compare levels. So a missed event, a controller restart, a network blip - none of it matters: the next pass sees the difference and fixes it. That is where the "self-healing" comes from. (systemd's Restart= from 2.10 reacts to one event, a process exiting; this checks the whole picture again and again.)
3. Fighting the loop always loses. Delete a pod of a Deployment and it comes back. Scale a ReplicaSet owned by a Deployment by hand and the Deployment scales it back. If something keeps "undoing" your change, look for the controller that owns the thing you changed.
Watching it happen
First create something to watch. k create deployment web --image=nginx:1.27 --replicas=3 makes a Deployment of 3 nginx pods; k get pods -l app=web lists only pods carrying the label app=web (the generator adds that label, 15.26 explains labels):
$ k create deployment web --image=nginx:1.27 --replicas=3 # AlreadyExists is fine
deployment.apps/web created
$ k get pods -l app=web
NAME READY STATUS RESTARTS AGE
web-7d9f5cbb6d-8xk2p 1/1 Running 0 3m
web-7d9f5cbb6d-n4v7q 1/1 Running 0 3m
web-7d9f5cbb6d-zq9hm 1/1 Running 0 3m
Now delete one pod and watch what happens. Normally the watch runs in a second terminal; with one terminal, delete without waiting and start the watch straight after (Ctrl+C to stop it):
$ k delete $(k get pods -l app=web -o name | head -1) --wait=false; k get pods -l app=web -w
pod "web-7d9f5cbb6d-8xk2p" deleted from default namespace
NAME READY STATUS RESTARTS AGE
web-7d9f5cbb6d-8xk2p 1/1 Running 0 3m
web-7d9f5cbb6d-n4v7q 1/1 Running 0 3m
web-7d9f5cbb6d-zq9hm 1/1 Running 0 3m
web-7d9f5cbb6d-8xk2p 1/1 Terminating 0 3m
web-7d9f5cbb6d-q2l8c 0/1 Pending 0 0s
web-7d9f5cbb6d-q2l8c 0/1 ContainerCreating 0 0s
web-7d9f5cbb6d-8xk2p 0/1 Terminating 0 3m
web-7d9f5cbb6d-q2l8c 1/1 Running 0 2s
Piece by piece:
k get pods -l app=web -o name | head -1- print the pods aspod/<name>and keep the first;$( )puts that into thek deletecommand.--wait=false- return immediately instead of waiting for the pod to be gone.-w- watch: print the list, then one new line each time a pod changes.
Now read the lines. The old pod goes Terminating (its graceful shutdown started - SIGTERM, then a wait, 3.18). The ReplicaSet controller immediately creates a replacement, because it only counts pods that are not terminating. The new pod is Pending (no node yet), then ContainerCreating (the scheduler picked a node, the kubelet is starting it), then Running.
The replacement has a new name. Pods are cattle, not pets: nobody repairs a pod, the loop replaces it.
Who owns what: ownerReferences
Every object a controller creates carries a pointer back to its creator, in metadata.ownerReferences. The jsonpath below prints kind/name of the first owner of the first pod (.items[0] = the first object in the list):
$ k get pods -l app=web -o jsonpath='{.items[0].metadata.ownerReferences[0].kind}/{.items[0].metadata.ownerReferences[0].name}{"\n"}'
ReplicaSet/web-7d9f5cbb6d
$ k get rs -l app=web -o jsonpath='{.items[0].metadata.ownerReferences[0].kind}/{.items[0].metadata.ownerReferences[0].name}{"\n"}'
Deployment/web
The pod is owned by a ReplicaSet; the ReplicaSet is owned by the Deployment. In the YAML it looks like:
metadata:
ownerReferences:
- apiVersion: apps/v1
blockOwnerDeletion: true
controller: true
kind: ReplicaSet
name: web-7d9f5cbb6d
uid: 1f5a0c3e-7b0b-4f1a-9f5e-2b7c0d8e4a11
controller: true means "this owner manages me". The uid is the owner's unique id (every object gets one when it is created). k describe shows the same as Controlled By: ReplicaSet/web-7d9f5cbb6d. The chain for a Deployment is always:
Deployment/web -> ReplicaSet/web-<template-hash> -> Pod/web-<template-hash>-<random>
The <template-hash> is a hash (a short fingerprint) of the pod description inside the Deployment. Change that description - a new image, say - and you get a new hash, so a new ReplicaSet. That is how updates work (15.16).
Garbage collection
The garbage collector (one of the controllers, 15.5) deletes objects whose owners are all gone. So deleting a Deployment deletes its ReplicaSets, which deletes their pods:
$ k delete deploy web
deployment.apps "web" deleted from default namespace
$ k get rs,pods -l app=web
No resources found in default namespace.
That is background cascading deletion (the default): the owner goes first, the garbage collector removes the dependents shortly after. --cascade picks another way:
k delete deploy web --cascade=foreground # dependents first, then the owner
k delete deploy web --cascade=orphan # delete ONLY the owner, leave the rest
--cascade=orphan is useful and dangerous: the ReplicaSet and its pods keep running with no owner - orphans. Recreate the Deployment with the same label rules and it adopts them (sets itself as their owner again) instead of creating new pods. That is the standard way to change a field that cannot be edited in place without restarting anything.
Adoption works both ways
A ReplicaSet owns every pod carrying its label that has no other controller. Create a bare pod with the label app=web next to a 3-replica ReplicaSet and the controller adopts it - and now it sees 4, so it deletes one. Usually yours, because the newest pods are deleted first. Labels are not decoration; they are the wiring (15.26).
What you can now do:
- explain spec vs status and the reconciliation loop
- predict what a controller does when you delete or change something it owns
- follow
ownerReferencesfrom a pod up to its Deployment, and choose how a delete cascades