Why StatefulSets
Deployments treat pods as interchangeable: random names, created and deleted in any order. Try running a 3-member database cluster that way and it breaks: members find each other by name, each keeps its own data on disk, and the first one must start before the others join it. A StatefulSet is the workload for that: pods with fixed names, a fixed start order and their own disks.
What you need to know already: Deployments and rolling updates (15.16), pods and their DNS settings (15.14), etcd's members and quorum (15.5), DNS A records and nslookup (8.18, 8.22), Docker volumes (11.22).
What a Deployment cannot give you
A Deployment's pods are interchangeable - perfect for stateless web servers and wrong for a database cluster, where each member has:
- an identity other members know it by ("I am member 2 of the group")
- its own data that must follow it, not be shared
- an order: the primary starts first, the copies join it
A StatefulSet provides exactly those three. It comes with a companion object, a Service - a stable name in front of pods (Ch 16 teaches Services). Here it is a special headless one, explained below:
apiVersion: v1
kind: Service
metadata:
name: db
spec:
clusterIP: None # headless: no virtual IP, DNS returns pod IPs
selector:
app: db
ports:
- port: 6379
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: db
spec:
serviceName: db # the headless Service that gives pods their DNS names
replicas: 3
selector:
matchLabels:
app: db
template:
metadata:
labels:
app: db
spec:
containers:
- name: redis
image: redis:7.4
(--- separates two objects in one YAML file. Redis is a small in-memory database; port 6379 is its port.) The StatefulSet looks like a Deployment - replicas, selector, template - plus serviceName.
Stable names, ordered start
# after applying the db manifest above (the next mission does it with `cache`)
k get pods -w
NAME READY STATUS RESTARTS AGE
db-0 0/1 ContainerCreating 0 0s
db-0 1/1 Running 0 2s
db-1 0/1 ContainerCreating 0 0s
db-1 1/1 Running 0 2s
db-2 0/1 ContainerCreating 0 0s
db-2 1/1 Running 0 2s
Names are <statefulset>-<ordinal>: an ordinal is the member number, 0 to replicas-1, and it never changes. With the default podManagementPolicy: OrderedReady, db-1 is not created until db-0 is Running and Ready, db-2 not until db-1 is. Scaling down removes the highest ordinal first, one at a time.
Delete a member and it comes back with the same name (and, with storage, the same disk):
k delete pod db-1
pod "db-1" deleted from default namespace
k get pods -l app=db
NAME READY STATUS RESTARTS AGE
db-0 1/1 Running 0 5m
db-1 1/1 Running 0 3s
db-2 1/1 Running 0 5m
Same name, new pod (AGE 3s). Each pod also gets labels naming itself, which lets you pick out one member:
statefulset.kubernetes.io/pod-name=db-1
apps.kubernetes.io/pod-index=1
podManagementPolicy: Parallel drops the ordering (all pods at once) while keeping the stable names - use it when members do not depend on each other's start order.
The headless Service and per-pod DNS
A normal Service gets one virtual IP and spreads traffic over its pods. A headless Service (clusterIP: None) has no IP of its own; its DNS name resolves to the pod IPs directly, and - the reason StatefulSets need one - each pod gets its own DNS name served by CoreDNS:
<pod>.<service>.<namespace>.svc.cluster.local
db-0.db.default.svc.cluster.local
Test it from a throwaway pod. k run -it --rm dns --image=busybox:1.36 --restart=Never -- nslookup db-0.db: run a pod dns interactively (-it), delete it when it ends (--rm), do not restart it (--restart=Never), and run nslookup db-0.db in it:
k run -it --rm dns --image=busybox:1.36 --restart=Never -- nslookup db-0.db
Server: 10.96.0.10
Address: 10.96.0.10:53
Name: db-0.db.default.svc.cluster.local
Address: 10.244.1.87
pod "dns" deleted from default namespace
Server is CoreDNS; the short name db-0.db was expanded with the search domains (8.27) to the full name, which points at db-0's pod IP.
That is how a Redis copy finds its primary (db-0.db), how an etcd member lists its peers. The IP changes when the pod is recreated; the name does not. serviceName must match an existing headless Service for this to work.
Later (Ch 16): Services in full - ClusterIP, headless, and how DNS names are built.
Stable storage: volumeClaimTemplates
spec:
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: [ReadWriteOnce]
resources:
requests:
storage: 10Gi
A PersistentVolumeClaim (PVC) is a pod's request for a disk: "10Gi, usable by one node at a time" (ReadWriteOnce). It is Kubernetes' version of a named Docker volume (11.22) that outlives the container. volumeClaimTemplates makes the controller create one PVC per pod - data-db-0, data-db-1, data-db-2 - and always mount data-db-N into db-N, even after it moves to another node.
The PVCs survive deleting the StatefulSet (the default persistentVolumeClaimRetentionPolicy: Retain), so data outlives the workload. This cluster has no disk provider set up yet, so such PVCs would stay Pending, and their pods with them - which is why the labs here run without storage.
Updates
updateStrategy: RollingUpdate (the default) replaces pods from the highest ordinal down, one at a time, each only after the previous is Ready. rollingUpdate.partition: N updates only ordinals >= N - a way to try a new version on one member first (a canary):
k patch sts db -p '{"spec":{"updateStrategy":{"rollingUpdate":{"partition":2}}}}'
k set image sts/db redis=redis:8.2 # only db-2 gets 8.2
k patch ... -p '<JSON>' changes just the fields in that JSON snippet (15.40 covers patch). sts is the short name for statefulsets. With OnDelete as the strategy nothing updates until you delete a pod yourself.
When a StatefulSet and when a Deployment
Use a StatefulSet when any of these is true: pods need a stable network identity, each pod needs its own persistent disk, or start/stop order matters. Databases, message brokers, consensus systems, anything clustered.
Otherwise a Deployment. And at a bank, seriously consider whether the database should be in Kubernetes at all versus a managed database from a cloud provider - operating stateful systems is where most of the pain lives.
What you can now do:
- explain what a StatefulSet adds: names, order, per-pod disks
- reach one member by its DNS name through a headless Service
- update one member at a time with a partition