OnCallReady

Lesson 15.19 · Kubernetes: Architecture & Workloads · 15 min read

StatefulSets: stable identity, ordered start, headless Services

In plain words

In a football team, players aren't interchangeable like identical toy soldiers. Each has a shirt number, their own locker with their own boots, and they walk onto the pitch in order, goalkeeper first. If number 2 gets injured and a substitute comes on, they wear shirt number 2 and use locker 2.

A StatefulSet gives pods exactly that: stable names (db-0, db-1, db-2), ordered start and stop (db-1 waits for db-0 to be Ready; scale-down removes the highest first), a stable DNS name per pod through a headless Service (db-0.db.default.svc.cluster.local), and one PersistentVolumeClaim per pod from volumeClaimTemplates (data-db-0) that follows it and survives deleting the StatefulSet. Updates go from the highest ordinal down, and partition allows canaries.

Why StatefulSets

Deployments treat pods as interchangeable: random names, created and deleted in any order. Try running a 3-member database cluster that way and it breaks: members find each other by name, each keeps its own data on disk, and the first one must start before the others join it. A StatefulSet is the workload for that: pods with fixed names, a fixed start order and their own disks.

What you need to know already: Deployments and rolling updates (15.16), pods and their DNS settings (15.14), etcd's members and quorum (15.5), DNS A records and nslookup (8.18, 8.22), Docker volumes (11.22).

What a Deployment cannot give you

A Deployment's pods are interchangeable - perfect for stateless web servers and wrong for a database cluster, where each member has:

A StatefulSet provides exactly those three. It comes with a companion object, a Service - a stable name in front of pods (Ch 16 teaches Services). Here it is a special headless one, explained below:

apiVersion: v1
kind: Service
metadata:
  name: db
spec:
  clusterIP: None          # headless: no virtual IP, DNS returns pod IPs
  selector:
    app: db
  ports:
  - port: 6379
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: db
spec:
  serviceName: db          # the headless Service that gives pods their DNS names
  replicas: 3
  selector:
    matchLabels:
      app: db
  template:
    metadata:
      labels:
        app: db
    spec:
      containers:
      - name: redis
        image: redis:7.4

(--- separates two objects in one YAML file. Redis is a small in-memory database; port 6379 is its port.) The StatefulSet looks like a Deployment - replicas, selector, template - plus serviceName.

Stable names, ordered start

# after applying the db manifest above (the next mission does it with `cache`)
k get pods -w
NAME   READY   STATUS              RESTARTS   AGE
db-0   0/1     ContainerCreating   0          0s
db-0   1/1     Running             0          2s
db-1   0/1     ContainerCreating   0          0s
db-1   1/1     Running             0          2s
db-2   0/1     ContainerCreating   0          0s
db-2   1/1     Running             0          2s

Names are <statefulset>-<ordinal>: an ordinal is the member number, 0 to replicas-1, and it never changes. With the default podManagementPolicy: OrderedReady, db-1 is not created until db-0 is Running and Ready, db-2 not until db-1 is. Scaling down removes the highest ordinal first, one at a time.

Delete a member and it comes back with the same name (and, with storage, the same disk):

k delete pod db-1
pod "db-1" deleted from default namespace
k get pods -l app=db
NAME   READY   STATUS    RESTARTS   AGE
db-0   1/1     Running   0          5m
db-1   1/1     Running   0          3s
db-2   1/1     Running   0          5m

Same name, new pod (AGE 3s). Each pod also gets labels naming itself, which lets you pick out one member:

statefulset.kubernetes.io/pod-name=db-1
apps.kubernetes.io/pod-index=1

podManagementPolicy: Parallel drops the ordering (all pods at once) while keeping the stable names - use it when members do not depend on each other's start order.

The headless Service and per-pod DNS

A normal Service gets one virtual IP and spreads traffic over its pods. A headless Service (clusterIP: None) has no IP of its own; its DNS name resolves to the pod IPs directly, and - the reason StatefulSets need one - each pod gets its own DNS name served by CoreDNS:

<pod>.<service>.<namespace>.svc.cluster.local
db-0.db.default.svc.cluster.local

Test it from a throwaway pod. k run -it --rm dns --image=busybox:1.36 --restart=Never -- nslookup db-0.db: run a pod dns interactively (-it), delete it when it ends (--rm), do not restart it (--restart=Never), and run nslookup db-0.db in it:

k run -it --rm dns --image=busybox:1.36 --restart=Never -- nslookup db-0.db
Server:		10.96.0.10
Address:	10.96.0.10:53

Name:	db-0.db.default.svc.cluster.local
Address: 10.244.1.87

pod "dns" deleted from default namespace

Server is CoreDNS; the short name db-0.db was expanded with the search domains (8.27) to the full name, which points at db-0's pod IP.

That is how a Redis copy finds its primary (db-0.db), how an etcd member lists its peers. The IP changes when the pod is recreated; the name does not. serviceName must match an existing headless Service for this to work.

Later (Ch 16): Services in full - ClusterIP, headless, and how DNS names are built.

Stable storage: volumeClaimTemplates

spec:
  volumeClaimTemplates:
  - metadata:
      name: data
    spec:
      accessModes: [ReadWriteOnce]
      resources:
        requests:
          storage: 10Gi

A PersistentVolumeClaim (PVC) is a pod's request for a disk: "10Gi, usable by one node at a time" (ReadWriteOnce). It is Kubernetes' version of a named Docker volume (11.22) that outlives the container. volumeClaimTemplates makes the controller create one PVC per pod - data-db-0, data-db-1, data-db-2 - and always mount data-db-N into db-N, even after it moves to another node.

The PVCs survive deleting the StatefulSet (the default persistentVolumeClaimRetentionPolicy: Retain), so data outlives the workload. This cluster has no disk provider set up yet, so such PVCs would stay Pending, and their pods with them - which is why the labs here run without storage.

Updates

updateStrategy: RollingUpdate (the default) replaces pods from the highest ordinal down, one at a time, each only after the previous is Ready. rollingUpdate.partition: N updates only ordinals >= N - a way to try a new version on one member first (a canary):

k patch sts db -p '{"spec":{"updateStrategy":{"rollingUpdate":{"partition":2}}}}'
k set image sts/db redis=redis:8.2        # only db-2 gets 8.2

k patch ... -p '<JSON>' changes just the fields in that JSON snippet (15.40 covers patch). sts is the short name for statefulsets. With OnDelete as the strategy nothing updates until you delete a pod yourself.

When a StatefulSet and when a Deployment

Use a StatefulSet when any of these is true: pods need a stable network identity, each pod needs its own persistent disk, or start/stop order matters. Databases, message brokers, consensus systems, anything clustered.

Otherwise a Deployment. And at a bank, seriously consider whether the database should be in Kubernetes at all versus a managed database from a cloud provider - operating stateful systems is where most of the pain lives.

What you can now do:

Why it helps

You'll meet StatefulSets for Kafka, Redis, Elasticsearch, Zookeeper and operator-managed databases, and they fail differently from Deployments. Situations: a pod stuck because its PVC is Pending; db-1 not starting because db-0 isn't Ready (ordered start working as designed); someone force-deleting db-0 on an unreachable node and ending up with two writers on the same data; a StatefulSet deleted and its PVCs still there, by design. In a bank you'll also face the question "should this database run in Kubernetes at all?", and knowing what StatefulSets do and don't solve lets you argue for a managed service when appropriate. "StatefulSet vs Deployment" is a standard interview question.

FAQ

Why do StatefulSets need a headless Service?

The headless Service (clusterIP: None), named in serviceName, is what gives each pod its own DNS name, <pod>.<service>.<namespace>.svc.cluster.local. The StatefulSet sets each pod's hostname and subdomain to match it. Clustered software uses those stable names to find peers, like a replica finding its primary at db-0.db, even though pod IPs change on recreation.

What happens to the data when I delete a StatefulSet?

By default the PVCs created from volumeClaimTemplates are kept (persistentVolumeClaimRetentionPolicy Retain), so the data survives, and a recreated StatefulSet with the same name reattaches data-db-0 to db-0. You must delete PVCs explicitly if you want the storage gone. The retention policy can be changed to Delete on scale-down or deletion.

Why is force-deleting a StatefulSet pod dangerous?

StatefulSets guarantee at most one pod per identity. Force-deleting removes the API object without confirming the container stopped, so if the node is unreachable, the old db-0 may still be running and writing while the controller creates a new db-0 elsewhere, two writers on one identity and possibly one volume. Only force-delete when you're sure the old instance is dead.

Can StatefulSet pods start in parallel?

Yes, with podManagementPolicy: Parallel, which drops the ordering for creation and deletion but keeps stable names and storage. Use it when members don't depend on each other's startup order. The default, OrderedReady, creates pods one at a time in ordinal order, each waiting for the previous to be Running and Ready.

How do I canary an update on a StatefulSet?

Set updateStrategy.rollingUpdate.partition: N: only pods with ordinal N or higher get the new template. With 3 replicas and partition 2, only db-2 updates. Check it, then lower the partition to continue. OnDelete is the other option: nothing updates until you delete pods yourself.

In an interview Junior

When would you use a StatefulSet instead of a Deployment?

When the pods are not interchangeable - any of:

So: databases, message brokers, consensus systems. Everything stateless - web servers, APIs - is a Deployment.

And ask whether the database belongs in the cluster at all: a managed database is often the better answer.

Also asked: What is a headless Service, and why does a StatefulSet need one? · What happens to a StatefulSet's PVCs when you delete it? · Why should you not force-delete a StatefulSet pod on a node you cannot reach?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.