OnCallReady

Lesson 16.37 · Kubernetes: Networking & Storage · 8 min read

Volumes: where the data lives, and for how long

In plain words

Think about where you keep your stuff at school. Writing on your desk is gone when the cleaner wipes it at the end of the lesson (the container's own filesystem). A shared tray on the group's table lasts until the group is split up (an emptyDir, shared by the pod's containers, gone with the pod). A drawer in one specific classroom stays there, but only helps if you're put in that same room again (hostPath, tied to one node). A proper locker in the school basement is yours wherever you sit (a persistentVolumeClaim).

In Kubernetes, the volume type decides how long data lives. A container restart wipes its writable layer; an emptyDir survives restarts but not pod deletion; hostPath is one node's disk; a PVC points at storage that exists independently of any pod.

A container's own filesystem is disposable

Everything a container writes to its image filesystem lives in its writable layer (11.22), and that layer is thrown away when the container restarts - not only when the pod is deleted. A crash-looping database that writes to /var/lib/data without a volume starts empty on every restart. Volumes are how data outlives a container; the volume's type decides for how long.

What you need to know already: mounts (4.18), Docker volumes, bind mounts and tmpfs (11.22), pods, restarts and CrashLoopBackOff (15.14), sidecars (15.35), ConfigMap and Secret volumes (15.29, 15.33), kubectl debug node (16.3), memory limits and cgroups (5.11).

Declaring a volume

A volume in Kubernetes is storage attached to a pod. It takes two parts in the pod spec: volumes: (what the storage is, with a name) and, in each container, volumeMounts: (where that named volume appears in the container's filesystem, like docker run -v):

spec:
  containers:
  - name: app
    volumeMounts:
    - {name: cache, mountPath: /cache}
    - {name: node, mountPath: /node}
    - {name: data, mountPath: /data}
  volumes:
  - name: cache
    emptyDir: {}                      # sizeLimit: 1Gi, medium: Memory
  - name: node
    hostPath: {path: /var/lib/lab-notes, type: DirectoryOrCreate}
  - name: data
    persistentVolumeClaim: {claimName: notes}

The three types used here, and how long their data lives:

volume            lives as long as          survives container   survives pod   follows the pod
                                            restart              deletion       to another node
container layer   the container             no                   no             no
emptyDir          the POD (on its node)     yes                  no             no
hostPath          the node's disk           yes                  yes*           no - it is THAT node's dir
PVC (network)     the PersistentVolume      yes                  yes            yes (within its topology)
configMap/secret  the object (read-only)    yes                  yes            yes

* only if the next pod lands on the same node. "Topology" = where the storage can be reached from, for example one zone (16.41).

emptyDir

An emptyDir is created empty when the pod starts on a node, and deleted with the pod. All containers of the pod can mount it - the standard way for a sidecar and the app to exchange files, and for scratch space, caches and unpacked files.

hostPath - and why it is a security hole

A hostPath volume mounts a directory of the node into the pod - a bind mount (11.22) from the node. Legitimate users are node agents: log shippers reading /var/log, network and storage plugins, the kube-proxy DaemonSet. For applications it is wrong twice:

Later (Ch 17): Pod Security Standards - the "baseline" profile forbids hostPath entirely.

configMap, secret, projected

Read-only views of API objects (15.29, 15.33): files that update when the object changes (except when mounted with subPath, a single file). A projected volume merges several sources (the pod's ServiceAccount token, a ConfigMap, a Secret) into one directory.

persistentVolumeClaim

The pod does not name a disk. It names a claim - a PersistentVolumeClaim (PVC): "I need 10Gi, read-write, of class lab-disk". The cluster binds that claim to a PersistentVolume (PV), an actual piece of storage (a cloud disk, a network file share, a storage-cluster volume). The data lives on the volume, which exists independently of any pod:

# writer/reader: two pods sharing the claim above (the volumes mission)
k exec -n data writer -- sh -c 'echo hello > /data/f.txt'
k delete pod writer -n data --grace-period=0 --force
k apply -f reader.yaml            # a new pod, same claim
k exec -n data reader -- cat /data/f.txt
hello

sh -c '...' runs a small shell command inside the container. --grace-period=0 --force deletes the pod immediately, without the usual shutdown wait. The new pod using the same claim still finds the file.

The rest of this part is about that indirection: who creates volumes (16.39), who may mount them where (16.41), and what happens when the claim goes away (16.44).

What you can now do:

Why it helps

The "database starts empty after every crash" incident is exactly this: data written to the container layer instead of a volume. You'll see it with a team's first stateful app, and you'll recognise it from the restart count.

hostPath is a security review item you'll raise constantly: a pod that mounts the node's / can read the kubelet's credentials and every secret on disk, which is why the Pod Security baseline profile forbids it. You'll also explain to teams why their hostPath data "disappeared" after a reschedule, and why emptyDir with medium: Memory counts against their memory limit. This lesson is the vocabulary for every storage conversation that follows.

FAQ

Does a container restart lose data in an emptyDir?

No. An emptyDir belongs to the pod: it survives container restarts and crashes and is deleted only when the pod is removed from the node. Data written to the container's own filesystem (its writable layer) is lost on every restart. That's why a crash-looping app that writes without a volume starts empty each time.

Why is hostPath considered dangerous?

Because it mounts the node's filesystem into the pod. A pod that can mount / or /etc/kubernetes can read PKI keys and kubelet credentials, write static pod manifests, or reach /run/containerd/containerd.sock and start anything. That's node compromise. The Pod Security Standards baseline profile forbids hostPath. Legitimate users are node agents: log shippers, CNI and CSI plugins.

What happens when an emptyDir exceeds its sizeLimit?

The kubelet evicts the pod. The limit isn't a filesystem quota that makes writes fail; it's checked periodically, and a pod over the limit is evicted and rescheduled with a fresh, empty emptyDir. With medium: Memory the emptyDir is a tmpfs, and what you write counts against the container's memory limit.

Do ConfigMap and Secret volumes update when I change the object?

Yes, eventually: the kubelet syncs the files, typically within about a minute. Two exceptions: files mounted with subPath never update, and environment variables from ConfigMaps or Secrets are fixed at container start. The app must also re-read the file; many only read config at startup.

Why does the pod reference a claim instead of the disk?

To separate what the app needs from how the cluster provides it. The pod says "claim notes"; the claim says "10Gi, read-write, class lab-disk"; the cluster binds it to a PersistentVolume backed by a real disk or export. The same manifest works in a cloud with its disks and on-prem with Ceph, and the data outlives any pod that uses it.

In an interview Junior

Which volume would you use for scratch space, for files shared between containers, and for a database?

A container's writable layer is thrown away on every container restart, so anything that matters goes in a volume (volumes: in the pod, volumeMounts: in the container):

Not hostPath: it ties the data to one node and, worse, a pod that can mount the node's filesystem can read the cluster's keys and other pods' secrets and take over the node. It is for node agents only.

Also asked: Why do security teams restrict hostPath? · A team's app loses its data every time the pod restarts. What do you check? · What is the difference between emptyDir and a PersistentVolumeClaim?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.