OnCallReady

Lesson 15.26 · Kubernetes: Architecture & Workloads · 20 min read

Labels, selectors, annotations, namespaces

In plain words

Imagine a school where children wear coloured stickers: "red team", "year 3", "choir". Teachers never call children by list; they say "all year-3 choir members to the hall", and whoever wears those stickers goes. Peel off a sticker and the child is no longer in that group. Notes pinned to their backpacks, like "allergic to nuts", are information for adults but nobody gathers children by them.

Labels are the stickers: key/value pairs that selectors match (-l app=web, tier in (frontend,backend), matchLabels, matchExpressions). ReplicaSets, Services, DaemonSets and NetworkPolicies all find their pods this way, so labels are the wiring. Annotations are the backpack notes: larger, not selectable, for tools and humans. Namespaces are separate schools: scopes for names, RBAC and quotas, but not a network boundary.

Why this matters

What you need to know already: objects and manifests (15.1, 15.3), pods (15.14), ReplicaSets and Deployments (15.16), DaemonSets (15.22), shell quoting (6.6), DNS names (8.16).

A Deployment has three pods. How does it know which three, out of hundreds in the cluster? Not by name, not by a list - by labels. One wrong letter in a label and a pod silently stops belonging to its Deployment, or a test pod gets swallowed by the wrong one. This lesson is the wiring of the whole cluster.

Labels are the wiring

A label is a key/value pair in an object's metadata.labels, like app: web. A selector is a query over labels: "every pod with app=web". Kubernetes has no other way to say "these pods belong to that thing": a ReplicaSet finds its pods by label selector, a DaemonSet too, and so do the objects of the next chapters.

Later (Ch 16-17): Services (a stable address in front of pods), NetworkPolicies and PodDisruptionBudgets all find their pods with the same selectors.

Get a label wrong and things silently stop being connected.

metadata:
  labels:
    app.kubernetes.io/name: checkout
    app.kubernetes.io/instance: checkout-prod
    app.kubernetes.io/version: "2.4.1"
    app.kubernetes.io/component: api
    app.kubernetes.io/part-of: shop
    app.kubernetes.io/managed-by: kubectl
    tier: backend
    env: prod

The app.kubernetes.io/* keys are the recommended labels - a naming convention most tools and dashboards understand (managed-by names the tool that deployed it). Keys can have a DNS-prefix (app.kubernetes.io/) and a name up to 63 characters; values up to 63 characters of [A-Za-z0-9._-], starting and ending alphanumeric (so no spaces, no slashes, and "2.4.1" quoted in YAML so it stays a string).

Selecting

-l (--selector) on get, delete, logs and most other commands filters by label (k is the kubectl alias from 15.3):

k get pods -l app=web                        # equality
k get pods -l app=web,tier=frontend          # AND
k get pods -l 'env!=prod'                    # not equal (or key missing)
k get pods -l 'tier in (frontend,backend)'   # set-based
k get pods -l 'env notin (prod)'
k get pods -l tier                           # key exists
k get pods -l '!tier'                        # key does not exist

Quote anything with spaces, parentheses or ! - the shell would eat them.

# shop = the labelled namespace of the next mission
k get pods -n shop -l 'tier in (frontend,backend)' -L tier,env
NAME                        READY   STATUS    RESTARTS   AGE   TIER       ENV
api-5d9c7b8f6d-2kx8p        1/1     Running   0          2m    backend    prod
web-6b8d9c7f5d-8xk2p        1/1     Running   0          2m    frontend   prod
web-6b8d9c7f5d-n4v7q        1/1     Running   0          2m    frontend   staging

-L key adds a column per label (here TIER and ENV); --show-labels adds one LABELS column with all of them. The other columns are the usual pod columns (15.14).

In manifests the same two flavours appear as matchLabels and matchExpressions:

selector:
  matchLabels:
    app: web
  matchExpressions:
  - key: tier
    operator: In              # In, NotIn, Exists, DoesNotExist
    values: [frontend, edge]

matchLabels is the equality form (every listed label must match); matchExpressions is the set-based form. When both are given, both must match. Some older objects only support the simple map form (selector: {app: web}).

Changing labels on the fly

k run NAME --image=IMG -l k=v,... starts a single pod with labels; k label OBJECT key=value sets a label on an existing object:

$ k run demo --image=nginx:1.27 -l app=demo,env=prod
pod/demo created
$ k label pod demo env=staging
error: 'env' already has a value (prod), and --overwrite is false
$ k label pod demo env=staging --overwrite
pod/demo labeled
$ k label pod demo env-
pod/demo unlabeled

KEY- removes. --overwrite is required to change an existing value - a guard against typos in automation.

The quarantine trick

Because a ReplicaSet owns pods that match its selector, changing a pod's label takes it out of the ReplicaSet:

$ k create deployment web --image=nginx:1.27 --replicas=3     # AlreadyExists is fine
deployment.apps/web created
$ k label $(k get pods -l app=web -o name | head -1) app=web-debug --overwrite
pod/web-6b8d9c7f5d-8xk2p labeled
$ k get pods -L app
NAME                   READY   STATUS              RESTARTS   AGE   APP
web-6b8d9c7f5d-8xk2p   1/1     Running             0          5m    web-debug
web-6b8d9c7f5d-n4v7q   1/1     Running             0          5m    web
web-6b8d9c7f5d-zq9hm   1/1     Running             0          5m    web
web-6b8d9c7f5d-q2l8c   0/1     ContainerCreating   0          1s    web

($(k get pods -l app=web -o name | head -1) picks the first pod's name, pod/web-..., with command substitution from Ch 6.)

The ReplicaSet released it (removed its ownerReference, the field in metadata that says which controller owns an object) and created a replacement to get back to 3. The misbehaving pod keeps running, no longer counted and no longer sent traffic, for you to exec into and debug at leisure. Delete it when you are done - nothing else will.

Adoption, and why a Deployment's selector has a hash in it

The reverse also happens. A controller adopts any pod matching its selector that has no other controller. A Deployment protects itself: the ReplicaSets it creates select on app=web,pod-template-hash=6b8d9c7f5d, so a stray pod with just app=web is not adopted by them. A DaemonSet or a bare ReplicaSet has no such hash in its selector - run a test pod with the same labels as a DaemonSet's pods and the DaemonSet adopts it, sees two pods on that node, and deletes one. Usually yours.

Annotations

Also key/value metadata, but not selectable and much larger (up to 256KB total). They are for tools and humans, not for wiring:

kubernetes.io/change-cause                     rollout history
kubectl.kubernetes.io/last-applied-configuration   kubectl apply's memory
kubectl.kubernetes.io/restartedAt              kubectl rollout restart
deployment.kubernetes.io/revision              the deployment controller
checksum/config: 3f7a...                       a hash of the config, to force rollouts (15.29)
k annotate deploy web owner="[email protected]" contact=@payments-oncall
k annotate deploy web owner-            # remove

Rule of thumb: if something needs to select by it, label; otherwise annotate.

Namespaces

A namespace is a folder for objects: a scope for names, and later for permissions and limits (who may do what, how much CPU a team gets - Ch 17). Four exist on every cluster built with kubeadm (the standard installer, Ch 18):

$ k get ns
NAME              STATUS   AGE
default           Active   12d
kube-node-lease   Active   12d      node heartbeats (Lease objects)
kube-public       Active   12d      readable by everyone, e.g. cluster-info
kube-system       Active   12d      the cluster's own components

Columns: NAME, STATUS (Active, or Terminating while being deleted), AGE.

Not everything is namespaced - nodes and namespaces themselves are cluster-scoped (one per cluster, no namespace), as are a few kinds you meet later. k api-resources --namespaced=false lists them.

Namespaces are not a network boundary: by default any pod can reach any pod in any namespace (Ch 16 adds the rules that change that). And names only need to be unique within one: shop/web and payments/web are different objects, and DNS tells them apart as web.shop.svc.cluster.local and web.payments.svc.cluster.local.

Stop typing -n all the time. k config set-context --current --namespace=shop changes the default namespace of your current kubeconfig context (15.1); k config get-contexts shows it in the NAMESPACE column, with * marking the current context:

$ k config set-context --current --namespace=shop
Context "kubernetes-admin@kubernetes" modified.
$ k config get-contexts
CURRENT   NAME                          CLUSTER      AUTHINFO           NAMESPACE
*         kubernetes-admin@kubernetes   kubernetes   kubernetes-admin   shop

Now k get pods means shop. The trap is forgetting you did it - k get pods returning "No resources found in shop namespace." is the reminder. Switch back before you go on:

$ k config set-context --current --namespace=default
Context "kubernetes-admin@kubernetes" modified.

Deleting a namespace deletes everything in it and waits for it to be empty (the namespace shows Terminating meanwhile). There is no confirmation prompt.

What you can now do:

Why it helps

A wrong label is one of the quietest failures in Kubernetes: a Service with a selector that matches nothing has no endpoints, and nothing errors. You'll debug that with k get pods --show-labels and -l. The quarantine trick, relabelling a misbehaving pod so it leaves its ReplicaSet and Service while you debug it, is a genuinely useful incident technique. Using the recommended app.kubernetes.io/* labels makes deploy tools, dashboards and cost tools work. And setting a default namespace on your context, then forgetting it, produces the "No resources found in shop namespace" moment everyone has once. Selector syntax is basic exam material.

FAQ

What is the difference between labels and annotations?

Labels are small key/value pairs used for selection: controllers, Services and policies find objects by them, and -l filters on them. Values are limited to 63 characters of a restricted set. Annotations can't be selected on and can be much larger (up to 256KB total); they hold information for tools and humans, like kubernetes.io/change-cause, last-applied-configuration or deployment.kubernetes.io/revision. If something must select by it, label; otherwise annotate.

Are namespaces a security boundary?

Partly. They scope names, RBAC permissions, ResourceQuotas and Pod Security settings, so they're the unit of multi-tenancy for policy. But they are not a network boundary: by default any pod can reach any pod in any namespace. You need NetworkPolicies for network isolation (chapter 16), and namespaces don't isolate nodes or the kernel.

Why does kubectl label refuse to change my label?

Changing an existing label's value requires --overwrite, as a guard against typos in automation: kubectl label pod x env=staging --overwrite. To remove a label, use the key with a trailing dash: kubectl label pod x env-. The same syntax applies to annotations.

What happens if I change a pod's label so it no longer matches its ReplicaSet?

The ReplicaSet releases it (removes its ownerReference) and creates a replacement, because it now sees one too few. The relabelled pod keeps running, outside the Service too if the Service's selector no longer matches, so you can exec into it and debug. Delete it yourself when done; no controller will.

Which objects are not namespaced?

Nodes, Namespaces themselves, PersistentVolumes, StorageClasses, ClusterRoles and ClusterRoleBindings, CRDs and a few others. kubectl api-resources --namespaced=false lists them. For these, -n is ignored, and RBAC for them needs ClusterRoles rather than namespaced Roles. PersistentVolumeClaims, by contrast, are namespaced, while the PersistentVolumes they bind to are not.

In an interview Junior

What are labels and selectors used for in Kubernetes?

A label is a key/value pair in metadata.labels (app: web); a selector is a query over labels (app=web). They are the wiring of the cluster: a ReplicaSet finds its pods by selector, a DaemonSet too, and so do Services. There is no other link.

Useful consequences: relabelling a pod takes it out of its ReplicaSet (it gets replaced, and you can debug the old one in peace); a stray pod with matching labels can be adopted - a Deployment's ReplicaSets add pod-template-hash to their selector to prevent that.

Annotations are metadata for tools and humans, not selectable. Namespaces scope names, not network traffic.

Also asked: What is the difference between a label and an annotation? · What is a namespace, and what does it not isolate? · How do you change the default namespace of your kubectl context?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.