OnCallReady

Lesson 17.9 · Kubernetes: Scheduling, Health & Security · 17 min read

LimitRange and ResourceQuota: per-pod defaults vs namespace totals

In plain words

Think of a family holiday budget. One rule says "each kid gets 10 euros of pocket money a day unless they ask for something else, and never more than 20 for one thing" (that's per person). A second rule says "the whole family may spend at most 500 euros this week" (that's the total). The first rule fills in a default and caps each purchase; the second only watches the total.

A LimitRange is the per-person rule: it fills in default requests and limits for containers that set none and enforces min and max per container. A ResourceQuota is the family budget: it caps the namespace's total requests, limits and object counts. Once a quota counts requests.cpu, every pod must declare it, which is why the two come as a pair.

Two admission plugins, two jobs

The problem. Ten teams share one cluster. One team forgets resources on every pod; another deploys 200 replicas by mistake and eats every node. The platform team needs defaults for the forgetful and a budget for the greedy - per namespace.

What you need to know already: namespaces (15.26), requests/limits and QoS (17.1, 17.3), the path from kubectl apply to a running pod (15.11), ReplicaSets (15.16).

Two words first:

LimitRangeResourceQuota
scopeeach container / pod / PVCthe whole namespace
doesfills in defaults, enforces min/max per objectcaps the sum (requests, limits, object counts)
whenadmission, on createadmission on create, plus a controller keeping status.used
typical use"every container gets 100m/128Mi unless it says otherwise""team-a may use at most 8 cores and 16Gi in total"

A platform team usually installs both in every tenant namespace: the LimitRange so nobody ends up BestEffort by accident, the quota so no team can eat the cluster.

LimitRange

apiVersion: v1
kind: LimitRange
metadata:
  name: defaults
  namespace: team-a
spec:
  limits:
  - type: Container
    defaultRequest:        # requests when the container sets none
      cpu: 100m
      memory: 128Mi
    default:               # limits when the container sets none
      cpu: 500m
      memory: 256Mi
    max:                   # no container may set more than this
      cpu: "1"
      memory: 1Gi
    min:
      memory: 32Mi

Read the spec: type: Container = the rules apply to each container; defaultRequest / default = the requests / limits filled in when missing; max / min = the allowed range. Defaults are applied by the LimitRanger admission plugin at create time. The pod records it in an annotation (a free-text note on an object, 15.26):

# an illustration: team-a with the LimitRange and quota above (the quota mission builds it)
kubectl run bare --image=nginx:1.27 -n team-a
pod/bare created
kubectl get pod bare -n team-a -o jsonpath='{.metadata.annotations}'
{"kubernetes.io/limit-ranger":"LimitRanger plugin set: cpu, memory request for container bare; cpu, memory limit for container bare"}
kubectl get pod bare -n team-a -o jsonpath='{.spec.containers[0].resources}'
{"limits":{"cpu":"500m","memory":"256Mi"},"requests":{"cpu":"100m","memory":"128Mi"}}

The ordering rules worth knowing:

min/max violations are admission errors on the pod:

$ kubectl run big --image=nginx:1.27 -n team-a --dry-run=client -o yaml > big.yaml   # then set limits.cpu: 2
$ kubectl apply -f big.yaml
Error from server (Forbidden): error when creating "big.yaml": pods "big" is forbidden: maximum cpu usage per Container is 1, but limit is 2
# an illustration: team-a with the LimitRange and quota above (the quota mission builds it)
kubectl describe limitrange defaults -n team-a
Name:       defaults
Namespace:  team-a
Type        Resource  Min   Max  Default Request  Default Limit  Max Limit/Request Ratio
----        --------  ---   ---  ---------------  -------------  -----------------------
Container   cpu       -     1    100m             500m           -
Container   memory    32Mi  1Gi  128Mi            256Mi          -

ResourceQuota

# an illustration: team-a with the LimitRange and quota above (the quota mission builds it)
kubectl create quota compute -n team-a --hard=requests.cpu=2,requests.memory=4Gi,limits.cpu=4,limits.memory=8Gi,pods=20
resourcequota/compute created

kubectl create quota NAME creates a ResourceQuota; --hard= lists the ceilings as key=value pairs: requests.cpu=2 = the requests of all pods together may add up to 2 cores, limits.memory=8Gi = all memory limits together at most 8Gi, pods=20 = at most 20 pods. kubectl get quota / describe quota show Used against Hard:

# an illustration: team-a with the LimitRange and quota above (the quota mission builds it)
kubectl get quota -n team-a
NAME      AGE   REQUEST                                                     LIMIT
compute   12m   pods: 5/20, requests.cpu: 500m/2, requests.memory: 640Mi/4Gi   limits.cpu: 2500m/4, limits.memory: 1280Mi/8Gi
kubectl describe quota compute -n team-a
Name:            compute
Namespace:       team-a
Resource         Used    Hard
--------         ----    ----
limits.cpu       2500m   4
limits.memory    1280Mi  8Gi
pods             5       20
requests.cpu     500m    2
requests.memory  640Mi   4Gi

Also countable: services, configmaps, secrets, persistentvolumeclaims, requests.storage, and any kind as count/deployments.apps, count/jobs.batch.

The side effect everyone trips on: once a quota covers requests.cpu (or limits.memory, etc.), every new pod must specify that resource, because the quota system cannot count what is not declared:

Error from server (Forbidden): pods "bare" is forbidden: failed quota: compute: must specify limits.cpu for: bare; limits.memory for: bare; requests.cpu for: bare; requests.memory for: bare

That is precisely why quota and LimitRange come as a pair: the LimitRange fills the values in, the quota then counts them.

Going over the total:

Error from server (Forbidden): pods "big" is forbidden: exceeded quota: compute, requested: requests.cpu=500m, used: requests.cpu=1800m, limited: requests.cpu=2

Where the error goes when a controller creates the pod

You never see those errors when a Deployment creates the pods - kubectl talked to the apiserver about the Deployment, which was fine. The ReplicaSet controller (the control-plane loop that creates pods for a ReplicaSet, 15.9) gets the error, and records it as an event on the ReplicaSet:

# an illustration: team-a with the LimitRange and quota above (the quota mission builds it)
kubectl get deploy many -n team-a
NAME   READY   UP-TO-DATE   AVAILABLE   AGE
many   4/6     4            4           2m
kubectl describe rs -n team-a -l app=many | tail -2
  Warning  FailedCreate      8s (x5 over 2m)  replicaset-controller  Error creating: pods "many-7c9f4d8b6c-" is forbidden: exceeded quota: compute, requested: pods=1,requests.cpu=250m, used: pods=4,requests.cpu=1, limited: pods=4,requests.cpu=1

(pods "many-7c9f4d8b6c-" - the name is still the generateName prefix: admission rejected it before a name was generated.) The same pattern applies to every admission rejection: Pod Security (lesson 17.39), quotas, a missing ServiceAccount (the identity a pod runs as, lesson 17.33). Deployment stuck below its replica count with no Pending pods = look at the ReplicaSet's events.

Quota and the cluster admin

A quota is namespaced and applies to everyone, cluster-admin (the all-powerful admin user) included. It is not RBAC (the permission system, lesson 17.30): RBAC says whether you may create a pod, quota says whether this one still fits. kubectl describe namespace team-a shows both the quotas and the limit ranges in one place - the first command when a team says "we cannot deploy".

What you can now do

Why it helps

Every multi-tenant cluster uses both, and you'll be the one setting them up per team namespace. The support ticket that follows is predictable: "we can't deploy, the Deployment shows 4/6 and there are no Pending pods". You'll know that admission rejected the pods, so the error lives on the ReplicaSet's events (FailedCreate ... exceeded quota), not on the Deployment or in kubectl's output.

You'll also see teams hit "must specify limits.cpu" the day a quota is introduced, and LimitRange defaults that reject pods whose request exceeds the default limit. Understanding the ordering rules lets you write defaults that don't break people. kubectl describe namespace showing both quotas and limit ranges is your first command for "we cannot deploy".

Commands in this lesson

kubectl

FAQ

Why does my pod fail with "must specify limits.cpu" after a quota was added?

A ResourceQuota that tracks a resource (limits.cpu, requests.memory, and so on) requires every new pod to declare it, because it can't count what isn't declared. Either set the values in the pod spec or add a LimitRange with default and defaultRequest so admission fills them in. That's why quotas and LimitRanges are deployed together.

My Deployment is stuck at 4/6 and there are no Pending pods. Where's the error?

On the ReplicaSet. kubectl only talked to the API server about the Deployment, which was fine. The ReplicaSet controller tries to create the pods and admission rejects them (quota, Pod Security, missing ServiceAccount), so it records FailedCreate events on the ReplicaSet: k describe rs -l app=<name>. No pod object exists, so there's nothing Pending to look at.

Does a LimitRange change existing pods?

No. Defaults and min/max are applied by the LimitRanger admission plugin when pods are created. Existing pods keep whatever they had. Create the LimitRange before the workloads, or roll the workloads afterwards so the new pods get the defaults. You can see what was applied in the pod's kubernetes.io/limit-ranger annotation.

Does a ResourceQuota apply to cluster-admins too?

Yes. A quota is an admission check on the namespace's totals, not an authorization rule. RBAC decides whether you may create a pod at all; the quota decides whether this pod still fits in the namespace. Cluster-admin gets the same "exceeded quota" error as anyone.

Can a quota limit things other than CPU and memory?

Yes. Object counts (pods, services, configmaps, secrets, persistentvolumeclaims, and any kind as count/deployments.apps), storage (requests.storage, and per StorageClass), and extended resources like GPUs. Quotas can also be scoped, for example to a PriorityClass, so only some pods can use a high priority.

In an interview Junior

What is the difference between a LimitRange and a ResourceQuota?

Both are namespaced admission rules, with different jobs:

They come as a pair: once a quota covers a resource, every new pod must declare it ("must specify limits.cpu"), and the LimitRange's defaults are what fill it in. Going over gives "exceeded quota".

The trap when a Deployment creates the pods: you see no error - the ReplicaSet controller gets it. A Deployment stuck below its replica count with no Pending pods: read the ReplicaSet's events. kubectl describe namespace shows the quotas and limit ranges in one place.

Also asked: A team says they cannot deploy anymore. How do you investigate? · Why must a namespace with a ResourceQuota usually also have a LimitRange? · Where do you see the error when a quota blocks a Deployment's pods?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.