OnCallReady

Lesson 17.28 · Kubernetes: Scheduling, Health & Security · 13 min read

The HorizontalPodAutoscaler

In plain words

Imagine an ice-cream stand where a manager checks the queue every few minutes. The rule is: each seller should be busy about half the time. If they're all busy 90% of the time, the manager calls in more sellers, as many as needed to bring it back to half. When the rush is over, the manager waits a while before sending anyone home, in case another rush comes.

That's the HorizontalPodAutoscaler: every 15 seconds it reads the pods' metric, computes ceil(current replicas x current / target), and writes the new spec.replicas. CPU utilization is measured as a percentage of the pods' requests, so without requests it can't work. behavior controls how fast it scales up and how long it waits (5 minutes by default) before scaling down.

A control loop over the replica count

The problem. At 09:00 traffic triples; at 03:00 it is almost zero. A fixed replica count is either too few in the morning or wasted money at night. The HorizontalPodAutoscaler (HPA) changes the replica count for you, based on measured load. ("Horizontal" = more or fewer pods; "vertical" would be bigger pods.)

What you need to know already: requests (17.1), metrics-server and kubectl top (17.1), the reconciliation loop and kube-controller-manager (15.5, 15.9), Deployments (15.16).

The HPA controller (a loop inside kube-controller-manager) runs every 15 seconds. For each HPA it reads the metric of the target's pods from the metrics API (metrics-server for CPU and memory), computes a desired replica count and writes it into the target's spec.replicas through the scale subresource (the small part of the API that changes only the replica count - what kubectl scale uses too). That is all: the Deployment does the rest. (ceil below = round up to the next whole number.)

desiredReplicas = ceil( currentReplicas x currentMetricValue / targetMetricValue )

With 3 pods at 90% CPU utilization and a 50% target: ceil(3 x 90 / 50) = ceil(5.4) = 6. If the ratio is within 10% of 1 (the default tolerance), nothing changes - that stops it chasing noise.

Utilization is a percentage of the pods' requests. 50% of a 200m request is 100m. So:

An HPA on CPU utilization cannot work without CPU requests. No request = no denominator.

# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl get hpa -n shop
NAME   REFERENCE        TARGETS              MINPODS   MAXPODS   REPLICAS   AGE
web    Deployment/web   cpu: <unknown>/50%   1         10        1          3m
kubectl describe hpa web -n shop | tail -3
  Warning  FailedGetResourceMetric       15s (x12 over 3m)  horizontal-pod-autoscaler  failed to get cpu utilization: missing request for cpu in container web of Pod web-5d8f7c9b4d-2kq9x
  Warning  FailedComputeMetricsReplicas  15s (x12 over 3m)  horizontal-pod-autoscaler  invalid metrics (1 invalid out of 1), first error is: failed to get cpu resource metric value: failed to get cpu utilization: missing request for cpu in container web of Pod web-5d8f7c9b4d-2kq9x

<unknown> is also what you see in the first ~30 seconds (no metrics yet) and when metrics-server is down (unable to fetch metrics from resource metrics API). The sidecar version of the trap: every container in the pod needs the request.

Creating one

# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl autoscale deployment php-apache --cpu=50% --min=1 --max=10
horizontalpodautoscaler.autoscaling/php-apache autoscaled

kubectl autoscale deployment NAME creates an HPA for that Deployment: --cpu=50% = keep average CPU at 50% of the request, --min=1 / --max=10 = never fewer / more replicas than that.

(--cpu-percent=50 still works but is deprecated since 1.34; --cpu=500m would target an average value instead of utilization.) The object it creates is autoscaling/v2:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: php-apache
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: php-apache
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 50

Metric types: Resource (cpu/memory of the pods), ContainerResource (one container only - useful with sidecars), Pods and Object (custom metrics - numbers your app reports, such as requests per second, served to the HPA by an extra add-on called a metrics adapter), External (a number from outside the cluster, such as the length of a message queue). With several metrics, the HPA computes a count for each and takes the highest.

Reading it

# an illustration: needs metrics-server and a php-apache Deployment (the HPA mission)
kubectl get hpa php-apache
NAME         REFERENCE               TARGETS        MINPODS   MAXPODS   REPLICAS   AGE
php-apache   Deployment/php-apache   cpu: 47%/50%   1         10        5          4m
kubectl describe hpa php-apache
Metrics:                                               ( current / target )
  resource cpu on pods  (as a percentage of request):  47% (94m) / 50%
Min replicas:                                          1
Max replicas:                                          10
Deployment pods:                                       5 current / 5 desired
Conditions:
  Type            Status  Reason              Message
  ----            ------  ------              -------
  AbleToScale     True    ReadyForNewScale    recommended size matches current size
  ScalingActive   True    ValidMetricFound    the HPA was able to successfully calculate a replica count from cpu resource utilization (percentage of request)
  ScalingLimited  False   DesiredWithinRange  the desired count is within the acceptable range
Events:
  Normal  SuccessfulRescale  74s  horizontal-pod-autoscaler  New size: 5; reason: cpu resource utilization (percentage of request) above target

The three conditions are the diagnosis: AbleToScale (can it read/write the target's scale?), ScalingActive (does it have valid metrics? False = look at the reason), ScalingLimited (is it pinned at min or max? TooManyReplicas = raise max or fix the app).

behavior: stopping the flapping

Load is spiky; scaling on every spike is expensive and scaling down too eagerly means scaling up again a minute later. behavior controls both directions:

spec:
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0        # react immediately (default)
      policies:                            # default: +100% or +4 pods per 15s, whichever is more
      - {type: Percent, value: 100, periodSeconds: 15}
      - {type: Pods, value: 4, periodSeconds: 15}
      selectPolicy: Max
    scaleDown:
      stabilizationWindowSeconds: 300      # default: use the HIGHEST recommendation of the last 5 min
      policies:
      - {type: Percent, value: 100, periodSeconds: 15}

The scale-down window is why load disappears and replicas stay up for five minutes - by design. The condition says so: AbleToScale True ScaleDownStabilized recent recommendations were higher than current one, applying the highest recent recommendation. Shorten it for bursty batch, lengthen it for services with slow startup (a JVM that takes 90s to become Ready should not be scaled down and up every two minutes).

Living with an HPA

What you can now do

Why it helps

Autoscaling is where resource settings, metrics and capacity meet, and it's easy to get wrong in ways that look like app problems. TARGETS: <unknown> with "missing request for cpu" is the classic ticket: the team forgot requests on a sidecar. A sawtooth in replica count after every deploy means the manifest sets replicas and fights the HPA. Replicas that stay up five minutes after load drops is the stabilization window doing its job.

On a platform team you'll tune behavior for slow-starting JVMs, pair the HPA with the cluster autoscaler so new pods have somewhere to go, and decide when CPU is the wrong signal and requests per second or queue depth (custom metrics, KEDA) should drive scaling. kubectl autoscale is standard CKA material.

FAQ

Why does my HPA show TARGETS <unknown>?

It has no usable metric. Most often a container in the pod has no CPU request, so utilization has no denominator ("missing request for cpu"); every container needs it, sidecars included. It's also normal in the first ~30 seconds after creation, and it happens when metrics-server is down ("unable to fetch metrics from resource metrics API"). k describe hpa shows which.

Is CPU utilization measured against the limit or the request?

The request. A 50% target on pods requesting 200m means an average of 100m per pod. This is why HPA behaviour changes when someone changes requests: lower requests make the same load look like higher utilization and trigger more replicas.

Why doesn't it scale down immediately when load drops?

The default scale-down stabilization window is 300 seconds: the HPA uses the highest recommendation of the last five minutes. That avoids scaling down and up again for spiky load. The condition says ScaleDownStabilized. Shorten it for bursty batch work; keep or lengthen it for slow-starting services.

Should I keep replicas in my Deployment manifest when using an HPA?

No. Every kubectl apply resets the replica count to the manifest's value, the HPA corrects it 15 seconds later, and you get a sawtooth. Remove replicas from the manifest (or have whatever applies it ignore that field) once an HPA owns the Deployment. kubectl scale fights the HPA the same way.

Can I use HPA and VPA together?

Not on the same resource. If VPA changes CPU requests while the HPA scales on CPU utilization, each changes the other's input and they fight. A common combination is VPA in recommendation mode (or on memory only) plus HPA on CPU or custom metrics. The HPA can also use several metrics and takes the highest resulting replica count.

In an interview Junior

How does the Horizontal Pod Autoscaler work?

A control loop in kube-controller-manager, every 15 seconds: read the metric of the target's pods (CPU and memory from metrics-server), compute

desiredReplicas = ceil(currentReplicas x currentValue / targetValue)

and write it into the Deployment's replica count through the scale subresource. 3 pods at 90% with a 50% target -> ceil(5.4) = 6. Within 10% of the target it does nothing.

The traps:

kubectl autoscale deployment web --cpu=50% --min=2 --max=10 creates one; describe hpa conditions (AbleToScale, ScalingActive, ScalingLimited) explain what it is doing.

Also asked: An HPA is not scaling the way the team expects. How do you debug it? · Why does an HPA on CPU need resource requests? · When is CPU a bad metric for autoscaling?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.