kubectl autoscale - Auto-scale a deployment, replica set, stateful set, or replication controller
kubectl autoscale (-f FILENAME | TYPE NAME | TYPE/NAME) [--min=MINPODS] --max=MAXPODS [--cpu=TARGET] [--memory=TARGET]
Options you will use
--max=N- upper limit for the number of pods. Required.
--min=N- lower limit; if not specified the server applies its default (1).
--cpu=70%|500m- target CPU: a percentage is average UTILIZATION of the pods' CPU requests; a quantity is an average VALUE. Without --cpu/--memory the server default is 80% CPU utilization.
--memory=60%|200Mi- target memory, same two forms
--cpu-percent=N- deprecated in 1.34: use --cpu=N%
--name=NAME- name of the HPA (default: the target's name)
--dry-run=client -o yaml- print the autoscaling/v2 HorizontalPodAutoscaler instead of creating it
Examples
$ kubectl autoscale deployment php-apache --cpu=50% --min=1 --max=10the upstream walkthrough
$ kubectl get hpa -wwatch TARGETS and REPLICAS change
$ kubectl describe hpa php-apacheConditions and Events explain every decision (or why it cannot decide)
Gotchas
- Utilization is measured against REQUESTS. A container with no cpu request makes the HPA print cpu: <unknown>/50% and the event "missing request for cpu in container ...".
- Scale-up has no stabilization by default; scale-down waits 300s (the highest recommendation of the last 5 minutes wins) to stop flapping.
Try kubectl autoscale in a real terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.