Why Jobs
Deployments, StatefulSets and DaemonSets keep pods running for ever. But a lot of work is meant to finish: a database migration, a nightly report, a backup, resizing a batch of images. Run those in a Deployment and they finish, exit 0, get restarted, and crash-loop (15.14). A Job runs pods until they succeed, then stops. A CronJob creates a Job on a schedule - Kubernetes' cron, or systemd timer (2.16).
What you need to know already: pods, restartPolicy, exit codes and CrashLoopBackOff timings (15.14), generators and $do (15.3), systemd timers and OnCalendar (2.16), kubectl logs (15.14).
Jobs
k create job pi --image=busybox:1.36 -- sh -c '...' creates a Job named pi whose pod prints two lines and exits:
$ k create job pi --image=busybox:1.36 -- sh -c 'echo computing; sleep 5; echo 3.14159'
job.batch/pi created
$ k get jobs,pods
NAME STATUS COMPLETIONS DURATION AGE
job.batch/pi Running 0/1 4s 4s
NAME READY STATUS RESTARTS AGE
pod/pi-tdrpn 1/1 Running 0 4s
$ k get jobs,pods
NAME STATUS COMPLETIONS DURATION AGE
job.batch/pi Complete 1/1 6s 10s
NAME READY STATUS RESTARTS AGE
pod/pi-tdrpn 0/1 Completed 0 10s
$ k logs job/pi
computing
3.14159
The Job columns: STATUS (Running, Complete, Failed), COMPLETIONS (successful pods / wanted), DURATION (how long it ran or has been running). The pod's STATUS Completed means its container exited 0. k logs job/pi reads the logs of the Job's pod without you looking up its name.
The completed pod stays (so you can read its logs) until the Job is deleted or ttlSecondsAfterFinished removes it. Pods get the labels job-name=pi and batch.kubernetes.io/job-name=pi, so k get pods -l job-name=pi finds them.
The fields:
apiVersion: batch/v1
kind: Job
metadata:
name: resize-images
spec:
completions: 6 # need 6 successful pods in total (default 1)
parallelism: 2 # run at most 2 at a time (default 1)
backoffLimit: 4 # give up after 4 failed retries (default 6)
activeDeadlineSeconds: 600 # kill everything after 10 minutes, whatever state
ttlSecondsAfterFinished: 3600 # delete the Job (and its pods) an hour after it ends
template:
spec:
restartPolicy: Never # REQUIRED: Never or OnFailure (Always is rejected)
containers:
- name: worker
image: busybox:1.36
command: ['sh', '-c', 'echo resizing; sleep 3']
- completions - how many pods must succeed in total.
- parallelism - how many may run at the same time.
- backoffLimit - how many failed retries before the Job gives up.
- activeDeadlineSeconds - a hard time limit for the whole Job.
- ttlSecondsAfterFinished - clean up automatically after it ends.
k create job has no flags for most of these: generate with $do, add them, apply.
# after applying the manifest above (the next mission runs one like it)
k get job resize-images -w
NAME STATUS COMPLETIONS DURATION AGE
resize-images Running 0/6 2s 2s
resize-images Running 2/6 5s 5s
resize-images Running 4/6 9s 9s
resize-images Complete 6/6 13s 13s
Two at a time, six in total. A Job with restartPolicy: Always is refused - it could never finish:
The Job "resize-images" is invalid: spec.template.spec.restartPolicy: Unsupported value: "Always": supported values: "OnFailure", "Never"
Never vs OnFailure - where retries happen
Never: a failed container fails the pod; the Job controller creates a new pod for the retry. You get a trail ofErrorpods, each with its own logs - good for debugging.OnFailure: the kubelet restarts the container in the same pod (CrashLoopBackOff timings, 15.14). One pod, rising RESTARTS, and--previouslogs only for the last attempt.
Either way, retries count against backoffLimit, with a doubling delay (10s, 20s, 40s ... capped at 6 minutes) between new pods. When it is exceeded:
# flaky = a Job whose command always exits 1, backoffLimit: 2
k get job flaky
NAME STATUS COMPLETIONS DURATION AGE
flaky Failed 0/1 71s 71s
k describe job flaky | tail -4
Warning BackoffLimitExceeded 2s job-controller Job has reached the specified backoff limit
(tail -4 = the last 4 lines, where the events are.) activeDeadlineSeconds is the other stop: reason DeadlineExceeded, "Job was active longer than specified deadline". It wins over backoffLimit.
CronJobs
A CronJob creates a Job on a schedule, written as the standard 5-field cron expression: minute, hour, day of month, month, day of week (* = every, */5 = every 5th, 1-5 = a range). It runs in the controller-manager's time zone - UTC on almost every cluster - unless you set timeZone:
*/5 * * * * every 5 minutes
0 2 * * * 02:00 every day
30 6 * * 1-5 06:30 on weekdays
@hourly = 0 * * * *
k create cronjob tick --image=busybox:1.36 --schedule='*/1 * * * *' -- date = a CronJob tick that runs date every minute. Quote the schedule: an unquoted * is a glob the shell expands into file names (6.6).
$ k create cronjob tick --image=busybox:1.36 --schedule='*/1 * * * *' -- date
cronjob.batch/tick created
$ k get cronjobs,jobs
NAME SCHEDULE TIMEZONE SUSPEND ACTIVE LAST SCHEDULE AGE
cronjob.batch/tick */1 * * * * <none> False 0 45s 2m10s
NAME STATUS COMPLETIONS DURATION AGE
job.batch/tick-29835121 Complete 1/1 1s 105s
job.batch/tick-29835122 Complete 1/1 1s 45s
The CronJob columns: SCHEDULE, TIMEZONE (<none> = the controller's, UTC), SUSPEND (True = paused), ACTIVE (Jobs running right now), LAST SCHEDULE (how long ago it last fired).
The Job name suffix is the scheduled time in minutes since the epoch (1970, like date +%s but in minutes) - which makes names unique per run and lets you see when each was due.
spec:
schedule: '0 2 * * *'
timeZone: Europe/Bucharest # optional, IANA name
concurrencyPolicy: Forbid # Allow (default) | Forbid | Replace
startingDeadlineSeconds: 300 # if a run is missed by more than this, skip it
successfulJobsHistoryLimit: 3 # default 3
failedJobsHistoryLimit: 1 # default 1
suspend: false
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers: [...]
jobTemplate is a whole Job spec (which contains a pod template). The history limits say how many finished Jobs to keep.
concurrencyPolicy answers "what if the previous run is still going?":
- Allow - start another alongside (fine for short jobs that can safely run twice, dangerous for backups)
- Forbid - skip the new run
- Replace - kill the old run and start the new one
Operational commands you will use:
k patch cronjob tick -p '{"spec":{"suspend":true}}' # pause the schedule
k create job manual-run --from=cronjob/tick # run it NOW, off-schedule
k get jobs --sort-by=.metadata.creationTimestamp
k patch ... -p '{"spec":{"suspend":true}}'- set one field (spec.suspend) without editing the rest (15.40).k create job NAME --from=cronjob/tick- make a Job from the CronJob's template right now. The standard way to test a CronJob without waiting; the Job gets the annotationcronjob.kubernetes.io/instantiate: manual.--sort-by=.metadata.creationTimestamp- list the Jobs oldest first.
What you can now do:
- write a Job with completions, parallelism and a retry limit
- choose Never vs OnFailure and find the failed attempt's logs
- write, suspend and manually trigger a CronJob