OnCallReady

Lesson 15.24 · Kubernetes: Architecture & Workloads · 17 min read

Jobs and CronJobs: run to completion, on a schedule

In plain words

Most workers in a shop are there all day: the cashier, the shelf-stacker. But some jobs are "do it once and you're done": count the stock tonight, clean the windows on Sundays. You hire someone for the task; when it's finished, they go home, and you keep their note saying what they counted. If they trip and fail, you let them try again a few times, then give up.

A Job runs pods until enough of them succeed (completions, parallelism), retries failures up to backoffLimit, and stops at activeDeadlineSeconds. restartPolicy must be Never (new pod per retry) or OnFailure (restart in place). A CronJob creates Jobs on a cron schedule, with concurrencyPolicy deciding what happens if the previous run is still going, and kubectl create job --from=cronjob/... to run one now.

Why Jobs

Deployments, StatefulSets and DaemonSets keep pods running for ever. But a lot of work is meant to finish: a database migration, a nightly report, a backup, resizing a batch of images. Run those in a Deployment and they finish, exit 0, get restarted, and crash-loop (15.14). A Job runs pods until they succeed, then stops. A CronJob creates a Job on a schedule - Kubernetes' cron, or systemd timer (2.16).

What you need to know already: pods, restartPolicy, exit codes and CrashLoopBackOff timings (15.14), generators and $do (15.3), systemd timers and OnCalendar (2.16), kubectl logs (15.14).

Jobs

k create job pi --image=busybox:1.36 -- sh -c '...' creates a Job named pi whose pod prints two lines and exits:

$ k create job pi --image=busybox:1.36 -- sh -c 'echo computing; sleep 5; echo 3.14159'
job.batch/pi created
$ k get jobs,pods
NAME           STATUS     COMPLETIONS   DURATION   AGE
job.batch/pi   Running    0/1           4s         4s

NAME           READY   STATUS    RESTARTS   AGE
pod/pi-tdrpn   1/1     Running   0          4s
$ k get jobs,pods
NAME           STATUS     COMPLETIONS   DURATION   AGE
job.batch/pi   Complete   1/1           6s         10s

NAME           READY   STATUS      RESTARTS   AGE
pod/pi-tdrpn   0/1     Completed   0          10s
$ k logs job/pi
computing
3.14159

The Job columns: STATUS (Running, Complete, Failed), COMPLETIONS (successful pods / wanted), DURATION (how long it ran or has been running). The pod's STATUS Completed means its container exited 0. k logs job/pi reads the logs of the Job's pod without you looking up its name.

The completed pod stays (so you can read its logs) until the Job is deleted or ttlSecondsAfterFinished removes it. Pods get the labels job-name=pi and batch.kubernetes.io/job-name=pi, so k get pods -l job-name=pi finds them.

The fields:

apiVersion: batch/v1
kind: Job
metadata:
  name: resize-images
spec:
  completions: 6              # need 6 successful pods in total (default 1)
  parallelism: 2              # run at most 2 at a time (default 1)
  backoffLimit: 4             # give up after 4 failed retries (default 6)
  activeDeadlineSeconds: 600  # kill everything after 10 minutes, whatever state
  ttlSecondsAfterFinished: 3600   # delete the Job (and its pods) an hour after it ends
  template:
    spec:
      restartPolicy: Never    # REQUIRED: Never or OnFailure (Always is rejected)
      containers:
      - name: worker
        image: busybox:1.36
        command: ['sh', '-c', 'echo resizing; sleep 3']

k create job has no flags for most of these: generate with $do, add them, apply.

# after applying the manifest above (the next mission runs one like it)
k get job resize-images -w
NAME            STATUS    COMPLETIONS   DURATION   AGE
resize-images   Running   0/6           2s         2s
resize-images   Running   2/6           5s         5s
resize-images   Running   4/6           9s         9s
resize-images   Complete  6/6           13s        13s

Two at a time, six in total. A Job with restartPolicy: Always is refused - it could never finish:

The Job "resize-images" is invalid: spec.template.spec.restartPolicy: Unsupported value: "Always": supported values: "OnFailure", "Never"

Never vs OnFailure - where retries happen

Either way, retries count against backoffLimit, with a doubling delay (10s, 20s, 40s ... capped at 6 minutes) between new pods. When it is exceeded:

# flaky = a Job whose command always exits 1, backoffLimit: 2
k get job flaky
NAME    STATUS   COMPLETIONS   DURATION   AGE
flaky   Failed   0/1           71s        71s
k describe job flaky | tail -4
  Warning  BackoffLimitExceeded  2s  job-controller  Job has reached the specified backoff limit

(tail -4 = the last 4 lines, where the events are.) activeDeadlineSeconds is the other stop: reason DeadlineExceeded, "Job was active longer than specified deadline". It wins over backoffLimit.

CronJobs

A CronJob creates a Job on a schedule, written as the standard 5-field cron expression: minute, hour, day of month, month, day of week (* = every, */5 = every 5th, 1-5 = a range). It runs in the controller-manager's time zone - UTC on almost every cluster - unless you set timeZone:

*/5 * * * *        every 5 minutes
0 2 * * *          02:00 every day
30 6 * * 1-5       06:30 on weekdays
@hourly            = 0 * * * *

k create cronjob tick --image=busybox:1.36 --schedule='*/1 * * * *' -- date = a CronJob tick that runs date every minute. Quote the schedule: an unquoted * is a glob the shell expands into file names (6.6).

$ k create cronjob tick --image=busybox:1.36 --schedule='*/1 * * * *' -- date
cronjob.batch/tick created
$ k get cronjobs,jobs
NAME                 SCHEDULE      TIMEZONE   SUSPEND   ACTIVE   LAST SCHEDULE   AGE
cronjob.batch/tick   */1 * * * *   <none>     False     0        45s             2m10s

NAME                      STATUS     COMPLETIONS   DURATION   AGE
job.batch/tick-29835121   Complete   1/1           1s         105s
job.batch/tick-29835122   Complete   1/1           1s         45s

The CronJob columns: SCHEDULE, TIMEZONE (<none> = the controller's, UTC), SUSPEND (True = paused), ACTIVE (Jobs running right now), LAST SCHEDULE (how long ago it last fired).

The Job name suffix is the scheduled time in minutes since the epoch (1970, like date +%s but in minutes) - which makes names unique per run and lets you see when each was due.

spec:
  schedule: '0 2 * * *'
  timeZone: Europe/Bucharest       # optional, IANA name
  concurrencyPolicy: Forbid        # Allow (default) | Forbid | Replace
  startingDeadlineSeconds: 300     # if a run is missed by more than this, skip it
  successfulJobsHistoryLimit: 3    # default 3
  failedJobsHistoryLimit: 1        # default 1
  suspend: false
  jobTemplate:
    spec:
      template:
        spec:
          restartPolicy: OnFailure
          containers: [...]

jobTemplate is a whole Job spec (which contains a pod template). The history limits say how many finished Jobs to keep.

concurrencyPolicy answers "what if the previous run is still going?":

Operational commands you will use:

k patch cronjob tick -p '{"spec":{"suspend":true}}'     # pause the schedule
k create job manual-run --from=cronjob/tick             # run it NOW, off-schedule
k get jobs --sort-by=.metadata.creationTimestamp

What you can now do:

Why it helps

Batch work is everywhere in platform teams: database migrations, nightly reports, backups, cleanup, certificate checks. Situations: a nightly backup CronJob with the default concurrencyPolicy: Allow starts a second backup while the first is still running and both fail; a Job with a sidecar never completes (fixed by native sidecars); a failed Job leaves one Error pod per retry, which is actually useful for debugging; you need to test a CronJob now without waiting for 2am. And schedules run in the controller-manager's time zone, UTC on almost every cluster, unless you set timeZone. exam tasks regularly ask for a Job or CronJob with specific completions and parallelism.

FAQ

What is the difference between restartPolicy Never and OnFailure in a Job?

With Never, a failed container fails the whole pod and the Job controller creates a new pod for the retry, leaving a trail of Error pods each with its own logs. With OnFailure, the kubelet restarts the container inside the same pod, so you see one pod with rising RESTARTS and only the last attempt's previous logs. Both count retries against backoffLimit. Always is rejected for Jobs.

What does concurrencyPolicy do in a CronJob?

It decides what happens if a run is due while the previous one is still running. Allow (the default) starts another alongside, fine for short idempotent jobs, dangerous for backups. Forbid skips the new run. Replace kills the old run and starts the new one. For most operational jobs, Forbid is the safe choice.

How do I run a CronJob immediately to test it?

kubectl create job manual-run --from=cronjob/NAME creates a Job from the CronJob's template right now, marked with the annotation cronjob.kubernetes.io/instantiate: manual. It's the standard way to test a CronJob, or to trigger an extra run, without waiting for the schedule.

Why are completed Job pods still around?

Deliberately, so you can read their logs. They stay until the Job is deleted, until ttlSecondsAfterFinished removes the Job and its pods, or, for CronJobs, until the history limits (successfulJobsHistoryLimit 3, failedJobsHistoryLimit 1) prune old Jobs. Set a TTL for ad-hoc Jobs to avoid clutter.

What time zone do CronJob schedules use?

The kube-controller-manager's local time zone, which is UTC on almost every cluster, managed ones included. Set spec.timeZone with an IANA name, like Europe/Bucharest, to run in local time. Be aware of daylight saving changes when scheduling around the transition hours.

In an interview Junior

What is the difference between a Job and a Deployment, and how do CronJobs work?

A Deployment keeps pods running for ever (restartPolicy: Always); a finished process gets restarted. A Job runs pods until they succeed, then stops - for migrations, reports, backups. Its restartPolicy must be OnFailure or Never.

Job fields: completions (successes needed), parallelism (at the same time), backoffLimit (failed retries before giving up), activeDeadlineSeconds (a hard time limit), ttlSecondsAfterFinished (clean-up). With Never each retry is a new pod with its own logs; with OnFailure the kubelet restarts the container in place.

A CronJob creates a Job on a schedule, a 5-field cron expression in UTC unless timeZone is set. concurrencyPolicy decides what happens if the last run is still going: Allow, Forbid (right for backups) or Replace.

Operating it: k create job manual-run --from=cronjob/tick runs it now; suspend: true pauses it; k logs job/NAME reads the logs.

Also asked: How would you design a nightly backup as a CronJob? · What does backoffLimit do? · What is the difference between restartPolicy Never and OnFailure in a Job?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.