OnCallReady

Lesson 33.26 · Ansible: Configuration as Code · 18 min read

Rolling changes: serial, health checks and delegate_to

In plain words

Imagine repainting the lanes of a busy road. If you close every lane at once, traffic stops. So you close one lane, paint it, check the paint is dry, open it again, then move to the next lane. If the paint in the first lane is the wrong colour, you stop there and only one lane was affected.

A rolling deploy does that with servers behind a load balancer: take one server out of the pool, change it, check it answers correctly, put it back, then the next one. If the check fails, Ansible stops the rollout and the other servers keep serving users.

The problem

Two web servers sit behind lb-1. A new release needs a new config and a restart. Ansible's default is to do every task on every host at the same time (5 in parallel): both web servers restart together, and for those seconds the load balancer has nothing to send traffic to. With a broken config, both stay down. A rolling change does a few hosts at a time, takes each one out of the load balancer first, checks it is healthy before putting it back, and stops the rollout at the first failure - so a bad release costs you one host, not the fleet.

What you need to know already: the handlers lesson (flush_handlers), the templates lesson, 2.10 (reload vs restart), 1.13 (addresses and ports).

serial: batches

serial: on a play splits its hosts into batches; the whole play (all tasks, handlers included) runs on one batch before the next starts:

- name: Rolling deploy
  hosts: web
  serial: 1            # one host at a time;  also: 2, "25%", or a ramp [1, 5, "50%"]

serial: [1, 5, "50%"] is a canary pattern: one host first, then five, then half the rest at a time. Each batch prints its own PLAY [...] banner.

Stopping a bad rollout

By default Ansible keeps going while at least one host in the batch survives. For a rollout that is the wrong instinct: if the new version fails on the first host, you want to stop there.

  serial: 1
  max_fail_percentage: 0     # any failure in a batch stops the whole play

max_fail_percentage: N aborts the play when more than N% of the current batch failed (Ansible prints NO MORE HOSTS LEFT and the batches after it never start). With serial and no max_fail_percentage, a batch where every host failed also stops the play. any_errors_fatal: true is the bigger hammer: the first failure on any host ends the play for all hosts, even in the middle of a batch.

$ cd ~/oncall-lab/labs/1a-ansible/try/rolling
$ cat roll.yml
- name: Rolling deploy
  hosts: web
  serial: 1
  max_fail_percentage: 0
  become: true

  pre_tasks:
    - name: Drain from the pool
      ansible.builtin.lineinfile:
        path: /etc/nginx/conf.d/upstream.conf
        regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
        line: "    server {{ ansible_facts['default_ipv4']['address'] }}:80 down;"
      delegate_to: lb-1
      notify: Reload lb

    - name: Apply the drain now
      ansible.builtin.meta: flush_handlers

  tasks:
    - name: Deploy the new site
      ansible.builtin.template:
        src: site.conf.j2
        dest: /etc/nginx/sites-available/default
        mode: "0644"
      notify: Reload nginx

    - name: Reload now, not at the end of the play
      ansible.builtin.meta: flush_handlers

  post_tasks:
    - name: Health check
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}/version"
        return_content: true
      register: health
      until: health.status == 200 and version in health.content
      retries: 3
      delay: 2
      delegate_to: localhost

    - name: Back into the pool
      ansible.builtin.lineinfile:
        path: /etc/nginx/conf.d/upstream.conf
        regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
        line: "    server {{ ansible_facts['default_ipv4']['address'] }}:80;"
      delegate_to: lb-1
      notify: Reload lb

  handlers:
    - name: Reload nginx
      ansible.builtin.service:
        name: nginx
        state: reloaded

    - name: Reload lb
      ansible.builtin.service:
        name: nginx
        state: reloaded
      delegate_to: lb-1
$ ansible-playbook roll.yml -e version=1.5.0-bad

PLAY [Rolling deploy] **********************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]

TASK [Drain from the pool] *****************************************************
changed: [web-1 -> lb-1]

RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-1 -> lb-1]

TASK [Deploy the new site] *****************************************************
changed: [web-1]

RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]

TASK [Health check] ************************************************************
FAILED - RETRYING: [web-1 -> localhost]: Health check (2 retries left).
FAILED - RETRYING: [web-1 -> localhost]: Health check (1 retries left).
fatal: [web-1 -> localhost]: FAILED! => {"attempts": 3, "changed": false, "connection": "close", "content": "1.4.2\n", "content_length": "6", "content_type": "text/plain", "cookies": {}, "cookies_string": "", "date": "Tue, 22 Sep 2026 20:00:13 GMT", "elapsed": 0, "msg": "OK (6 bytes)", "redirected": false, "server": "nginx/1.26.3 (Ubuntu)", "status": 200, "url": "http://web-1/version"}

NO MORE HOSTS LEFT *************************************************************

PLAY RECAP *********************************************************************
web-1                      : ok=5    changed=4    unreachable=0    failed=1    skipped=0    rescued=0    ignored=0

web-1 failed its health check, and web-2 was never touched - the bad version cost one host.

delegate_to: acting on another host

Taking a web server out of the load balancer is a change on lb-1, made while the play is working on web-1. delegate_to: runs one task on a different host, while everything else (variables, inventory_hostname) is still about the current host:

pre_tasks:
  - name: Drain {{ inventory_hostname }} from the pool
    ansible.builtin.lineinfile:
      path: /etc/nginx/conf.d/upstream.conf
      regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
      line: "    server {{ ansible_facts['default_ipv4']['address'] }}:80 down;"
    delegate_to: lb-1
    notify: Reload lb

The output shows it as changed: [web-1 -> lb-1]: the task belongs to web-1, it ran on lb-1. Two relatives:

pre_tasks, post_tasks and health checks

A play has three task lists that run in order, with handlers flushed after each: pre_tasks, then roles and tasks, then post_tasks. A rolling deploy uses them as drain / deploy / check-and-enable:

- name: Rolling deploy
  hosts: web
  serial: 1
  max_fail_percentage: 0
  become: true

  pre_tasks:
    - name: Drain from the pool            # delegate_to: lb-1, as above

  tasks:
    - name: Deploy the new site
      ansible.builtin.template: { src: site.conf.j2, dest: /etc/nginx/sites-available/default }
      notify: Reload nginx
    - name: Reload now, not at the end of the play
      ansible.builtin.meta: flush_handlers

  post_tasks:
    - name: Wait for the port
      ansible.builtin.wait_for:
        port: 80
        timeout: 30
    - name: Health check
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}/version"
        return_content: true
      register: health
      until: health.status == 200 and version in health.content
      retries: 5
      delay: 2
      delegate_to: localhost
    - name: Back into the pool
      ansible.builtin.lineinfile: { ... "down" removed ... }
      delegate_to: lb-1
      notify: Reload lb

If the health check fails, the host fails, max_fail_percentage: 0 stops the play, and the drained host is still out of the pool - the load balancer keeps sending traffic to the healthy one. That is the safety the whole pattern exists for.

$ ansible-playbook roll.yml -e version=1.5.0

PLAY [Rolling deploy] **********************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]

TASK [Drain from the pool] *****************************************************
ok: [web-1 -> lb-1]

TASK [Deploy the new site] *****************************************************
changed: [web-1]

RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]

TASK [Health check] ************************************************************
ok: [web-1 -> localhost]

TASK [Back into the pool] ******************************************************
changed: [web-1 -> lb-1]

RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-1 -> lb-1]

PLAY [Rolling deploy] **********************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-2]

TASK [Drain from the pool] *****************************************************
changed: [web-2 -> lb-1]

RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-2 -> lb-1]

TASK [Deploy the new site] *****************************************************
changed: [web-2]

RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-2]

TASK [Health check] ************************************************************
ok: [web-2 -> localhost]

TASK [Back into the pool] ******************************************************
changed: [web-2 -> lb-1]

RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-2 -> lb-1]

PLAY RECAP *********************************************************************
web-1                      : ok=7    changed=4    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=8    changed=6    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
$ curl -s lb-1/version
1.5.0

Restarting the hosts themselves

Kernel updates need reboots; the same pattern applies. ansible.builtin.reboot reboots the host, waits for SSH to come back and runs a test command:

- name: Reboot one at a time
  hosts: web
  serial: 1
  become: true
  tasks:
    - name: Reboot
      ansible.builtin.reboot:
        reboot_timeout: 300

throttle: 1 on a single task limits only that task to one host at a time, without batching the whole play.

In an interview: "serial makes batches, max_fail_percentage (or any_errors_fatal) stops the rollout at the first failure, delegate_to runs the drain and enable steps on the load balancer, and a uri check with until/retries in post_tasks decides whether the host goes back in. A bad release then costs one host, not the fleet."

What you can now do

Why it helps

By default Ansible changes every host at the same time, which turns one bad value into a full outage. Rolling changes are the pattern that limits the damage of any mistake to a single server, and it is the pattern behind most safe releases on VM fleets.

You will use it for releases, kernel patches with reboots and config changes on load-balanced tiers. It is also a favourite interview question, because it combines serial, delegate_to, handlers and health checks into one design.

Commands in this lesson

cd cat ansible-playbook curl

FAQ

What does serial do?

It splits the hosts of a play into batches and runs the whole play (pre_tasks, tasks, handlers, post_tasks) on one batch before starting the next. serial: 1 is one host at a time; serial: "25%" is a quarter of the hosts; a list like [1, 5, "50%"] starts with a single canary and grows.

How does Ansible stop a bad rollout?

max_fail_percentage sets how many failed hosts in a batch are tolerated before the play stops; 0 means any failure stops it. any_errors_fatal: true does the same for the whole play. Without either, a batch with one failed host still lets the next batch start.

What does delegate_to do?

It runs a task on another host while keeping the current host's variables. In a rolling deploy, the drain task runs on lb-1 but uses the web host's IP address, so it can take exactly that host out of the pool. delegate_to: localhost runs a health check from the control node.

How do I wait for a service to be healthy?

Use the uri module against a health or version URL, register the result, and add until: (a condition on the status and content), retries: and delay:. Ansible repeats the check until it passes or the retries run out, then fails the host. wait_for is the simpler check for "the port is open".

How do I reboot servers safely with Ansible?

The ansible.builtin.reboot module reboots the host, waits for it to come back and checks it answers, all in one task. Put it in a play with serial and a health check afterwards, and only one host is ever down at a time.

In an interview Junior

How would you deploy a change to web servers behind a load balancer without downtime?

A rolling deploy: serial: 1 so the play runs in batches of one host, and max_fail_percentage: 0 so the first failure stops it. pre_tasks drain the host from the pool, a lineinfile that marks it down, delegated to lb-1 with delegate_to (it runs on lb-1), plus a reload. tasks deploy the change and flush the handlers. post_tasks run a health check with uri and until / retries / delay, then put the host back. If the check fails, the host stays drained, the rest are never touched, and users only reach healthy servers.

Also asked: What is the difference between serial and forks? · What does run_once do, and when would you use it? · How would you roll back a bad release with Ansible?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.