The problem
Two web servers sit behind lb-1. A new release needs a new config and a restart. Ansible's default is to do every task on every host at the same time (5 in parallel): both web servers restart together, and for those seconds the load balancer has nothing to send traffic to. With a broken config, both stay down. A rolling change does a few hosts at a time, takes each one out of the load balancer first, checks it is healthy before putting it back, and stops the rollout at the first failure - so a bad release costs you one host, not the fleet.
What you need to know already: the handlers lesson (flush_handlers), the templates lesson, 2.10 (reload vs restart), 1.13 (addresses and ports).
serial: batches
serial: on a play splits its hosts into batches; the whole play (all tasks, handlers included) runs on one batch before the next starts:
- name: Rolling deploy
hosts: web
serial: 1 # one host at a time; also: 2, "25%", or a ramp [1, 5, "50%"]
serial: [1, 5, "50%"] is a canary pattern: one host first, then five, then half the rest at a time. Each batch prints its own PLAY [...] banner.
Stopping a bad rollout
By default Ansible keeps going while at least one host in the batch survives. For a rollout that is the wrong instinct: if the new version fails on the first host, you want to stop there.
serial: 1
max_fail_percentage: 0 # any failure in a batch stops the whole play
max_fail_percentage: N aborts the play when more than N% of the current batch failed (Ansible prints NO MORE HOSTS LEFT and the batches after it never start). With serial and no max_fail_percentage, a batch where every host failed also stops the play. any_errors_fatal: true is the bigger hammer: the first failure on any host ends the play for all hosts, even in the middle of a batch.
$ cd ~/oncall-lab/labs/1a-ansible/try/rolling
$ cat roll.yml
- name: Rolling deploy
hosts: web
serial: 1
max_fail_percentage: 0
become: true
pre_tasks:
- name: Drain from the pool
ansible.builtin.lineinfile:
path: /etc/nginx/conf.d/upstream.conf
regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
line: " server {{ ansible_facts['default_ipv4']['address'] }}:80 down;"
delegate_to: lb-1
notify: Reload lb
- name: Apply the drain now
ansible.builtin.meta: flush_handlers
tasks:
- name: Deploy the new site
ansible.builtin.template:
src: site.conf.j2
dest: /etc/nginx/sites-available/default
mode: "0644"
notify: Reload nginx
- name: Reload now, not at the end of the play
ansible.builtin.meta: flush_handlers
post_tasks:
- name: Health check
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}/version"
return_content: true
register: health
until: health.status == 200 and version in health.content
retries: 3
delay: 2
delegate_to: localhost
- name: Back into the pool
ansible.builtin.lineinfile:
path: /etc/nginx/conf.d/upstream.conf
regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
line: " server {{ ansible_facts['default_ipv4']['address'] }}:80;"
delegate_to: lb-1
notify: Reload lb
handlers:
- name: Reload nginx
ansible.builtin.service:
name: nginx
state: reloaded
- name: Reload lb
ansible.builtin.service:
name: nginx
state: reloaded
delegate_to: lb-1
$ ansible-playbook roll.yml -e version=1.5.0-bad
PLAY [Rolling deploy] **********************************************************
TASK [Gathering Facts] *********************************************************
ok: [web-1]
TASK [Drain from the pool] *****************************************************
changed: [web-1 -> lb-1]
RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-1 -> lb-1]
TASK [Deploy the new site] *****************************************************
changed: [web-1]
RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]
TASK [Health check] ************************************************************
FAILED - RETRYING: [web-1 -> localhost]: Health check (2 retries left).
FAILED - RETRYING: [web-1 -> localhost]: Health check (1 retries left).
fatal: [web-1 -> localhost]: FAILED! => {"attempts": 3, "changed": false, "connection": "close", "content": "1.4.2\n", "content_length": "6", "content_type": "text/plain", "cookies": {}, "cookies_string": "", "date": "Tue, 22 Sep 2026 20:00:13 GMT", "elapsed": 0, "msg": "OK (6 bytes)", "redirected": false, "server": "nginx/1.26.3 (Ubuntu)", "status": 200, "url": "http://web-1/version"}
NO MORE HOSTS LEFT *************************************************************
PLAY RECAP *********************************************************************
web-1 : ok=5 changed=4 unreachable=0 failed=1 skipped=0 rescued=0 ignored=0
web-1 failed its health check, and web-2 was never touched - the bad version cost one host.
delegate_to: acting on another host
Taking a web server out of the load balancer is a change on lb-1, made while the play is working on web-1. delegate_to: runs one task on a different host, while everything else (variables, inventory_hostname) is still about the current host:
pre_tasks:
- name: Drain {{ inventory_hostname }} from the pool
ansible.builtin.lineinfile:
path: /etc/nginx/conf.d/upstream.conf
regexp: '^\s*server {{ ansible_facts["default_ipv4"]["address"] }}:80'
line: " server {{ ansible_facts['default_ipv4']['address'] }}:80 down;"
delegate_to: lb-1
notify: Reload lb
The output shows it as changed: [web-1 -> lb-1]: the task belongs to web-1, it ran on lb-1. Two relatives:
delegate_to: localhostruns on the control node - the usual way to call an API, or to check a URL from outside.run_once: trueruns a task for the first host of the batch only, and gives the result to all of them. Withdelegate_toit is "do this once, there": a database migration, a message to the team chat.
pre_tasks, post_tasks and health checks
A play has three task lists that run in order, with handlers flushed after each: pre_tasks, then roles and tasks, then post_tasks. A rolling deploy uses them as drain / deploy / check-and-enable:
- name: Rolling deploy
hosts: web
serial: 1
max_fail_percentage: 0
become: true
pre_tasks:
- name: Drain from the pool # delegate_to: lb-1, as above
tasks:
- name: Deploy the new site
ansible.builtin.template: { src: site.conf.j2, dest: /etc/nginx/sites-available/default }
notify: Reload nginx
- name: Reload now, not at the end of the play
ansible.builtin.meta: flush_handlers
post_tasks:
- name: Wait for the port
ansible.builtin.wait_for:
port: 80
timeout: 30
- name: Health check
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}/version"
return_content: true
register: health
until: health.status == 200 and version in health.content
retries: 5
delay: 2
delegate_to: localhost
- name: Back into the pool
ansible.builtin.lineinfile: { ... "down" removed ... }
delegate_to: lb-1
notify: Reload lb
- wait_for waits for a port to open (or close:
state: stopped), or a file to appear. It fails aftertimeoutseconds:Timeout when waiting for 10.0.5.11:80. - uri makes an HTTP request and fails unless the status is in
status_code(default[200]).return_content: truekeeps the body. - until / retries / delay retries a task until a condition holds: here, until the host answers 200 with the new version. Each retry prints
FAILED - RETRYING: [web-1]: Health check (4 retries left).
If the health check fails, the host fails, max_fail_percentage: 0 stops the play, and the drained host is still out of the pool - the load balancer keeps sending traffic to the healthy one. That is the safety the whole pattern exists for.
$ ansible-playbook roll.yml -e version=1.5.0
PLAY [Rolling deploy] **********************************************************
TASK [Gathering Facts] *********************************************************
ok: [web-1]
TASK [Drain from the pool] *****************************************************
ok: [web-1 -> lb-1]
TASK [Deploy the new site] *****************************************************
changed: [web-1]
RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]
TASK [Health check] ************************************************************
ok: [web-1 -> localhost]
TASK [Back into the pool] ******************************************************
changed: [web-1 -> lb-1]
RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-1 -> lb-1]
PLAY [Rolling deploy] **********************************************************
TASK [Gathering Facts] *********************************************************
ok: [web-2]
TASK [Drain from the pool] *****************************************************
changed: [web-2 -> lb-1]
RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-2 -> lb-1]
TASK [Deploy the new site] *****************************************************
changed: [web-2]
RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-2]
TASK [Health check] ************************************************************
ok: [web-2 -> localhost]
TASK [Back into the pool] ******************************************************
changed: [web-2 -> lb-1]
RUNNING HANDLER [Reload lb] ****************************************************
changed: [web-2 -> lb-1]
PLAY RECAP *********************************************************************
web-1 : ok=7 changed=4 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
web-2 : ok=8 changed=6 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
$ curl -s lb-1/version
1.5.0
Restarting the hosts themselves
Kernel updates need reboots; the same pattern applies. ansible.builtin.reboot reboots the host, waits for SSH to come back and runs a test command:
- name: Reboot one at a time
hosts: web
serial: 1
become: true
tasks:
- name: Reboot
ansible.builtin.reboot:
reboot_timeout: 300
throttle: 1 on a single task limits only that task to one host at a time, without batching the whole play.
In an interview: "serial makes batches, max_fail_percentage (or any_errors_fatal) stops the rollout at the first failure, delegate_to runs the drain and enable steps on the load balancer, and a uri check with until/retries in post_tasks decides whether the host goes back in. A bad release then costs one host, not the fleet."
What you can now do
- Batch a play with
serial(numbers, percentages, ramps) and stop it withmax_fail_percentageorany_errors_fatal. - Run tasks on another host with
delegate_to, once withrun_once, on the control node withdelegate_to: localhost. - Build drain / deploy / check / enable with
pre_tasks,flush_handlers,wait_for,uri+until, andpost_tasks. - Reboot hosts one at a time with the
rebootmodule.