The ticket: "Short 502 bursts on the shop, every 30 minutes, since Tuesday." The 502s line up with the scheduled run of the web config playbook. Nobody has changed the site config since Tuesday. Run it by hand and it agrees that the config is fine, then restarts nginx anyway:
$ ansible-playbook site.yml
PLAY [Web servers] *************************************************************
TASK [Gathering Facts] *********************************************************
ok: [web-1]
ok: [web-2]
TASK [Back up the current site config] *****************************************
changed: [web-1]
changed: [web-2]
TASK [Site config] *************************************************************
ok: [web-1]
ok: [web-2]
TASK [Check the nginx config] **************************************************
changed: [web-1]
changed: [web-2]
RUNNING HANDLER [Restart nginx] ************************************************
changed: [web-1]
changed: [web-2]
PLAY RECAP *********************************************************************
web-1 : ok=5 changed=3 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
web-2 : ok=5 changed=3 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0Site config: ok means the template module compared the rendered file with the one on disk and found them equal. And still: changed=3 and a restart. The restart is visible on the host:
$ ansible web-1 -a 'systemctl show nginx -p ActiveEnterTimestamp'
web-1 | CHANGED | rc=0 >>
ActiveEnterTimestamp=Tue 2026-09-22 19:40:04 UTC
$ ansible-playbook site.yml
...
$ ansible web-1 -a 'systemctl show nginx -p ActiveEnterTimestamp'
web-1 | CHANGED | rc=0 >>
ActiveEnterTimestamp=Tue 2026-09-22 20:00:07 UTCA restart drops connections for a moment, and the load balancer in front turns that into a few 502 Bad Gateway responses. 48 small outages a day, from a job that changes nothing.
What is happening: Ansible cannot see inside a command
Most modules are idempotent: they check the current state, change it only if needed, and report ok or changed honestly. template, copy, file, apt and service all work like that.
command and shell cannot. Ansible has no idea what cp or nginx -t did, so it assumes the worst: every command task reports changed. Even the ad-hoc check above says CHANGED for a read-only systemctl show. With -v you can see it in the result:
$ ansible-playbook site.yml -v | grep -A1 'TASK \[Check the nginx config\]'
TASK [Check the nginx config] **************************************************
changed: [web-1] => {"changed": true, "cmd": ["nginx", "-t"], "delta": "0:00:00.000317", "end": "2026-09-22 20:00:10.940580", "msg": "", "rc": 0, "start": "2026-09-22 20:00:10.940580", "stderr": "nginx: the configuration file /etc/nginx/nginx.conf syntax is ok\nnginx: configuration file /etc/nginx/nginx.conf test is successful", "stderr_lines": ["nginx: the configuration file /etc/nginx/nginx.conf syntax is ok", "nginx: configuration file /etc/nginx/nginx.conf test is successful"], "stdout": "", "stdout_lines": []}Now the second half. A handler runs when a task that notifies it reports changed. Here is the playbook:
tasks:
- name: Back up the current site config
ansible.builtin.shell: cp /etc/nginx/sites-available/default /root/default.bak
notify: Restart nginx
- name: Site config
ansible.builtin.template:
src: site.conf.j2
dest: /etc/nginx/sites-available/default
mode: "0644"
notify: Restart nginx
- name: Check the nginx config
ansible.builtin.command: nginx -t
notify: Restart nginx
handlers:
- name: Restart nginx
ansible.builtin.service:
name: nginx
state: restartedTwo always-changed tasks notify the handler, so the handler runs on every run. It runs once at the end of the play, however many tasks notified it, but once is enough to drop connections. And it restarts, where a reload would have been enough.
Diagnosis
- Run it twice in a row. The second run of a correct playbook says
changed=0. Anything that sayschangedtwice is either really changing something every time (a timestamp in a template, alatestpackage) or lying. - Look at which tasks say
changed, and at their module.command,shell,rawandscriptare the usual suspects. - Look at what they notify. A
RUNNING HANDLERon a run where the config task saidokis the symptom in one line.
--check --diff does not catch this, by the way. In check mode, command tasks are skipped, so the dry run looks clean while the real run restarts:
$ ansible-playbook site.yml --check --diff
...
TASK [Back up the current site config] *****************************************
skipping: [web-1]
skipping: [web-2]
...
PLAY RECAP *********************************************************************
web-1 : ok=2 changed=0 unreachable=0 failed=0 skipped=2 rescued=0 ignored=0
web-2 : ok=2 changed=0 unreachable=0 failed=0 skipped=2 rescued=0 ignored=0The fix
Three changes, each with a reason:
tasks:
- name: Site config
ansible.builtin.template:
src: site.conf.j2
dest: /etc/nginx/sites-available/default
mode: "0644"
backup: true
notify: Reload nginx
- name: Check the nginx config
ansible.builtin.command: nginx -t
changed_when: false
handlers:
- name: Reload nginx
ansible.builtin.service:
name: nginx
state: reloaded- The
cpbackup is gone.backup: trueon the template keeps a timestamped copy only when the file actually changes. nginx -tonly reads, so it says so withchanged_when: falseand notifies nothing. For commands that do change things, give Ansible a way to know:creates:/removes:, orchanged_whenon the output (changed_when: "'created' in result.stdout").- The handler reloads. nginx re-reads its config with no dropped connections.
Prove it with two clean runs, then one real change:
$ ansible-playbook site.yml
...
PLAY RECAP *********************************************************************
web-1 : ok=3 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
web-2 : ok=3 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
$ sed -i 's/keepalive_timeout 65/keepalive_timeout 60/' templates/site.conf.j2 && ansible-playbook site.yml
...
TASK [Site config] *************************************************************
changed: [web-1]
changed: [web-2]
TASK [Check the nginx config] **************************************************
ok: [web-1]
ok: [web-2]
RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]
changed: [web-2]Do not "fix" it by deleting the notify. Then a real config change never reaches the running nginx, and you have swapped this incident for one where the change never applies.
Keeping it from coming back
- Run twice in CI. Molecule's idempotence step does exactly this: converge, converge again, fail if anything reports
changed.ansible-lintflagscommand/shelltasks withoutchanged_when(ruleno-changed-when). - Prefer a module over a command.
ansible.builtin.copywithremote_src: trueinstead ofcp,ansible.builtin.lineinfileinstead ofsed -i. - Reload by default, restart when you must (a new listen port, a new module), and make that a separate handler.
- Validate before the reload. For a whole
nginx.conf, the template module'svalidate: nginx -t -c %schecks the new file before it replaces the old one. A site file undersites-available/is not a complete config on its own, so there keep thenginx -ttask (withchanged_when: false) ahead of the handler, so a broken file fails the run instead of the reload. - For the cases where a restart cannot be avoided, roll it:
serialplus a health check, so one server at a time leaves the pool. The same idea as a Kubernetes rolling update with readiness.
When the run does not even start, it is usually the connection: see "UNREACHABLE!" and "Missing sudo password".
Practise it
The Ansible chapter has this as Incident: every deploy run restarts nginx: two web servers, the playbook above, and you are done when two runs in a row report changed=0 and a real template change still reloads nginx. Before it, the mission Reload only when the config changed and the drill changed=0 on the second run build the habit.