OnCallReady

Lesson 33.14 · Ansible: Configuration as Code · 17 min read

Handlers, tags, check and diff

In plain words

A handler is like a note on the fridge: "if anyone changes the oven settings, restart the oven before dinner". Ten people may change a setting during the afternoon, but the oven is restarted once, at the end, and only if someone actually changed something. If nobody touched the settings, nobody restarts anything.

Tags are labels on the steps of a recipe, so you can say "only do the config steps today". Check mode is a rehearsal: Ansible walks through the recipe and says what it would change, without changing anything, and diff mode shows the exact lines.

The problem

You change nginx's config. nginx has to reload to pick it up (2.10: a daemon reads its config at start or on reload). But you do not want a reload on every run - a reload or restart on every run means every run is a small outage, and "every run" might be every 30 minutes from a scheduler. You want: reload if, and only if, the config changed. Ansible's answer is handlers. The same "only when it matters" idea runs through the other tools of this lesson: tags (run only part of a playbook) and check mode (run without changing anything).

What you need to know already: the playbook and template lessons (changed vs ok, template), 2.10 (reload vs restart), 2.9 (reading systemctl status).

Handlers

A handler is a task that only runs when another task notifies it, and only if that task reported changed:

- name: Web servers
  hosts: web
  become: true
  tasks:
    - name: Site config
      ansible.builtin.template:
        src: site.conf.j2
        dest: /etc/nginx/sites-available/default
        validate: nginx -t -c /etc/nginx/nginx.conf
      notify: Reload nginx              # the handler's name, exactly

  handlers:
    - name: Reload nginx
      ansible.builtin.service:
        name: nginx
        state: reloaded

The rules:

The output marks them RUNNING HANDLER:

$ cd ~/oncall-lab/labs/1a-ansible/try/handlers
$ ansible-playbook web.yml

PLAY [Web servers] *************************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]
ok: [web-2]

TASK [Install nginx] ***********************************************************
changed: [web-1]
changed: [web-2]

TASK [Site config] *************************************************************
changed: [web-1]
changed: [web-2]

TASK [nginx runs] **************************************************************
ok: [web-1]
ok: [web-2]

RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]
changed: [web-2]

PLAY RECAP *********************************************************************
web-1                      : ok=5    changed=3    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=5    changed=3    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

When a handler does not run

The trap that bites every team once: a task changed the config and notified the handler, then a later task failed on that host. A failed host stops - and its pending handlers are dropped. The config file is new, nginx still runs the old one. Worse, on the next run the template task is ok (the file is already right), so nothing notifies the handler ever again. The change is stuck.

Three tools:

An incident later in the chapter is exactly this.

changed_when and failed_when

Every task can decide for itself what "changed" and "failed" mean, from its registered result:

- name: Apply the migrations
  ansible.builtin.command: /opt/app/bin/migrate
  register: mig
  changed_when: "'Applied' in mig.stdout"      # changed only if it applied something
  failed_when: mig.rc not in [0, 3]            # rc 3 = "nothing to do" for this tool
  notify: Restart app

changed_when: false is the standard line for commands that only read (nginx -t, cat, systemctl is-active). Without it, every run reports changed - and if such a task notifies a handler, every run restarts the service.

Tags

Tags label tasks so you can run part of a playbook:

- name: Install nginx
  ansible.builtin.apt:
    name: nginx
  tags: [packages]

- name: Site config
  ansible.builtin.template: { src: site.conf.j2, dest: /etc/nginx/sites-available/default }
  tags: [config]
--tags config               only tasks tagged config (plus those tagged always)
--skip-tags packages        everything except packages
--list-tags                 what tags exist
--tags never,debug          tasks tagged never only run when asked for by name

Two special tags: always runs unless skipped by name (good for gathering facts or a sanity check), never runs only when its other tag is asked for (good for a dangerous "reset" task). Tags on a play, a block or a role apply to everything inside.

$ ansible-playbook web.yml --list-tags

playbook: web.yml

  play #1 (web): Web servers	TAGS: []
      TASK TAGS: [always, config, packages]
$ ansible-playbook web.yml --tags config

PLAY [Web servers] *************************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]
ok: [web-2]

TASK [Site config] *************************************************************
ok: [web-1]
ok: [web-2]

TASK [nginx runs] **************************************************************
ok: [web-1]
ok: [web-2]

PLAY RECAP *********************************************************************
web-1                      : ok=3    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=3    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

Check mode and diff mode

Check mode (--check, -C) runs every task in "what would you do?" mode: modules compare the desired state with the host and report changed without changing anything. Diff mode (--diff, -D) prints what changed (or would change) in files. Together they are the review step before a risky change:

$ sed -i 's/keepalive_timeout 65/keepalive_timeout 30/' templates/site.conf.j2
$ ansible-playbook web.yml --check --diff -l web-1

PLAY [Web servers] *************************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]

TASK [Install nginx] ***********************************************************
ok: [web-1]

TASK [Site config] *************************************************************
changed: [web-1]
--- before: /etc/nginx/sites-available/default
+++ after: /home/learner/oncall-lab/labs/1a-ansible/try/handlers/templates/site.conf.j2
@@ -4,7 +4,7 @@
     server_name web-1;
     root /var/www/html;
     index index.html;
-    keepalive_timeout 65;
+    keepalive_timeout 30;

     location / {
         try_files $uri $uri/ =404;


TASK [nginx runs] **************************************************************
ok: [web-1]

RUNNING HANDLER [Reload nginx] *************************************************
changed: [web-1]

PLAY RECAP *********************************************************************
web-1                      : ok=5    changed=2    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

The diff shows the exact lines, the task says changed, the handler shows as notified

Two more ways to run part of a playbook, for long ones:

--start-at-task "Site config"   skip everything before that task
--step                          ask before every task: (N)o/(y)es/(c)ontinue

What you can now do

Why it helps

Restarting services for no reason is how configuration runs cause outages: a restart on every run means dropped connections on every run. Handlers make restarts happen only on a real change, once. The other half of the lesson is when a handler does not run: the classic "the config is on disk but the service still runs the old one" incident.

--check --diff is how changes get reviewed before they run, and tags let you run the one part of a large playbook you need during an incident, without touching packages or users.

Commands in this lesson

cd ansible-playbook sed

FAQ

When do handlers run?

At the end of the play, after all tasks, once per handler, and only if a task that notified it reported changed on that host. You can make them run earlier with ansible.builtin.meta: flush_handlers, for example before a health check that needs the new config already loaded.

Why did my handler not run?

Usually because the notifying task reported ok, not changed, which is correct. The surprising case is a later task failing on that host: the play stops for the host and its pending handlers are dropped. --force-handlers or force_handlers: true makes them run anyway; otherwise the next run will not notify them again.

What does changed_when do?

It overrides the changed status of a task with a condition. changed_when: false makes a read-only command (a version check, nginx -t) report ok, so it never triggers handlers or noise. A condition such as changed_when: "'created' in result.stdout" makes a command report changed only when it really changed something.

What do the special tags always and never mean?

A task tagged always runs even when you select other tags with --tags, unless you skip it explicitly. A task tagged never runs only when you ask for it by one of its tags. They are useful for a sanity check that should always run, and for a dangerous task that must be requested on purpose.

Can I trust --check completely?

No. command and shell are skipped in check mode, so anything that depends on their results is a guess, and tasks that depend on an earlier change can fail because that change did not happen. Treat --check --diff as a good preview of file and package changes, not as proof that the real run will succeed.

In an interview Junior

How do you restart a service only when its configuration changed?

The config task (template or copy) notifies a handler: notify: Reload nginx, and under handlers: a task with that name that uses the service module with state: reloaded. The handler runs only if the task reported changed, only once even if several tasks notified it, and at the end of the play (or earlier at meta: flush_handlers). If a later task failed on that host, its pending handlers are dropped, so the new config sits on disk unapplied; --force-handlers prevents that. I also prefer a reload over a restart, and use changed_when: false on read-only commands so they never trigger handlers by accident.

Also asked: What does --check --diff show, and what can it miss? · How would you run only the config part of a big playbook? · What is the difference between notify and listen?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.