OnCallReady

Lesson 33.8 · Ansible: Configuration as Code · 23 min read

Variables, facts and precedence

In plain words

Variables are the blanks in a form. The form says "port: ___", and the blank gets filled from somewhere: a default printed on the form, a note for all web servers, a note for one particular server, or a sticky note someone slapped on at the last minute. When there are several notes for the same blank, there is a strict rule for which one wins, and the last-minute sticky note (-e on the command line) always wins.

Facts are the blanks the server fills in itself: how much memory it has, its IP address, its operating system. Ansible asks each server at the start of a play.

The problem

The same playbook has to configure web servers that are almost, but not quite, the same: production listens on port 80, the staging box on 8080; db-1 has twice the memory of the web hosts, so it gets more worker processes. You do not want one playbook per host. You want one playbook and variables: values defined in the right place, filled in per host when the play runs. The hard part is not defining them - it is knowing which definition wins when there are several.

What you need to know already: the inventory and playbook lessons, 6.12 (variables and defaults in bash, ${x:-default}), 7.11 (navigating nested data).

Using a variable

Inside a playbook, {{ name }} is replaced by the variable's value (this is Jinja2, the template language of the next lesson):

- name: Show the port
  ansible.builtin.debug:
    msg: "{{ inventory_hostname }} listens on {{ app_port }}"

Remember the YAML rule: a value that starts with {{ must be quoted. Values can be any YAML type - numbers, lists, dictionaries - and {{ var }} on its own keeps that type: "{{ app_port }}" is the number 8080, not the string "8080".

ansible.builtin.debug prints a message (msg:) or a variable (var:, no braces) and is how you look at values while writing a playbook.

Where variables come from

From the least specific to the most specific, the places you will actually use:

group_vars/all.yml          every host                       (a file per group)
group_vars/web.yml          hosts in group web
host_vars/web-2.yml         one host
vars: in the play           this play only
vars_files: in the play     a file loaded by this play
register: / set_fact:       values a task produced at run time
-e / --extra-vars           the command line

group_vars/ and host_vars/ are directories next to the inventory file (or next to the playbook). Ansible loads group_vars/<group>.yml for every group a host is in, and host_vars/<host>.yml for the host. A group's file can also be a directory (group_vars/db/vars.yml + group_vars/db/vault.yml) - you will use that for secrets.

The lab project for this lesson has all of them:

$ cd ~/oncall-lab/labs/1a-ansible/try/vars
$ head -n 20 group_vars/all.yml group_vars/web.yml host_vars/web-2.yml
==> group_vars/all.yml <==
# every host
app_port: 8080
app_env: prod

==> group_vars/web.yml <==
# the web group
app_port: 80

==> host_vars/web-2.yml <==
# just web-2
app_port: 8081
$ cat show.yml
- name: Which port?
  hosts: web
  gather_facts: false
  tasks:
    - name: Show the port
      ansible.builtin.debug:
        msg: "{{ inventory_hostname }} listens on {{ app_port }} ({{ app_env }})"
$ ansible-playbook show.yml

PLAY [Which port?] *************************************************************

TASK [Show the port] ***********************************************************
ok: [web-1] => {
    "msg": "web-1 listens on 80 (prod)"
}
ok: [web-2] => {
    "msg": "web-2 listens on 8081 (prod)"
}

PLAY RECAP *********************************************************************
web-1                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

Precedence: who wins

Ansible's documentation lists 22 levels. You need the shape, from lowest to highest:

 1  role defaults                    (roles/x/defaults/main.yml: meant to be overridden)
 2  inventory group vars             ([group:vars], group_vars/all, then child groups)
 3  inventory host vars              (host line, host_vars/<host>)
 4  facts                            (gathered from the host)
 5  play vars, vars_files
 6  role vars                        (roles/x/vars/main.yml)
 7  block vars, task vars
 8  set_fact, registered vars
 9  role and include parameters
10  extra vars (-e)                  ALWAYS win

Three rules cover almost every real case:

Watch the precedence in action. web-2 has its own value in host_vars; -e overrides every host:

$ ansible-playbook show.yml -e app_port=9000

PLAY [Which port?] *************************************************************

TASK [Show the port] ***********************************************************
ok: [web-1] => {
    "msg": "web-1 listens on 9000 (prod)"
}
ok: [web-2] => {
    "msg": "web-2 listens on 9000 (prod)"
}

PLAY RECAP *********************************************************************
web-1                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

When two values fight and you cannot tell why, ansible-inventory --host web-2 shows the inventory's view, and debug: var=app_port inside the play shows the final one.

In an interview: "Variables merge by precedence: role defaults are the weakest, then inventory group vars, host vars, play vars, role vars, task vars, set_fact and registered vars, and extra vars with -e always win. In practice: defaults in the role, environment differences in group_vars, exceptions in host_vars, and -e only for one-offs."

Facts: what the host tells you

Before the first task, the Gathering Facts step runs the setup module on every host and stores what it finds in ansible_facts: the hostname, the OS and version, CPUs, memory, disks, IP addresses, the Python. You can ask for them ad-hoc, with filter= to keep it short:

$ ansible web-1 -m ansible.builtin.setup -a 'filter=ansible_processor_vcpus'
web-1 | SUCCESS => {
    "ansible_facts": {
        "ansible_processor_vcpus": 1
    },
    "changed": false
}
$ ansible db-1 -m ansible.builtin.setup -a 'filter=ansible_memtotal_mb'
db-1 | SUCCESS => {
    "ansible_facts": {
        "ansible_memtotal_mb": 3911
    },
    "changed": false
}

In a playbook you read them as ansible_facts['memtotal_mb'] (no ansible_ prefix inside the dictionary). The ones you will use most:

ansible_facts['hostname']                     web-1
ansible_facts['distribution'], ['distribution_version']   Ubuntu, 26.04
ansible_facts['os_family']                    Debian (the family, for apt vs dnf choices)
ansible_facts['processor_vcpus']              1
ansible_facts['memtotal_mb']                  1966
ansible_facts['default_ipv4']['address']      10.0.5.11

Older playbooks use the same facts as top-level variables: ansible_distribution, ansible_memtotal_mb. That still works in ansible-core 2.20, with a deprecation warning, and stops working in 2.24. Write ansible_facts[...] in new code:

$ ansible-playbook facts.yml

PLAY [Facts] *******************************************************************

TASK [Gathering Facts] *********************************************************
ok: [web-1]
ok: [db-1]

TASK [Size of each host] *******************************************************
ok: [web-1] => {
    "msg": "web-1: 1 vCPU, 1966 MB"
}
ok: [db-1] => {
    "msg": "db-1: 2 vCPU, 3911 MB"
}

TASK [The old way (deprecated in 2.20)] ****************************************
[DEPRECATION WARNING]: INJECT_FACTS_AS_VARS default to `True` is deprecated, top-level facts will not be auto injected after the change. This feature will be removed from ansible-core version 2.24.
Origin: facts.yml:8:7

6         msg: "{{ ansible_facts['hostname'] }}: {{ ansible_facts['processor_vcpus'] }} vCPU, {{ ansible_facts['memtotal_mb'] }} MB"
7
8     - name: The old way (deprecated in 2.20)
        ^ column 7

Use `ansible_facts["fact_name"]` (no `ansible_` prefix) instead.

ok: [web-1] => {
    "msg": "Ubuntu 26.04"
}
ok: [db-1] => {
    "msg": "Ubuntu 26.04"
}

PLAY RECAP *********************************************************************
db-1                       : ok=3    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-1                      : ok=3    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0

Magic variables are filled in by Ansible itself and cannot be set:

inventory_hostname          the host's name in the inventory (web-1)
group_names                 the groups this host is in: ['prod', 'web']
groups                      every group and its hosts: groups['web'] = ['web-1', 'web-2']
hostvars                    every host's variables: hostvars['db-1']['ansible_facts']...
ansible_play_hosts          the hosts still running this play

hostvars is how one host's config can mention another's address - the load balancer listing its web servers, say.

register, when and loop

A task's result can be kept in a variable with register. It is a dictionary: for a command, rc, stdout, stdout_lines, stderr; every result has changed and failed.

- name: Read the kernel version
  ansible.builtin.command: uname -r
  register: kernel
  changed_when: false

- name: Only on the new kernel
  ansible.builtin.debug:
    msg: "running {{ kernel.stdout }}"
  when: kernel.stdout is search('^7\.')

when takes a condition - an expression, not in {{ }} (it is already one). Since ansible-core 2.19 it must evaluate to a real boolean: when: kernel.stdout (a string) is an error, when: kernel.stdout | length > 0 is right. Tests make the common checks readable: is defined, is search('x'), result is changed, result is failed.

loop runs a task once per item, with the item in item:

- name: Base packages
  ansible.builtin.apt:
    name: "{{ item }}"
  loop: [htop, tree, jq]

The output has one line per item - changed: [web-1] => (item=tree) - and a registered loop result has a results list, one entry per item. loop_control: {label: "{{ item.name }}"} keeps the output short when items are big dictionaries. (Many modules, apt included, also take a list directly, which is faster than a loop.)

set_fact creates a variable from an expression, per host, for the rest of the run:

- name: Work out the worker count
  ansible.builtin.set_fact:
    workers: "{{ [ansible_facts['processor_vcpus'] * 2, 8] | min }}"

What you can now do

Why it helps

"Why does web-1 have the wrong port?" is one of the most common Ansible questions in real teams, and the answer is always precedence: a value set in a stronger place than you thought. Knowing the ladder and how to see the value Ansible actually uses for a host turns an hour of guessing into a minute.

Facts make playbooks adapt to the host instead of hard-coding values: a worker count from the CPU count, a memory limit from the RAM, a config that lists every web server's IP. register, when and loop are how a playbook reacts to what it finds.

Commands in this lesson

cd head cat ansible-playbook ansible

FAQ

group_vars or the inventory file?

Both set group variables. Files under group_vars/ and host_vars/ next to the inventory or the playbook are easier to read, review and diff than long variable sections in an INI file, and they beat the inventory file's own group variables. Most teams keep the inventory file for host names and groups only.

What are facts and how do I see them?

Facts are what Ansible learns about a host when a play gathers them: ansible_facts['memtotal_mb'], ansible_facts['processor_vcpus'], ansible_facts['default_ipv4']['address'] and many more. ansible web-1 -m ansible.builtin.setup prints them all; add -a 'filter=ansible_mem*' to narrow it down. In a playbook, gather_facts: true (the default) collects them at the start of every play.

Why use ansible_facts['x'] rather than ansible_x?

Ansible 2.20 still injects facts as top-level variables such as ansible_hostname, but warns that this is deprecated and will stop. Reading them through ansible_facts works now and later. The lab's examples use the new form, so the deprecation warnings never appear in your runs.

What does register give me?

It saves a task's result in a variable: rc, stdout, stdout_lines, changed, failed and module-specific fields. A later task can test it with when: (result.rc != 0, 'error' in result.stdout) or print it with debug. Use debug: var=result once to see what a module returns before you write conditions on it.

How do loops work?

loop: takes a list and runs the task once per item, with the item in {{ item }}. With a list of dictionaries you use item.name, item.shell and so on. loop_control: label: keeps the output short. Many modules also accept a list directly (apt name: [a, b]), which is faster than a loop.

In an interview Junior

Explain Ansible variable precedence in a few sentences.

Variables can come from many places and Ansible merges them by precedence. Weakest first: role defaults, then inventory variables (group, then host), group_vars and host_vars files, then facts, then play vars and vars_files, role vars, set_fact and registered results, and finally extra vars. Two rules cover most cases: The more specific place wins (a host beats its group, a child group beats its parent), and The playbook beats the inventory. And -e beats everything. To see the value Ansible really uses for one host I run ansible web-1 -m ansible.builtin.debug -a var=app_port.

Also asked: Where would you put a value that should be easy to override? · What are facts, and how do you use one in a task? · What is the difference between register and set_fact?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.