OnCallReady

Lesson 33.1 · Ansible: Configuration as Code · 15 min read

Why configuration management

In plain words

Think of three ways to make sure every classroom in a school has the same posters. You could hand a teacher a list of steps ("pin poster A, then B"), which goes wrong the moment a poster is already there. You could build every classroom identical from scratch each year, which is reliable but slow. Or you could describe how the wall should look and have someone check each room and fix only the differences.

Servers have the same three options: a script of steps, a golden image built once and copied, or configuration management. Ansible is the third: you describe the desired state, it compares, and it changes only what differs. Running it twice is safe.

The problem

Picture the fleet this chapter gives you: a load balancer lb-1, two web servers web-1 and web-2, a database host db-1. All four are Ubuntu 26.04 servers you reach over SSH, exactly like the box you have been using since chapter 1.

Now someone asks for one small change: "set vm.swappiness=10 on every server and add the new on-call engineer's SSH key". You know how to do it on one machine - sysctl, an editor, authorized_keys (5.3, 1.17). Four machines means four SSH sessions, four chances to typo, and no record of what you did. Forty machines means a script. And the script is where the real trouble starts.

What you need to know already: 1.1 and 1.17 (SSH, keys, known_hosts), 1.11 (sudo, apt), 2.1 and 2.5 (systemd units and services), 6.1 (bash scripts, set -euo pipefail), 7.11 (structured data and jq).

Three ways to get a server into shape

A script (setup.sh copied over and run) says how: "append this line, create this user, restart that service". Run it twice and it appends the line twice, fails because the user exists, restarts a service that did not need it. Run it on a server that is halfway there and nobody knows what happens. Scripts are imperative: a list of steps.

A golden image bakes everything into the disk image the server boots from. Every new server is identical. But the servers already running do not change when the image does: you replace them, which is cheap for some machines and not an option for a database server with 2 TB of data.

Configuration management (Ansible, Puppet, Chef, Salt) says what: "this line must be in /etc/sysctl.conf", "user deploy must exist with this key", "nginx must be installed and running". The tool looks at the server, compares it with what you described, and changes only what differs. That is declarative: you describe the desired state, the tool works out the steps.

Later (Ch 10): container images are the golden-image idea taken all the way: the server (or the container) is replaced, never changed.

Idempotency: the property that matters

An operation is idempotent when doing it twice has the same effect as doing it once. echo "vm.swappiness=10" >> /etc/sysctl.conf is not: the second run adds a second line. "Make sure the line vm.swappiness=10 is in the file" is: the second run finds it and does nothing.

Ansible's modules are written to be idempotent, and every result says which case it was:

ok: [web-1]        the host was already in the desired state: nothing was done
changed: [web-2]   something was different and Ansible changed it

That gives you a free test that no script has: run it twice. The second run of a good playbook reports changed=0 on every host. When it does not, something in your playbook is not idempotent - and that is a bug, not a cosmetic issue, because "changed" is also what triggers restarts later in this chapter.

In an interview: "Idempotent means I can run it again safely: the second run finds everything already in the desired state, changes nothing and reports ok. That is what lets you run configuration management on a schedule, and it is why a shell: task that always reports changed is a smell."

Push and pull

Pull tools (Puppet, Chef, Salt in its usual setup) put an agent on every server. The agent wakes up every 30 minutes, asks a central server for its configuration and applies it. You need the agent installed, a server to run, and certificates between them - but drift is corrected all the time.

Push is Ansible's default: nothing runs on the servers until you run a command on the control node (the machine where Ansible is installed - here, the lab box). It connects to each managed node over plain SSH, does the work and disconnects. Ansible is agentless: the managed nodes only need two things they already have:

control node (oncall-lab)                     managed nodes
  ansible-playbook site.yml   --- ssh --->    web-1: python3 module.py -> JSON back
                              --- ssh --->    web-2: python3 module.py -> JSON back
                              --- ssh --->    db-1:  ...

The module runs on the host, does its comparison there (is nginx installed? is the line in the file?) and sends back a small JSON result: changed or not, plus details. Ansible prints it as one ok: or changed: line.

ansible-pull turns Ansible around for the cases where pull wins (thousands of laptops, servers that are not always reachable): each node clones a git repository with the playbooks and runs them against itself on a timer.

The lab fleet

The hosts exist while you work on this chapter. They are full Ubuntu servers: their own filesystem, users, systemd, packages, SSH server and IP address on 10.0.5.0/24. Everything you learnt about Linux works on them:

$ ssh web-1 'grep PRETTY /etc/os-release; hostname -I; systemctl is-active ssh'
PRETTY_NAME="Ubuntu 26.04.1 LTS"
10.0.5.11
active

The first ssh to each host asks about its host key; answer yes (1.17 explains why it asks). The box's /etc/hosts has a block with their names, so web-1 resolves:

$ grep -A6 'managed hosts' /etc/hosts
# BEGIN managed hosts (lab)
10.0.5.10       lb-1 lb-1.lab
10.0.5.11       web-1 web-1.lab
10.0.5.12       web-2 web-2.lab
10.0.5.21       db-1 db-1.lab
# END managed hosts (lab)

Ansible on Ubuntu 26.04

Ubuntu ships two packages:

You install the second one (sudo apt install -y ansible) in the first mission. Then:

$ ansible --version
ansible [core 2.20.1]
  config file = None
  configured module search path = ['/home/learner/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
  ansible python module location = /usr/lib/python3/dist-packages/ansible
  ansible collection location = /home/learner/.ansible/collections:/usr/share/ansible/collections
  executable location = /usr/bin/ansible
  python version = 3.14.4 (main, Aug 14 2026, 11:02:17) [GCC 15.2.0] (/usr/bin/python3.14)
  jinja version = 3.1.6
  pyyaml version = 6.0.2 (with libyaml v0.2.5)

Read it top to bottom: the core version, which ansible.cfg it found (none yet), where it looks for modules and collections, and the Python it runs on. When something behaves differently from a blog post, the first question is "which version?" - 2.19 rewrote templating and error messages, and most articles online predate it.

What Ansible is not

Later (Ch 12): Terraform is the provisioning tool, and the chapter's last lesson puts the tools side by side.

What you can now do

Why it helps

Every team has the server that "works, but nobody knows why", because someone fixed it by hand two years ago. This lesson gives you the words to explain why that happens (drift, imperative scripts) and what fixes it (a declarative, idempotent description that is re-run). Those words come up in every conversation about tooling, in design reviews and in interviews.

It also sets up the lab: a control node (your box) and four managed hosts reached over SSH, the same shape as a real fleet. Knowing which machine runs what saves a lot of confusion when an error appears on one host but not on another.

Commands in this lesson

ssh grep ansible

FAQ

Why not just keep using bash scripts?

A script says how, step by step, and assumes a starting point. Run it on a half-configured server, or twice, and it appends a line again, fails on a directory that exists, or restarts a service for nothing. Making every step check first is possible but it is exactly what Ansible modules already do, tested on many systems.

What is drift?

Drift is the difference that builds up between servers that should be identical: a package upgraded by hand on one, a config line edited during an incident on another. Nobody sees it until something behaves differently on one host. Re-running the same idempotent playbook regularly, and checking with --check --diff, is how Ansible finds and removes drift.

What is the difference between push and pull?

Push: the control node connects to the hosts and applies the changes, which is how ansible-playbook works. Pull: each host fetches the configuration and applies it to itself on a schedule, which ansible-pull can do and some other tools do by default. Push is simpler to start with and gives you a clear moment when a change happens.

What does the control node need?

Python, Ansible itself (apt install ansible on Ubuntu), an SSH key the managed hosts accept, and the project files: inventory, playbooks, roles and ansible.cfg. In the lab that is your box, where the hosts' keys are already trusted and your key is already on the hosts.

Is a golden image better than configuration management?

They solve different parts. A golden image is fast and identical at start, but every change means building and rolling out a new image. Configuration management changes servers in place and suits long-lived machines. Many teams use both: Ansible roles build the image, and Ansible also handles the servers that live for years.

In an interview Junior

What does idempotent mean in configuration management, and why does it matter?

Idempotent means running the same thing again on a host that is already right changes nothing. An Ansible module compares the current state with the desired state and acts only on a difference, so the first run reports changed and the second run reports changed=0. It matters because you can re-run a playbook safely at any time to remove drift, and the recap becomes a real change report: if something says changed, something was really different. A script that appends a line or restarts a service on every run is not idempotent, so it cannot be re-run safely and its output tells you nothing.

Also asked: What is the difference between declarative and imperative configuration? · What is the difference between a golden image and configuration management? · What does agentless mean for Ansible?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.