OnCallReady

Lesson 33.30 · Ansible: Configuration as Code · 21 min read

Debugging and linting

In plain words

When a recipe goes wrong, a good cook reads the note the kitchen left: which step, which ingredient, what happened. Ansible's errors are the same kind of note: what went wrong, and an Origin line that points at the file, line and column, with a little arrow under the exact spot.

Some problems happen before cooking even starts: Ansible cannot get into the kitchen at all (UNREACHABLE), or it gets in but is not allowed to use the oven (a sudo problem). ansible-lint is a strict friend who reads your recipe before you cook and points out the mistakes you will regret later.

The problem

A playbook that worked yesterday fails today on one host, at 2 a.m., in a job log. The output is dense - JSON on one line, a host that is "UNREACHABLE", a task that failed with a message about something you never touched. Debugging Ansible is mostly reading: knowing which part of the output answers which question, which verbosity level shows what, and which of a handful of usual suspects it is. The second half of this lesson is the tool that catches many of those mistakes before they run: ansible-lint.

What you need to know already: the whole chapter so far, 1.17 (ssh -v, known_hosts), 1.11 (sudo), 6.22 (shellcheck - ansible-lint is the same idea).

Reading a failure

fatal: [web-1]: FAILED! => {"changed": false, "msg": "No package matching 'ngnix' is available"}

fatal = the host stopped here. The JSON is the module's result: read msg first. For command / shell the useful keys are rc, stderr and cmd. A failure that says Task failed: Finalization of task args ... is not the module - it is a template in the task's arguments that could not be rendered (usually 'x' is undefined), and the Origin: lines point at the exact spot.

Verbosity adds detail in layers:

-v      every task's result as JSON, on one line
-vv     the config file in use, the task's file:line ("task path:")
-vvv    the SSH connection commands, the module file sent, the result indented,
        and the module's arguments ("invocation") - secrets included, see no_log
-vvvv   ssh -vvv: key exchange, which keys were offered, why authentication failed

Start with -v. Go to -vvv for "is it even running the command I think", -vvvv for SSH problems.

UNREACHABLE: the usual suspects

UNREACHABLE means Ansible could not run anything at all - the problem is the connection, before any module. The message is SSH's own error after Failed to connect to the host via ssh:, and there are only a few of them:

$ cd ~/oncall-lab/labs/1a-ansible/try/debug
$ cat inventory.ini
[web]
web-1
web-3

[lb]
lb-1 ansible_host=10.0.5.99

[db]
db-1 ansible_user=deploy
$ ansible all -m ping
web-1 | SUCCESS => {
    "ansible_facts": {
        "discovered_interpreter_python": "/usr/bin/python3.14"
    },
    "changed": false,
    "ping": "pong"
}
web-3 | UNREACHABLE! => {
    "changed": false,
    "msg": "Failed to connect to the host via ssh: ssh: Could not resolve hostname web-3: Temporary failure in name resolution",
    "unreachable": true
}
db-1 | UNREACHABLE! => {
    "changed": false,
    "msg": "Failed to connect to the host via ssh: deploy@db-1: Permission denied (publickey).",
    "unreachable": true
}
lb-1 | UNREACHABLE! => {
    "changed": false,
    "msg": "Failed to connect to the host via ssh: ssh: connect to host 10.0.5.99 port 22: Connection timed out",
    "unreachable": true
}
Could not resolve hostname web-3: Temporary failure in name resolution
        DNS / /etc/hosts: the name does not exist (a typo, or a host not built yet)
connect to host 10.0.5.99 port 22: Connection timed out
        nothing answers: wrong IP, host down, firewall dropping
connect to host web-1 port 22: Connection refused
        the host is up, nothing listens on that port: sshd down or ansible_port wrong
learner@db-1: Permission denied (publickey).
        connected, but the key was refused: wrong user, key not in authorized_keys,
        bad permissions on ~/.ssh on the host (4.7)
Host key verification failed.
        unknown host key and nobody to answer the prompt, or
        "REMOTE HOST IDENTIFICATION HAS CHANGED" - the host was rebuilt (ssh-keygen -R)

The checks, in order: ansible-inventory --host NAME (what address and user does Ansible use?), then the same with plain SSH - ssh -v user@address. If plain SSH fails, Ansible is not the problem.

become problems

Connected fine, then:

fatal: [db-1]: FAILED! => {"msg": "Missing sudo password"}

The task has become: true, the user's sudo needs a password, and Ansible was not given one. Either pass it with -K (--ask-become-pass, prompts BECOME password:), or set up a NOPASSWD sudoers rule for the automation user - the usual choice for a dedicated deploy account. Incorrect sudo password means -K got the wrong one.

Making a playbook explain itself

- name: Show what we have
  ansible.builtin.debug:
    var: hostvars[inventory_hostname]['ansible_facts']['default_ipv4']
    verbosity: 1                       # only printed with -v or more

- name: Refuse to run with nonsense input
  ansible.builtin.assert:
    that:
      - app_port | int > 1024
      - app_env in ['dev', 'staging', 'prod']
    fail_msg: "app_port={{ app_port }} app_env={{ app_env }} is not a valid combination"

- name: Stop here on purpose
  ansible.builtin.fail:
    msg: "db-1 is in maintenance"
  when: inventory_hostname in groups['maintenance'] | default([])

assert turns assumptions into checks with a readable message - put them at the top of a role to catch bad variables before anything changes. Every that: expression must be a real boolean in 2.19+.

After fixing a problem halfway through a long playbook, --start-at-task "Task name" resumes there, and --limit @site.retry-style reruns are replaced by -l with the failed hosts. --step asks before each task, for the careful walk through.

ansible-lint

ansible-lint reads playbooks and roles (without running them) and reports patterns that are known to cause trouble. Every finding has a rule id:

$ ansible-lint messy.yml
WARNING  Listing 14 violation(s) that are fatal
name[play]: All plays should be named.
messy.yml:1 Play: web

yaml[truthy]: Truthy value should be one of [false, true]
messy.yml:2

name[missing]: All tasks should be named.
messy.yml:4 Task/Handler: shell apt-get install -y nginx

fqcn[action-core]: Use FQCN for builtin module actions (shell).
messy.yml:4 Use `ansible.builtin.shell` or `ansible.legacy.shell` instead.

no-free-form: Avoid using free-form when calling module actions. (shell)
messy.yml:4 Task/Handler: shell apt-get install -y nginx

no-changed-when: Commands should not change things if nothing needs doing.
messy.yml:4 Task/Handler: shell apt-get install -y nginx

command-instead-of-module: apt-get used in place of apt module
messy.yml:4 Task/Handler: shell apt-get install -y nginx

command-instead-of-shell: Use shell only when shell functionality is required.
messy.yml:4 Task/Handler: shell apt-get install -y nginx

name[casing]: All names should start with an uppercase letter.
messy.yml:5 Task/Handler: copy the index page

fqcn[action-core]: Use FQCN for builtin module actions (copy).
messy.yml:5 Use `ansible.builtin.copy` or `ansible.legacy.copy` instead.

no-free-form: Avoid using free-form when calling module actions. (copy)
messy.yml:5 Task/Handler: copy the index page

risky-file-permissions: File permissions unset or incorrect.
messy.yml:5 Task/Handler: copy the index page

literal-compare: Don't compare to literal True/False.
messy.yml:7 Task/Handler: Restart nginx

no-handler: Tasks that run when changed should likely be handlers.
messy.yml:7 Task/Handler: Restart nginx

Read documentation for instructions on how to ignore specific rule violations.

# Rule Violation Summary

  1 command-instead-of-module profile:basic tags:command-shell,idiom
  1 command-instead-of-shell profile:basic tags:command-shell,idiom
  1 literal-compare profile:basic tags:idiom
  1 name[casing] profile:basic tags:idiom
  1 name[missing] profile:basic tags:idiom
  1 name[play] profile:basic tags:idiom
  2 no-free-form profile:basic tags:syntax,risk
  1 yaml[truthy] profile:basic tags:formatting,yaml
  1 risky-file-permissions profile:safety tags:unpredictability
  1 no-changed-when profile:shared tags:command-shell,idempotency
  1 no-handler profile:shared tags:idiom
  2 fqcn[action-core] profile:production tags:formatting

Failed: 14 failure(s), 0 warning(s) in 1 files processed of 1 encountered. Last profile that met the validation criteria was 'min'.

Each finding: rule-id: message, then file:line and the task. The ones you will meet first:

name[missing]              a task without a name
name[casing]               a name that does not start with a capital letter
fqcn[action-core]          apt instead of ansible.builtin.apt
no-changed-when            command/shell without changed_when (or creates/removes)
command-instead-of-module  shell: apt-get install ... instead of the apt module
command-instead-of-shell   shell: where command would do (no pipes, no globs)
risky-shell-pipe           a pipe in shell without set -o pipefail (6.3)
risky-file-permissions     copy/template/file without mode:
risky-octal                mode: 644 (decimal) instead of "0644"
no-handler                 when: result.changed - should be a handler
literal-compare            when: x == True  -  write when: x
yaml[truthy]               yes/no instead of true/false
var-naming[no-role-prefix] a role variable without the role's name as a prefix

Rules belong to profiles that build on each other: min (it parses) < basic < moderate < safety < shared < production. The last line says which one the code meets; --profile production makes the strictest one the goal. Teams run it in the job that tests the playbooks (exit code 2 = violations), with a .ansible-lint file:

# .ansible-lint
profile: production
skip_list:
  - yaml[line-length]
warn_list:
  - experimental

For one deliberate exception, a comment on the line: # noqa: no-changed-when. Fix before you silence: most findings are real bugs waiting for the right day.

What you can now do

Why it helps

Most of an Ansible engineer's time goes into reading failures. Knowing where to look first (the message, then the Origin line), how to tell a connection problem from a task failure, and how to make a playbook explain itself with -v, debug and assert turns long sessions into short ones.

ansible-lint is what teams run before every merge. Its rule ids are worth knowing: they point at real bugs, not only style, such as unquoted modes, latest package versions and handlers written as when: changed tasks.

Commands in this lesson

cd cat ansible ansible-lint

FAQ

What is the difference between UNREACHABLE and FAILED?

UNREACHABLE means Ansible never got a working SSH session, so no module ran: host key problems, a wrong address or user, a refused key. FAILED means it connected and the task, or become, failed: a module error, a non-zero return code, a missing sudo password. They need different fixes, so read which one it is first.

How much verbosity should I use?

-v shows each task's result, which is enough for most failures. -vv adds task paths and more detail, -vvv shows the SSH connection and the module arguments, -vvvv adds connection debugging. Start low: high verbosity buries the useful line and may print secrets.

How do I test an SSH problem outside Ansible?

Run the same connection by hand: ssh -v web-2 true, as the same user and with the same key. ssh prints each step of the connection, the host key check and the authentication, so you see exactly where it stops. Fix it there, then rerun the playbook.

What does Missing sudo password mean?

The login worked, but become needs sudo and sudo on that host asks for a password Ansible does not have. Run with -K (--ask-become-pass) to type it, or give the automation user a NOPASSWD rule in /etc/sudoers.d, which is the usual setup for unattended runs.

Should I silence ansible-lint rules?

Fix first. Most findings are real problems in disguise: free-form k=v arguments that break on spaces, state: latest that upgrades on a random day, missing file modes. A # noqa or a skip_list entry is fine for a reviewed exception, with a comment saying why.

In an interview Junior

A playbook fails. How do you work out why?

First, the kind of failure: UNREACHABLE (no SSH session: host key, address, user or key; I reproduce it with ssh -v host true) or FAILED (connected, the task or become failed). Then I read the message and its Origin line: file, line and column with a caret under the spot, which points at YAML mistakes, unknown modules and bad parameters. Verbosity: -v shows the task result, -vvv the connection and arguments. I add debug with var= or msg, or assert to check my assumptions, and run ansible-lint, which reports each problem with a rule id at a profile level, before the next run.

Also asked: What does --syntax-check catch, and what does it miss? · How would you debug a variable that has an unexpected value? · What would make Ansible say Permission denied (publickey) for one host only?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.