The problem
A playbook that worked yesterday fails today on one host, at 2 a.m., in a job log. The output is dense - JSON on one line, a host that is "UNREACHABLE", a task that failed with a message about something you never touched. Debugging Ansible is mostly reading: knowing which part of the output answers which question, which verbosity level shows what, and which of a handful of usual suspects it is. The second half of this lesson is the tool that catches many of those mistakes before they run: ansible-lint.
What you need to know already: the whole chapter so far, 1.17 (ssh -v, known_hosts), 1.11 (sudo), 6.22 (shellcheck - ansible-lint is the same idea).
Reading a failure
fatal: [web-1]: FAILED! => {"changed": false, "msg": "No package matching 'ngnix' is available"}
fatal = the host stopped here. The JSON is the module's result: read msg first. For command / shell the useful keys are rc, stderr and cmd. A failure that says Task failed: Finalization of task args ... is not the module - it is a template in the task's arguments that could not be rendered (usually 'x' is undefined), and the Origin: lines point at the exact spot.
Verbosity adds detail in layers:
-v every task's result as JSON, on one line
-vv the config file in use, the task's file:line ("task path:")
-vvv the SSH connection commands, the module file sent, the result indented,
and the module's arguments ("invocation") - secrets included, see no_log
-vvvv ssh -vvv: key exchange, which keys were offered, why authentication failed
Start with -v. Go to -vvv for "is it even running the command I think", -vvvv for SSH problems.
UNREACHABLE: the usual suspects
UNREACHABLE means Ansible could not run anything at all - the problem is the connection, before any module. The message is SSH's own error after Failed to connect to the host via ssh:, and there are only a few of them:
$ cd ~/oncall-lab/labs/1a-ansible/try/debug
$ cat inventory.ini
[web]
web-1
web-3
[lb]
lb-1 ansible_host=10.0.5.99
[db]
db-1 ansible_user=deploy
$ ansible all -m ping
web-1 | SUCCESS => {
"ansible_facts": {
"discovered_interpreter_python": "/usr/bin/python3.14"
},
"changed": false,
"ping": "pong"
}
web-3 | UNREACHABLE! => {
"changed": false,
"msg": "Failed to connect to the host via ssh: ssh: Could not resolve hostname web-3: Temporary failure in name resolution",
"unreachable": true
}
db-1 | UNREACHABLE! => {
"changed": false,
"msg": "Failed to connect to the host via ssh: deploy@db-1: Permission denied (publickey).",
"unreachable": true
}
lb-1 | UNREACHABLE! => {
"changed": false,
"msg": "Failed to connect to the host via ssh: ssh: connect to host 10.0.5.99 port 22: Connection timed out",
"unreachable": true
}
Could not resolve hostname web-3: Temporary failure in name resolution
DNS / /etc/hosts: the name does not exist (a typo, or a host not built yet)
connect to host 10.0.5.99 port 22: Connection timed out
nothing answers: wrong IP, host down, firewall dropping
connect to host web-1 port 22: Connection refused
the host is up, nothing listens on that port: sshd down or ansible_port wrong
learner@db-1: Permission denied (publickey).
connected, but the key was refused: wrong user, key not in authorized_keys,
bad permissions on ~/.ssh on the host (4.7)
Host key verification failed.
unknown host key and nobody to answer the prompt, or
"REMOTE HOST IDENTIFICATION HAS CHANGED" - the host was rebuilt (ssh-keygen -R)
The checks, in order: ansible-inventory --host NAME (what address and user does Ansible use?), then the same with plain SSH - ssh -v user@address. If plain SSH fails, Ansible is not the problem.
become problems
Connected fine, then:
fatal: [db-1]: FAILED! => {"msg": "Missing sudo password"}
The task has become: true, the user's sudo needs a password, and Ansible was not given one. Either pass it with -K (--ask-become-pass, prompts BECOME password:), or set up a NOPASSWD sudoers rule for the automation user - the usual choice for a dedicated deploy account. Incorrect sudo password means -K got the wrong one.
Making a playbook explain itself
- name: Show what we have
ansible.builtin.debug:
var: hostvars[inventory_hostname]['ansible_facts']['default_ipv4']
verbosity: 1 # only printed with -v or more
- name: Refuse to run with nonsense input
ansible.builtin.assert:
that:
- app_port | int > 1024
- app_env in ['dev', 'staging', 'prod']
fail_msg: "app_port={{ app_port }} app_env={{ app_env }} is not a valid combination"
- name: Stop here on purpose
ansible.builtin.fail:
msg: "db-1 is in maintenance"
when: inventory_hostname in groups['maintenance'] | default([])
assert turns assumptions into checks with a readable message - put them at the top of a role to catch bad variables before anything changes. Every that: expression must be a real boolean in 2.19+.
After fixing a problem halfway through a long playbook, --start-at-task "Task name" resumes there, and --limit @site.retry-style reruns are replaced by -l with the failed hosts. --step asks before each task, for the careful walk through.
ansible-lint
ansible-lint reads playbooks and roles (without running them) and reports patterns that are known to cause trouble. Every finding has a rule id:
$ ansible-lint messy.yml
WARNING Listing 14 violation(s) that are fatal
name[play]: All plays should be named.
messy.yml:1 Play: web
yaml[truthy]: Truthy value should be one of [false, true]
messy.yml:2
name[missing]: All tasks should be named.
messy.yml:4 Task/Handler: shell apt-get install -y nginx
fqcn[action-core]: Use FQCN for builtin module actions (shell).
messy.yml:4 Use `ansible.builtin.shell` or `ansible.legacy.shell` instead.
no-free-form: Avoid using free-form when calling module actions. (shell)
messy.yml:4 Task/Handler: shell apt-get install -y nginx
no-changed-when: Commands should not change things if nothing needs doing.
messy.yml:4 Task/Handler: shell apt-get install -y nginx
command-instead-of-module: apt-get used in place of apt module
messy.yml:4 Task/Handler: shell apt-get install -y nginx
command-instead-of-shell: Use shell only when shell functionality is required.
messy.yml:4 Task/Handler: shell apt-get install -y nginx
name[casing]: All names should start with an uppercase letter.
messy.yml:5 Task/Handler: copy the index page
fqcn[action-core]: Use FQCN for builtin module actions (copy).
messy.yml:5 Use `ansible.builtin.copy` or `ansible.legacy.copy` instead.
no-free-form: Avoid using free-form when calling module actions. (copy)
messy.yml:5 Task/Handler: copy the index page
risky-file-permissions: File permissions unset or incorrect.
messy.yml:5 Task/Handler: copy the index page
literal-compare: Don't compare to literal True/False.
messy.yml:7 Task/Handler: Restart nginx
no-handler: Tasks that run when changed should likely be handlers.
messy.yml:7 Task/Handler: Restart nginx
Read documentation for instructions on how to ignore specific rule violations.
# Rule Violation Summary
1 command-instead-of-module profile:basic tags:command-shell,idiom
1 command-instead-of-shell profile:basic tags:command-shell,idiom
1 literal-compare profile:basic tags:idiom
1 name[casing] profile:basic tags:idiom
1 name[missing] profile:basic tags:idiom
1 name[play] profile:basic tags:idiom
2 no-free-form profile:basic tags:syntax,risk
1 yaml[truthy] profile:basic tags:formatting,yaml
1 risky-file-permissions profile:safety tags:unpredictability
1 no-changed-when profile:shared tags:command-shell,idempotency
1 no-handler profile:shared tags:idiom
2 fqcn[action-core] profile:production tags:formatting
Failed: 14 failure(s), 0 warning(s) in 1 files processed of 1 encountered. Last profile that met the validation criteria was 'min'.
Each finding: rule-id: message, then file:line and the task. The ones you will meet first:
name[missing] a task without a name
name[casing] a name that does not start with a capital letter
fqcn[action-core] apt instead of ansible.builtin.apt
no-changed-when command/shell without changed_when (or creates/removes)
command-instead-of-module shell: apt-get install ... instead of the apt module
command-instead-of-shell shell: where command would do (no pipes, no globs)
risky-shell-pipe a pipe in shell without set -o pipefail (6.3)
risky-file-permissions copy/template/file without mode:
risky-octal mode: 644 (decimal) instead of "0644"
no-handler when: result.changed - should be a handler
literal-compare when: x == True - write when: x
yaml[truthy] yes/no instead of true/false
var-naming[no-role-prefix] a role variable without the role's name as a prefix
Rules belong to profiles that build on each other: min (it parses) < basic < moderate < safety < shared < production. The last line says which one the code meets; --profile production makes the strictest one the goal. Teams run it in the job that tests the playbooks (exit code 2 = violations), with a .ansible-lint file:
# .ansible-lint
profile: production
skip_list:
- yaml[line-length]
warn_list:
- experimental
For one deliberate exception, a comment on the line: # noqa: no-changed-when. Fix before you silence: most findings are real bugs waiting for the right day.
What you can now do
- Read a failure:
fatal,msg,rc/stderr, theFinalization of task argscase. - Choose the verbosity level for the question you have.
- Diagnose UNREACHABLE from SSH's message, and become errors (
Missing sudo password). - Use
debug,assertandfailto make a playbook check its own inputs. - Run ansible-lint, read its rule ids and profiles, configure it, and use
# noqasparingly.