OnCallReady

AnsibleLinuxNetworking · 6 min read

Ansible "UNREACHABLE!" and "Missing sudo password": what each error means and how to fix it

UNREACHABLE means Ansible never got an SSH session; Missing sudo password means it got in and become failed. Six real causes, their exact errors and fixes.

One playbook, three hosts, one harmless task with become: true. Two different runs:

terminal
$ ansible-playbook site.yml

PLAY [Who runs things] *********************************************************

TASK [Effective user] **********************************************************
ok: [web-1]
ok: [web-2]
fatal: [db-1]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: deploy@db-1: Permission denied (publickey).", "unreachable": true}

PLAY RECAP *********************************************************************
db-1                       : ok=0    changed=0    unreachable=1    failed=0    skipped=0    rescued=0    ignored=0
web-1                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
terminal
$ ansible-playbook site.yml

PLAY [Who runs things] *********************************************************

TASK [Effective user] **********************************************************
fatal: [web-2]: FAILED! => {"msg": "Missing sudo password"}
ok: [web-1]
ok: [db-1]

PLAY RECAP *********************************************************************
db-1                       : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-1                      : ok=1    changed=0    unreachable=0    failed=0    skipped=0    rescued=0    ignored=0
web-2                      : ok=0    changed=0    unreachable=0    failed=1    skipped=0    rescued=0    ignored=0

They look alike, and they point at completely different layers. The first word after fatal: tells you which.

What is happening: two steps, two kinds of failure

Every task on every host runs in two steps:

  1. Connect. Ansible runs plain OpenSSH (ssh) to the host, as the user from the inventory or your login, with your keys. If no working session comes out of this, the host is UNREACHABLE, and the recap counts it under unreachable=. Nothing ran on it.
  2. Run the module, and with become: true, run it through sudo first. If that step fails, it is FAILED, counted under failed=. The connection was fine.

So UNREACHABLE! is an SSH problem, and the text after Failed to connect to the host via ssh: is ssh's own error, passed through word for word. Missing sudo password is a sudo problem on a host you reached without any trouble.

Diagnosis: test the two layers apart

ping (the Ansible module, not ICMP) logs in and runs Python. It needs no root. The same with -b adds become:

terminal
$ ansible web-2 -m ansible.builtin.ping
web-2 | SUCCESS => {
    "ansible_facts": {
        "discovered_interpreter_python": "/usr/bin/python3.14"
    },
    "changed": false,
    "ping": "pong"
}
$ ansible web-2 -b -m ansible.builtin.ping
web-2 | FAILED! => {
    "msg": "Missing sudo password"
}

SSH works, sudo wants a password. Pass it once per run with -K (--ask-become-pass):

terminal
$ ansible-playbook site.yml -K
BECOME password:

PLAY [Who runs things] *********************************************************

TASK [Effective user] **********************************************************
ok: [web-1]
ok: [db-1]
ok: [web-2]

For automation, the usual fix is a reviewed NOPASSWD rule in /etc/sudoers.d/ for the automation user, on every host the same way. The real bug here is that web-2 differs from its siblings. (ansible_become_password from an encrypted vault file also works, but then keep that secret out of the logs: see the password in the job log.)

When the first ping fails, the problem is SSH. Read the message, and if it is not clear, run the same connection by hand: ssh -v <host> true. Add -vvvv to ansible to see the exact ssh command it ran. These are the causes you will meet most often, each with the message it prints.

1. A user that does not exist there

output
fatal: [db-1]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: deploy@db-1: Permission denied (publickey).", "unreachable": true}

Permission denied (publickey) on one host means a different user or key for that host. Look at its inventory line:

terminal
$ cat inventory.ini
[web]
web-1
web-2

[db]
db-1 ansible_user=deploy

A leftover ansible_user=deploy, and there is no deploy account on db-1. Fix the inventory, not the command line. An -e ansible_user=... workaround hides it until the next person runs the playbook.

2. A stale address

output
fatal: [web-2]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: ssh: connect to host 10.0.5.42 port 22: Connection timed out", "unreachable": true}
terminal
$ cat inventory.ini
[web]
web-1
web-2 ansible_host=10.0.5.42

ansible_host overrides how the name is reached. Is 10.0.5.42 the address you expect? Here it is an old one. A timeout looks exactly like a dead host, so check the address before you check the host. If the name itself does not resolve, the error says Could not resolve hostname, and that is a DNS problem.

3. An unknown host key

output
The authenticity of host 'web-1 (10.0.5.11)' can't be established.
ED25519 key fingerprint is SHA256:u3zC0wq5/mUWjxBz7AnIDq2g2TdC+hDhWiWpnCJr7n9.
This key is not known by any other names.
Are you sure you want to continue connecting (yes/no/[fingerprint])? no
fatal: [web-1]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: Host key verification failed.", "unreachable": true}

A new host that is not in ~/.ssh/known_hosts, and the prompt got a no (in CI there is nobody to say yes). Check the fingerprint against the host's console or provisioning output, accept it once with ssh web-1 true, then rerun. Turning off host_key_checking makes the error go away by removing the check that protects you.

4. A changed host key

output
fatal: [web-1]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\n@    WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED!     @\n@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\nIT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!\n...\nOffending ED25519 key in /home/learner/.ssh/known_hosts:5\n  remove with:\n  ssh-keygen -f '/home/learner/.ssh/known_hosts' -R 'web-1'\nHost key for web-1 has changed and you have requested strict checking.\nHost key verification failed.", "unreachable": true}

The host was rebuilt (or something is in the way). Confirm the rebuild first, then remove the old line and accept the new key like a first contact:

terminal
$ ssh-keygen -R web-1
# Host web-1 found: line 5
/home/learner/.ssh/known_hosts updated.
Original contents retained as /home/learner/.ssh/known_hosts.old
$ ssh web-1 true
The authenticity of host 'web-1 (10.0.5.11)' can't be established.
ED25519 key fingerprint is SHA256:u3zC0wq5/mUWjxBz7AnIDq2g2TdC+hDhWiWpnCJr7n9.
This key is not known by any other names.
Are you sure you want to continue connecting (yes/no/[fingerprint])? yes
Warning: Permanently added 'web-1' (ED25519) to the list of known hosts.

5. A private key others can read

When every host is unreachable at once, suspect your side:

output
fatal: [web-1]: UNREACHABLE! => {"changed": false, "msg": "Failed to connect to the host via ssh: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\n@         WARNING: UNPROTECTED PRIVATE KEY FILE!          @\n@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\nPermissions 0644 for '/home/learner/.ssh/id_ed25519' are too open.\nIt is required that your private key files are NOT accessible by others.\nThis private key will be ignored.\nLoad key \"/home/learner/.ssh/id_ed25519\": bad permissions\nlearner@web-1: Permission denied (publickey).", "unreachable": true}
...
NO MORE HOSTS LEFT *************************************************************

OpenSSH ignores a private key that group or others can read. A copy or an archive extract usually breaks the mode. chmod 600 ~/.ssh/id_ed25519 fixes it. NO MORE HOSTS LEFT means every host has dropped out of the play, so Ansible stopped.

6. Missing sudo password

Covered above: a FAILED, not an UNREACHABLE. -K for a person at a keyboard, a NOPASSWD rule for automation.

Keeping it from coming back

  • Read the recap columns. unreachable= is the network, keys and accounts. failed= is become and the tasks. Alert on them separately in scheduled runs.
  • Pre-load known hosts from a trusted source (the provisioning output, ssh-keyscan from inside the network) instead of disabling the check.
  • Keep host overrides rare. Every ansible_host and ansible_user on a single inventory line is a future surprise. Put shared settings in group_vars.
  • Make hosts alike. The sudo rule, the automation user and its key should come from one role, so one host cannot drift. Once the run connects, the next thing to check is that it only changes what it must.

Practise it

The Ansible chapter's drill Why can Ansible not reach it? builds a fresh fault on every round, from the six above: a host key, an address, a user, a key mode or sudo. You read the error, report the cause and get one clean run on all three hosts. The lesson Debugging and linting goes through the verbosity levels up to -vvvv and maps each SSH error to its cause.

OnCallReady is free, with no ads and no tracking. RSS · All posts