Ansible: Configuration as Code: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 33 of the course.
What is Ansible and how does it work? Junior
Ansible is an agentless configuration management tool. From a control node it connects to the managed nodes in an inventory over SSH and runs playbooks: YAML lists of plays, each a group of hosts and a list of tasks, each task calling a module such as ansible.builtin.apt or ansible.builtin.template. Modules are idempotent: they compare the host with the desired state and act only on a difference, so a second run reports changed=0. Variables and facts make one playbook fit many hosts, Jinja2 templates render config files, handlers restart services only when their config changed, roles package reusable pieces, and ansible-vault keeps secrets encrypted in git.
Also asked: What does idempotent mean, and how do you prove a playbook is idempotent? · What is the difference between an ad-hoc command and a playbook? · How does Ansible connect to a host, and what does the host need?
What does idempotent mean in configuration management, and why does it matter? Junior
Idempotent means running the same thing again on a host that is already right changes nothing. An Ansible module compares the current state with the desired state and acts only on a difference, so the first run reports changed and the second run reports changed=0. It matters because you can re-run a playbook safely at any time to remove drift, and the recap becomes a real change report: if something says changed, something was really different. A script that appends a line or restarts a service on every run is not idempotent, so it cannot be re-run safely and its output tells you nothing.
Also asked: What is the difference between declarative and imperative configuration? · What is the difference between a golden image and configuration management? · What does agentless mean for Ansible?
Learn it: 33.1 Why configuration management
What is an inventory, and how do you target a subset of hosts? Junior
The inventory lists the managed hosts, grouped: in INI a [web] section, a parent group with [prod:children], and host variables on the host line or in host_vars/ (group variables in a [group:vars] section or group_vars/). Every host is also in the group all. To target a subset you give a pattern: a group name, a host, or combinations such as web:!web-2 or prod:&canary, in an ad-hoc command (ansible web -m ansible.builtin.ping), a play's hosts:, or --limit. I check the pattern first with --list-hosts, which shows the hosts without running anything, and ansible-inventory --graph shows the whole tree.
Also asked: What is the difference between ansible and ansible-playbook? · How do you set a variable for one host only? · What does the host key prompt mean the first time Ansible connects?
Learn it: 33.2 Inventory and ad-hoc commands
What is a playbook made of, and how do you read its output? Junior
A playbook is a list of plays. Each play has hosts: (a pattern), options such as become: true, and tasks; each task has a name and calls one module with its arguments, for example ansible.builtin.apt with name: nginx and state: present. Ansible runs it task by task: one TASK banner per task, then a line per host, ok if the host already matched, changed if the module changed something, failed or skipping otherwise. The PLAY RECAP at the end sums ok, changed, unreachable and failed per host. A second run of a correct playbook shows changed=0, which proves it is idempotent.
Also asked: Why is the command module not idempotent, and how do you fix that? · What does gather_facts do, and when would you turn it off? · What does become: true do?
Learn it: 33.5 Playbooks: plays, tasks, modules and idempotency
Explain Ansible variable precedence in a few sentences. Junior
Variables can come from many places and Ansible merges them by precedence. Weakest first: role defaults, then inventory variables (group, then host), group_vars and host_vars files, then facts, then play vars and vars_files, role vars, set_fact and registered results, and finally extra vars. Two rules cover most cases: The more specific place wins (a host beats its group, a child group beats its parent), and The playbook beats the inventory. And -e beats everything. To see the value Ansible really uses for one host I run ansible web-1 -m ansible.builtin.debug -a var=app_port.
Also asked: Where would you put a value that should be easy to override? · What are facts, and how do you use one in a task? · What is the difference between register and set_fact?
Learn it: 33.8 Variables, facts and precedence
How do you manage a config file that differs slightly per host? Junior
With the ansible.builtin.template module and a Jinja2 template in templates/: the file is written once with expressions such as {{ app_port }} or {{ ansible_facts['processor_vcpus'] }}, Filters like default(30), join or upper, and Statements such as for and if. Ansible renders it per host and writes it only if the result differs, so it stays idempotent; --diff shows the change first. I set mode as a quoted string, put # {{ ansible_managed }} at the top, and add validate: (nginx -t -c %s, visudo -cf %s) so a broken file is refused before it replaces the working one. A handler then reloads the service only when the file changed.
Also asked: What is the difference between copy, template and lineinfile? · How would you list every web server's IP in a load balancer config? · What does the validate parameter do?
Learn it: 33.11 Templates (Jinja2) and files
How do you restart a service only when its configuration changed? Junior
The config task (template or copy) notifies a handler: notify: Reload nginx, and under handlers: a task with that name that uses the service module with state: reloaded. The handler runs only if the task reported changed, only once even if several tasks notified it, and at the end of the play (or earlier at meta: flush_handlers). If a later task failed on that host, its pending handlers are dropped, so the new config sits on disk unapplied; --force-handlers prevents that. I also prefer a reload over a restart, and use changed_when: false on read-only commands so they never trigger handlers by accident.
Also asked: What does --check --diff show, and what can it miss? · How would you run only the config part of a big playbook? · What is the difference between notify and listen?
Learn it: 33.14 Handlers, tags, check and diff
What is an Ansible role, and how is it laid out? Junior
Roles are reusable units with a fixed directory layout, created with ansible-galaxy role init: tasks/main.yml (the work), handlers/, templates/ and files/, defaults/main.yml (the knobs, lowest precedence, meant to be overridden), vars/main.yml (internal values, high precedence), and meta/main.yml (dependencies and platforms). A play applies them with roles:, or with import_role (static) and include_role (dynamic). A project keeps thin playbooks such as site.yml that map roles to groups, and pins external collections and roles from Ansible Galaxy in requirements.yml so every run uses the same versions.
Also asked: Why would an inventory variable not override a role variable? · What is the difference between a role and a collection? · How do you share a role between several projects?
How do you handle secrets in Ansible? Junior
With ansible-vault: secrets go in an encrypted file such as group_vars/db/vault.yml (vault_db_password), and a plain vars.yml maps them (db_password: "{{ vault_db_password }}"), so they are encrypted at rest in git but still easy to find and review. The vault password lives outside the repository, in a 600 file referenced by vault_password_file, or an executable that reads a secrets manager. At run time the value is plain, so every task that uses it gets no_log: true, because -v, -vvv, debug and --diff can print it into a job log. If a secret ever leaks, I rotate it.
Also asked: How would you rotate the vault password? · What is the difference between encrypt and encrypt_string? · Why is an environment variable not a safe place for a secret?
Learn it: 33.23 Secrets: ansible-vault and no_log
How would you deploy a change to web servers behind a load balancer without downtime? Junior
A rolling deploy: serial: 1 so the play runs in batches of one host, and max_fail_percentage: 0 so the first failure stops it. pre_tasks drain the host from the pool, a lineinfile that marks it down, delegated to lb-1 with delegate_to (it runs on lb-1), plus a reload. tasks deploy the change and flush the handlers. post_tasks run a health check with uri and until / retries / delay, then put the host back. If the check fails, the host stays drained, the rest are never touched, and users only reach healthy servers.
Also asked: What is the difference between serial and forks? · What does run_once do, and when would you use it? · How would you roll back a bad release with Ansible?
Learn it: 33.26 Rolling changes: serial, health checks and delegate_to
A playbook fails. How do you work out why? Junior
First, the kind of failure: UNREACHABLE (no SSH session: host key, address, user or key; I reproduce it with ssh -v host true) or FAILED (connected, the task or become failed). Then I read the message and its Origin line: file, line and column with a caret under the spot, which points at YAML mistakes, unknown modules and bad parameters. Verbosity: -v shows the task result, -vvv the connection and arguments. I add debug with var= or msg, or assert to check my assumptions, and run ansible-lint, which reports each problem with a rule id at a profile level, before the next run.
Also asked: What does --syntax-check catch, and what does it miss? · How would you debug a variable that has an unexpected value? · What would make Ansible say Permission denied (publickey) for one host only?
Learn it: 33.30 Debugging and linting
When would you use Ansible, and when would you not? Junior
I use Ansible for Layer 2, configuration.: the operating system of servers that already exist, packages, users, files, services, and for ordered changes such as rolling restarts. It fits Long-lived servers, Bootstrapping new machines, Building images, and Day-2 operations as code. It keeps no state of its own; the host is the state. I would not use it to own what another tool owns: Layer 1, provisioning. belongs to a tool that keeps a state of the resources it created, and the runtime of applications to an orchestrator. The rule is that every setting has exactly one owner, otherwise tools fight and cause drift.
Also asked: What is configuration drift, and how do you detect it? · What is the difference between mutable and immutable infrastructure? · How would you run Ansible for a team rather than from your laptop?
Learn it: 33.34 Ansible next to the other tools: who owns what
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.