Vault & Secrets Management: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 32 of the course.
How would you introduce Vault for a platform with many services, so that no service holds a long-lived shared password? Mid
Run Vault as a highly available cluster with auto-unseal and audit devices, then give every workload its own identity: the Kubernetes auth method for pods (ServiceAccount + namespace bound to a role), AppRole for VMs and pipelines, OIDC for humans. Each role maps to a small policy on its own paths (kv/data/<team>/<app>/*, database/creds/<app>-ro).
Replace static database passwords with dynamic secrets: database/creds/<role> creates a user per client with a lease, and rotate-root removes the bootstrap password. Deliver secrets with Vault Agent (re-authentication, templates, renewal) or External Secrets for static config. Manage mounts, roles and policies in Terraform, and keep the audit log in the incident tooling.
Also asked: What happens to your applications when Vault is sealed or down? · How do you rotate a secret that many services share? · How would you migrate secrets from Kubernetes Secrets into Vault?
Why use a secrets manager instead of environment variables or Kubernetes Secrets? Mid
A secrets manager gives one encrypted store with identity-based access: each client logs in, gets a token, and policies decide which paths it may read. Every request lands in an audit trail, and secrets can be short-lived or dynamic, created per client and revoked when the lease ends.
Environment variables leak through /proc/<pid>/environ, crash dumps, child processes and logs, and never expire. Kubernetes Secrets are base64, readable by anyone with get secrets, and have no expiry or read audit. Neither lets you answer "who read this password last week?" or revoke one client's access without changing the secret for everyone.
Also asked: What is the difference between encoding and encryption? · Where have you seen secrets leak in a real project? · What are the downsides of running your own secrets manager?
Vault restarted and every application that uses it is failing. Why, and what do you do? Mid
A restarted Vault comes up sealed: it has the encrypted data but not the root key in memory, so every request except status and health returns 503 "Vault is sealed". I confirm it with vault status (Sealed true, exit code 2) or sys/health returning 503, then get enough key holders - three of five here - to run vault operator unseal with their shares until Sealed is false.
Afterwards I check why it restarted (the unit's journal, OOM, a host reboot) and, long term, move to auto-unseal with a cloud KMS or HSM so a reboot at night does not need people.
Also asked: What is Shamir secret sharing and why does Vault use it? · What would you do with the root token after initialising Vault? · How do recovery keys differ from unseal keys?
Learn it: 32.3 How Vault works: the barrier, seal and unseal, storage, HA
A service worked for weeks and suddenly gets permission denied from Vault. Where do you look? Mid
At the token's lifetime first. I find the token's accessor (from the audit log or the app's logs) and run vault token lookup -accessor <accessor>: if it is gone or the expire_time has passed, it most likely hit its max TTL - 768h by default - because the app only renewed and never re-authenticated. Renewal can extend the TTL only up to the max.
The fix is in the client: log in again through its auth method before the max is reached, ideally with Vault Agent doing it automatically. If the token is still valid, the next suspects are a changed policy or the path (data/ on KV v2).
Also asked: What is the difference between a service token and a batch token? · How do you revoke every token a compromised pipeline created? · Why should the root token not be used for daily work?
Learn it: 32.5 Tokens and auth methods: TTLs, renewal, the token tree
In KV v2, what is the difference between vault kv delete, undelete, destroy and metadata delete? Mid
KV v2 keeps versions of every secret. vault kv delete is a soft delete: it marks the latest version deleted (deletion_time set), and vault kv undelete -versions=N brings it back. vault kv destroy -versions=N removes the data of those versions permanently, but the history in metadata remains. vault kv metadata delete removes the whole secret: every version and the metadata.
Each maps to its own API path - delete/, undelete/, destroy/, metadata/ - so a policy needs separate rules for each, and after a leak you destroy the leaked version rather than only deleting it.
Also asked: What is the difference between vault kv put and vault kv patch? · How would you prevent two deploy jobs from overwriting the same secret? · How do you roll a secret back to an earlier version?
Learn it: 32.8 KV v2: versions, paths and the put that wipes
A token has a policy that allows secret/app/*, but vault kv get secret/app/db is denied. Why? Mid
Because the mount is KV v2, and the CLI hides the real path: vault kv get secret/app/db sends GET /v1/secret/data/app/db. The policy rule path "secret/app/*" never matches a real KV v2 request, so Vault denies by default. The policy must say path "secret/data/app/*" { capabilities = ["read"] }, plus secret/metadata/app/* with list to list.
I would prove it rather than guess: vault token capabilities <token> secret/data/app/db returns deny before and read after the fix, and vault kv get -output-policy secret/app/db prints exactly the rule the command needs.
Also asked: How do policies combine when a token has several of them? · How would you give each team access only to its own path without one policy per team? · What are parameter constraints in a policy used for?
Learn it: 32.11 Policies: paths, capabilities, and the rule that wins
How does an automated job get secrets from Vault without a long-lived token stored in its configuration? Mid
With AppRole: the job's configuration contains only the role ID, which is not secret on its own. A trusted orchestrator - the system that launches the job - requests a short-lived, single-use secret ID with -wrap-ttl, so the job receives a wrapping token instead of the value. The job runs vault unwrap, logs in with vault write auth/approle/login role_id=... secret_id=..., and gets a token of about ten minutes with a narrow policy. Afterwards the secret ID is worthless.
Where the platform signs its own job tokens, the JWT/OIDC auth method removes the secret ID entirely.
Also asked: What does the trusted orchestrator pattern protect against? · How would you detect that a wrapped secret was intercepted? · Which settings would you put on an AppRole role for a nightly job?
What is a dynamic secret, and why is it better than a regularly rotated static password? Mid
A dynamic secret is created for one client at the moment it asks: vault read database/creds/<role> makes Vault run the role's creation statements and return a new database user and password with a lease. The client renews the lease while it runs; when it expires or is revoked - also when the token that owns it is revoked - Vault drops the user.
Compared with a rotated shared password: no credential is shared, every client is traceable in the database's own logs, a leak is limited to one credential for hours, and rotation is just the next read. After configuring the connection, rotate-root removes the last human-known password.
Also asked: What happens to the database users when Vault is unavailable for an hour? · How would you revoke every credential issued for one role? · When would you choose a static role over dynamic credentials?
Learn it: 32.18 Dynamic secrets: database credentials and leases
How would you protect card numbers in a database so that a stolen dump is useless? Mid
Encrypt them with Vault's transit engine: the app sends the base64 plaintext to transit/encrypt/<key> and stores only the returned vault:v1:... ciphertext; to read, it calls transit/decrypt/<key>. The key never leaves Vault, so the database and its backups hold nothing usable on their own.
The app's policy allows only update on encrypt and decrypt for that one key, and every call is audited. Rotation (transit/keys/<key>/rotate) adds versions; transit/rewrap upgrades old rows to the newest version without the app seeing the plaintext, and min_decryption_version retires old versions.
Also asked: Why would you use an intermediate CA in Vault rather than issuing from the root? · How do you check which CA issued a certificate and when it expires? · What is the difference between encryption in transit and encryption as a service?
Learn it: 32.21 PKI and transit: certificates on demand, encryption as a service
An application reads its database password from Vault at start-up and stops working after a few weeks. What happened, and how do you fix it properly? Mid
Its token, or the credential's lease, reached its max TTL. Renewal extends a token only up to the max; after that it must log in again, which the app never does, and when the token expires its leases are revoked, so the database user is dropped.
The structural fix is Vault Agent: auto-auth with AppRole or Kubernetes renews the token and re-authenticates before the max, and templates render the credentials to a file and re-render when they change - a renewed lease, new dynamic credentials - running a command so the app reloads. The app reads a file and never handles tokens.
Also asked: How would you run Vault Agent next to a service on a plain VM? · What does the agent do when Vault is unreachable for a while? · How do you make an application reload when its secrets file changes?
Learn it: 32.24 Vault Agent: auto-auth, templates, and apps that never see Vault
How does a pod in Kubernetes get a secret from Vault without a token baked into its manifest? Mid
It proves its identity with the ServiceAccount token Kubernetes already mounts in every pod. Vault's kubernetes auth method receives that JWT on auth/kubernetes/login with a role name, sends a TokenReview to the API server to confirm it is valid and whose it is, then checks the role: is the ServiceAccount name in bound_service_account_names and the namespace in bound_service_account_namespaces? If yes it issues a short-lived Vault token with the role's policies.
The reviewer must be allowed to create TokenReviews (system:auth-delegator). Delivery is then the Agent Injector (annotations on the pod template; an init container and sidecar render files under /vault/secrets), ESO (a Kubernetes Secret), the CSI driver, or the app calling Vault itself.
Also asked: What is "secret zero" and how do platforms avoid storing it? · What are the trade-offs of syncing Vault secrets into Kubernetes Secrets? · A pod fails to start after you added Vault annotations. How do you debug it?
Learn it: 32.27 Vault on Kubernetes: the auth method, Agent Injector, ESO and CSI
You are on call for a Vault cluster. What do you monitor, and how would you find out who read a particular secret? Mid
I watch the seal status and health: /v1/sys/health (200 active, 429 standby, 503 sealed, 501 not initialised) and vault status, which exits 2 when sealed; raft peers and autopilot's failure tolerance; certificate and token expiry; disk under the audit devices, because Vault blocks requests when no audit device can write.
For "who read it": the audit log has two JSON lines per request. I filter the response lines for the API path - secret/data/... for KV v2 - with jq and read auth.display_name, policies, time and remote address. To prove which value was returned, I compute its HMAC with vault write sys/audit-hash/file input=... and grep the log for it. Backups are raft snapshots with tested restores.
Also asked: How would you set up unsealing so that a reboot at 3 am does not page anyone? · How do you upgrade a three-node Vault cluster without downtime? · What is the difference between rekey and generate-root?
Learn it: 32.30 Operating Vault: health, audit devices, snapshots, the seal, upgrades, OpenBao
You manage Vault with Terraform. How do you keep secrets out of the Terraform state? Mid
I let Terraform own Vault's configuration - vault_mount, vault_auth_backend, vault_policy, AppRole and Kubernetes roles, database connections - and keep secret values out of it, because state stores every attribute in plain text: sensitive = true only hides the plan output, a vault_kv_secret_v2 value is readable in terraform.tfstate with jq.
Values come from elsewhere: dynamic secrets that no human sees, rotate-root right after configuring a database connection, or people and rotation jobs writing KV. Where Terraform must pass a secret, on Terraform 1.11+ I use write-only arguments (data_json_wo with a version) or an ephemeral resource. And the state backend itself is encrypted and access-controlled.
Also asked: How does the Vault provider authenticate when Terraform runs in a build pipeline? · What would you put in a Terraform module for a team that needs Vault access? · How do you detect and handle drift between Terraform and Vault?
Learn it: 32.34 Terraform and Vault: the vault provider, and secrets in state
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.