The problem: Vault is now tier 0
Once the apps read their passwords from Vault, Vault is the one service every other service depends on. When it is sealed, down or refusing requests, logins fail, dynamic database users are not issued, certificates are not renewed - and the incident is "everything is broken" rather than "Vault is broken". Operating Vault is mostly about four things: knowing its state, knowing who did what, being able to get it back and changing it safely.
What you need to know already: systemd units and journald (2.1, 2.30), curl and HTTP status codes (9.21), TLS (9.15), jq (7.11), etcd snapshots as the Kubernetes equivalent of a backup (18.20), and this chapter's lessons on the seal, the barrier and the unseal keys, tokens and accessors, and policies.
The words you need first
- Audit device - a place Vault writes one JSON line per request and one per response: who (token, policies, display name), what (path, operation) and from where. Values are not written in clear text but as HMACs.
- HMAC - a keyed hash:
hmac-sha256:plus 64 hex characters. The same input and the same key always give the same HMAC, but you cannot get the input back. Vault keeps the key, so it can tell you the HMAC of a value you suspect - that is how you search an audit log for a leaked password without the log containing it. - Snapshot - a consistent copy of the whole integrated storage (raft), taken while Vault runs. It is still encrypted by the barrier.
- Auto-unseal - Vault asks a key service (a cloud KMS, an HSM, or another Vault's transit engine) to decrypt its root key at start, instead of waiting for humans with unseal key shares.
Knowing its state
vault status exits 0 unsealed, 2 sealed, 1 on error - scripts and monitoring can use the exit code without parsing the table:
Load balancers and monitoring use the unauthenticated health endpoint /v1/sys/health, which answers with the state in the status code:
| code | meaning |
|---|---|
| 200 | initialised, unsealed, active |
| 429 | unsealed standby (an HA follower) |
| 472 / 473 | disaster-recovery secondary / performance standby (Enterprise) |
| 501 | not initialised |
| 503 | sealed |
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
200
$ curl -s https://127.0.0.1:8200/v1/sys/health | jq .
{
"initialized": true,
"sealed": false,
"standby": false,
"performance_standby": false,
"replication_performance_mode": "disabled",
"replication_dr_mode": "disabled",
"server_time_utc": 1790107204,
"version": "2.1.1",
"enterprise": false,
"echo_duration_ms": 0,
"clock_skew_ms": 0,
"replication_primary_canary_age_ms": 0,
"removed_from_cluster": false,
"cluster_name": "oncall-lab",
"cluster_id": "88511e62-10ff-4f01-b02c-650f939ce6f8"
}
Query parameters change the codes for a health check that only wants "can it serve?": ?standbyok=true turns a standby's 429 into 200, ?sealedcode=200 and friends remap the others. A load balancer in front of three nodes typically sends traffic only to the 200 one.
On an HA cluster, vault operator members and (with integrated storage) vault operator raft list-peers and vault operator raft autopilot state show the nodes, the leader and whether losing one more node would lose the quorum (Failure Tolerance):
$ vault operator members
Host Name API Address Cluster Address Active Node Version Upgrade Version Redundancy Zone Last Echo
--------- ----------- --------------- ----------- ------- --------------- --------------- ---------
oncall-lab https://127.0.0.1:8200 https://127.0.0.1:8201 true 2.1.1 2.1.1 n/a n/a
$ vault operator raft autopilot state
Healthy: true
Failure Tolerance: 0
Leader: oncall-lab
Voters:
oncall-lab
The lab is one node, so Failure Tolerance: 0: production integrated storage runs 3 or 5 voters (a raft quorum, like etcd's: 3 nodes survive the loss of 1, 5 survive 2). The server log is the journal of vault.service: journalctl -u vault. Lines to know: core: vault is sealed, core: vault is unsealed, core: seal configuration missing, not initialized, expiration: revoked lease, and every [ERROR].
Knowing who did what: audit devices
Without an audit device, Vault records nothing about who read a secret. Enabling one is a single command - and the first try usually teaches the same lesson:
The server process must be able to create and append to the file. Give the vault user its own directory, not world-readable (the log holds every token's HMAC and every path anyone touched):
$ sudo mkdir -p /var/log/vault && sudo chown vault:vault /var/log/vault && sudo chmod 750 /var/log/vault
$ vault audit enable file file_path=/var/log/vault/audit.log
Success! Enabled the file audit device at: file/
$ vault audit list -detailed
Path Type Description Replication Options
---- ---- ----------- ----------- -------
file/ file n/a replicated file_path=/var/log/vault/audit.log
Enabling an audit device needs sudo on sys/audit/<path>: an attacker who could disable auditing could cover their tracks.
Reading the log
Every request writes two lines, "type": "request" and "type": "response". The response of a KV read, shortened:
$ sudo tail -n 1 /var/log/vault/audit.log | jq '{type, who: .auth.display_name, policies: .auth.policies, path: .request.path, op: .request.operation, from: .request.remote_address, value: .response.data.data}'
{
"type": "response",
"who": "token-lab-admin",
"policies": [
"default",
"lab-admin"
],
"path": "secret/data/payments/api-key",
"op": "read",
"from": "127.0.0.1",
"value": {
"API_KEY": "hmac-sha256:50d3072051d308b352d30a4653d30bd954d30d6c55d30eff56d3109257d31225"
}
}
auth.display_nameis the human handle:token-<display-name>for tokens,userpass-alicefor a userpass login,kubernetes-shop-ordersfor a pod;auth.metadataholds the rest (username, role, ServiceAccount).auth.client_tokenandauth.accessorare HMAC'd too: compare them withvault write sys/audit-hash/file input=<accessor>to match a token.request.pathis the API path:secret/data/...for a KV v2 read.
"Who read the payments key yesterday?" is one jq filter:
$ sudo jq -c 'select(.type=="response" and .request.path=="secret/data/payments/api-key") | {time, op: .request.operation, who: .auth.display_name}' /var/log/vault/audit.log
{"time":"2026-09-22T20:00:03.604340076Z","op":"update","who":"token-lab-admin"}
{"time":"2026-09-22T20:00:03.704131976Z","op":"read","who":"token-lab-admin"}
"Was this leaked key ever returned by Vault?": ask Vault for the HMAC of the leaked value with the audit device's key, then search the log for that HMAC:
$ vault write sys/audit-hash/file input=pk_live_7f3a9c21e8
Key Value
--- -----
hash hmac-sha256:50d3072051d308b352d30a4653d30bd954d30d6c55d30eff56d3109257d31225
It matches the API_KEY value in the response above, so that read returned exactly this key.
The audit device can stop Vault
Vault refuses to answer a request it could not audit. If every enabled audit device fails (the disk is full, the file was made read-only), every request fails with a 500 and the log says no audit backend succeeded in logging the request. That is a feature (no unaudited access) and an outage cause. So: enable two devices (a file and a socket/syslog one), watch the disk, and rotate the file with logrotate plus systemctl reload vault (a SIGHUP makes Vault reopen the file).
Getting it back: snapshots
With integrated storage, a backup is a raft snapshot: one file with the whole storage, taken without stopping Vault. It needs a token with sudo on sys/storage/raft/snapshot (or root):
$ vault operator raft snapshot save /tmp/backup.snap; ls -l /tmp/backup.snap
-rw-r--r-- 1 learner learner 9007 Sep 22 20:00 /tmp/backup.snap
(simulator) A real snapshot is a gzip archive; the lab's is a text file with the same role.
Three things decide whether that file saves you:
- It is useless without the unseal keys (or the auto-unseal key) of the cluster it came from. The data inside is still encrypted by the barrier. Keep the key holders' shares and the snapshots in different places, both safe.
- Restore replaces everything - every secret, policy, token and lease goes back to the moment of the snapshot.
vault operator raft snapshot restore FILEon the same cluster;-forcefor a snapshot from another cluster (then that cluster's unseal keys are the ones that work). Tokens issued after the snapshot stop working. - An untested backup is a hope. Restore into a scratch Vault on a schedule and read one secret back. Vault Enterprise has automated snapshots; with Community you run
snapshot savefrom a systemd timer (2.16) and copy the file off the box.
The seal, operationally
- The emergency brake.
vault operator sealstops all service at once - for a suspected compromise. It needssudoonsys/seal. Getting back needs the key holders. - Key holders. With Shamir,
vault operator initprinted N shares once; any T of them unseal. Spread them across people and places;-pgp-keyscan encrypt each share to one holder's PGP key so nobody ever sees another's share. A share kept from an older initialisation does not fail on entry: it fails when the threshold is reached, withcipher: message authentication failed, and the progress resets. - Rekey (
vault operator rekey) makes new shares - a new N or T, or after a key holder leaves - from a quorum of the current ones. Generate-root (vault operator generate-root) makes a new root token from a quorum, for the day you need one: revoke it again afterwards. Since Vault 2.0 both also need a valid Vault token, unless the server'senable_unauthenticated_accesssetting allows it. - Auto-unseal. A
sealstanza invault.hcl(seal "awskms",seal "azurekeyvault",seal "gcpckms",seal "transit",seal "pkcs11") makes a restarted server unseal itself;vault operator initthen prints recovery keys, used for rekey and generate-root but unable to unseal. It moves the risk: whoever controls that KMS key controls the unseal. (simulator) The lab's server uses Shamir keys; the stanza is described here, not run.
# shown, not run in the lab: rekey and generate-root need a quorum of key holders
$ vault operator rekey -init -key-shares=5 -key-threshold=3
$ vault operator rekey # each key holder, one share at a time
$ vault operator generate-root -init
$ vault operator generate-root # each key holder; then -decode with the OTP
Changing it safely: upgrades
Vault ships a minor version about every four months (2.0 in April 2026, 2.1 in September 2026) and patch releases in between. An upgrade is: read the release notes' important changes page, take a snapshot, replace the binary, restart - for HA, the standbys first and the active node last (it hands leadership over). vault status and vault operator members show the version per node.
What Vault 2.0 changed that bites old scripts and configs:
sys/rekey,sys/generate-rootand the DR operation token now need a token;- request paths that are not clean (
//,/./,/../) are rejected; - HCL configuration and policies with a duplicated attribute fail to parse instead of warning;
- the official container image no longer needs the
IPC_LOCKcapability (setdisable_mlock = trueif swap is on).
A rolling upgrade of a three-node raft cluster, step by step:
vault operator raft autopilot state- healthy, failure tolerance at least 1, every voter caught up. Do not start an upgrade on a cluster that is already degraded.vault operator raft snapshot saveand copy it off the box.- One standby at a time: install the new package,
systemctl restart vault, unseal it (or let auto-unseal do it), and wait untilvault statuson it shows the new version and itsRaft Applied Indexhas caught up with the leader's. - Last, the active node:
vault operator step-downhands leadership to an upgraded standby, then upgrade the old leader like the others. - Check the dependants, not just Vault: logins, a secret read, a dynamic credential.
Never skip the snapshot because "it is only a patch release", and never upgrade the leader first: a mixed-version cluster is supported in one direction only, newer standbys behind an older leader until it steps down.
Licensing, and OpenBao
Since August 2023 (Vault 1.15) Vault is under the Business Source License (BSL 1.1): free to use, including in production, but not to offer as a competing service. HashiCorp is part of IBM since 2025. OpenBao is the community fork of the last MPL-licensed Vault, run by the Linux Foundation (OpenSSF), with the CLI bao and the same API for the parts they share - version 2.7.1 in October 2026. Commands in this chapter work the same with bao in place of vault for the core features (KV, policies, tokens, auth methods, database, PKI, transit); enterprise features (namespaces in Vault, replication, Sentinel) differ. When a job ad says "Vault", knowing that OpenBao exists and why (licence, governance) is part of the answer.
In an interview: "Vault is tier 0: I monitor sys/health and the seal status, run two audit devices because Vault blocks when auditing fails, take raft snapshots and test restores, keep unseal key shares and snapshots apart, and upgrade standbys first after reading the important-changes notes."
What you can now do
- Read Vault's state from
vault statusexit codes,/v1/sys/healthstatus codes,vault operator membersand the journal. - Enable a file audit device with the right ownership, find who read a path with
jq, and match a suspected value withsys/audit-hash. - Explain why a failing audit device stops Vault, and what to do about it.
- Take a raft snapshot and say what a restore needs (and replaces).
- Explain Shamir key holders, rekey, generate-root, auto-unseal and recovery keys.
- Plan an upgrade, and say what BSL and OpenBao mean for a team choosing Vault.