OnCallReady

Lesson 32.30 · Vault & Secrets Management · 26 min read

Operating Vault: health, audit devices, snapshots, the seal, upgrades, OpenBao

In plain words

Imagine the bank vault in a small town: when it is closed and locked, every shop that needs cash has to wait. So the town keeps an eye on three things - is the vault open, who took what out of it, and is there a copy of the ledger somewhere else if the building burns down.

Running Vault is the same. sys/health and vault status tell you whether it is unsealed and serving. An audit device writes one line for every request, with values replaced by HMACs so the log is useful without containing the secrets. A raft snapshot is the copy of the ledger - useless without the unseal keys, priceless with them. And because a restarted Vault starts locked, someone with keys, or an auto-unseal key service, has to open it again.

The problem: Vault is now tier 0

Once the apps read their passwords from Vault, Vault is the one service every other service depends on. When it is sealed, down or refusing requests, logins fail, dynamic database users are not issued, certificates are not renewed - and the incident is "everything is broken" rather than "Vault is broken". Operating Vault is mostly about four things: knowing its state, knowing who did what, being able to get it back and changing it safely.

What you need to know already: systemd units and journald (2.1, 2.30), curl and HTTP status codes (9.21), TLS (9.15), jq (7.11), etcd snapshots as the Kubernetes equivalent of a backup (18.20), and this chapter's lessons on the seal, the barrier and the unseal keys, tokens and accessors, and policies.

The words you need first

Knowing its state

vault status exits 0 unsealed, 2 sealed, 1 on error - scripts and monitoring can use the exit code without parsing the table:

Load balancers and monitoring use the unauthenticated health endpoint /v1/sys/health, which answers with the state in the status code:

codemeaning
200initialised, unsealed, active
429unsealed standby (an HA follower)
472 / 473disaster-recovery secondary / performance standby (Enterprise)
501not initialised
503sealed
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
200
$ curl -s https://127.0.0.1:8200/v1/sys/health | jq .
{
  "initialized": true,
  "sealed": false,
  "standby": false,
  "performance_standby": false,
  "replication_performance_mode": "disabled",
  "replication_dr_mode": "disabled",
  "server_time_utc": 1790107204,
  "version": "2.1.1",
  "enterprise": false,
  "echo_duration_ms": 0,
  "clock_skew_ms": 0,
  "replication_primary_canary_age_ms": 0,
  "removed_from_cluster": false,
  "cluster_name": "oncall-lab",
  "cluster_id": "88511e62-10ff-4f01-b02c-650f939ce6f8"
}

Query parameters change the codes for a health check that only wants "can it serve?": ?standbyok=true turns a standby's 429 into 200, ?sealedcode=200 and friends remap the others. A load balancer in front of three nodes typically sends traffic only to the 200 one.

On an HA cluster, vault operator members and (with integrated storage) vault operator raft list-peers and vault operator raft autopilot state show the nodes, the leader and whether losing one more node would lose the quorum (Failure Tolerance):

$ vault operator members
Host Name     API Address               Cluster Address           Active Node    Version    Upgrade Version    Redundancy Zone    Last Echo
---------     -----------               ---------------           -----------    -------    ---------------    ---------------    ---------
oncall-lab    https://127.0.0.1:8200    https://127.0.0.1:8201    true           2.1.1      2.1.1              n/a                n/a
$ vault operator raft autopilot state
Healthy:                         true
Failure Tolerance:               0
Leader:                          oncall-lab
Voters:
   oncall-lab

The lab is one node, so Failure Tolerance: 0: production integrated storage runs 3 or 5 voters (a raft quorum, like etcd's: 3 nodes survive the loss of 1, 5 survive 2). The server log is the journal of vault.service: journalctl -u vault. Lines to know: core: vault is sealed, core: vault is unsealed, core: seal configuration missing, not initialized, expiration: revoked lease, and every [ERROR].

Knowing who did what: audit devices

Without an audit device, Vault records nothing about who read a secret. Enabling one is a single command - and the first try usually teaches the same lesson:

The server process must be able to create and append to the file. Give the vault user its own directory, not world-readable (the log holds every token's HMAC and every path anyone touched):

$ sudo mkdir -p /var/log/vault && sudo chown vault:vault /var/log/vault && sudo chmod 750 /var/log/vault
$ vault audit enable file file_path=/var/log/vault/audit.log
Success! Enabled the file audit device at: file/
$ vault audit list -detailed
Path     Type    Description    Replication    Options
----     ----    -----------    -----------    -------
file/    file    n/a            replicated     file_path=/var/log/vault/audit.log

Enabling an audit device needs sudo on sys/audit/<path>: an attacker who could disable auditing could cover their tracks.

Reading the log

Every request writes two lines, "type": "request" and "type": "response". The response of a KV read, shortened:

$ sudo tail -n 1 /var/log/vault/audit.log | jq '{type, who: .auth.display_name, policies: .auth.policies, path: .request.path, op: .request.operation, from: .request.remote_address, value: .response.data.data}'
{
  "type": "response",
  "who": "token-lab-admin",
  "policies": [
    "default",
    "lab-admin"
  ],
  "path": "secret/data/payments/api-key",
  "op": "read",
  "from": "127.0.0.1",
  "value": {
    "API_KEY": "hmac-sha256:50d3072051d308b352d30a4653d30bd954d30d6c55d30eff56d3109257d31225"
  }
}

"Who read the payments key yesterday?" is one jq filter:

$ sudo jq -c 'select(.type=="response" and .request.path=="secret/data/payments/api-key") | {time, op: .request.operation, who: .auth.display_name}' /var/log/vault/audit.log
{"time":"2026-09-22T20:00:03.604340076Z","op":"update","who":"token-lab-admin"}
{"time":"2026-09-22T20:00:03.704131976Z","op":"read","who":"token-lab-admin"}

"Was this leaked key ever returned by Vault?": ask Vault for the HMAC of the leaked value with the audit device's key, then search the log for that HMAC:

$ vault write sys/audit-hash/file input=pk_live_7f3a9c21e8
Key     Value
---     -----
hash    hmac-sha256:50d3072051d308b352d30a4653d30bd954d30d6c55d30eff56d3109257d31225

It matches the API_KEY value in the response above, so that read returned exactly this key.

The audit device can stop Vault

Vault refuses to answer a request it could not audit. If every enabled audit device fails (the disk is full, the file was made read-only), every request fails with a 500 and the log says no audit backend succeeded in logging the request. That is a feature (no unaudited access) and an outage cause. So: enable two devices (a file and a socket/syslog one), watch the disk, and rotate the file with logrotate plus systemctl reload vault (a SIGHUP makes Vault reopen the file).

Getting it back: snapshots

With integrated storage, a backup is a raft snapshot: one file with the whole storage, taken without stopping Vault. It needs a token with sudo on sys/storage/raft/snapshot (or root):

$ vault operator raft snapshot save /tmp/backup.snap; ls -l /tmp/backup.snap
-rw-r--r-- 1 learner learner 9007 Sep 22 20:00 /tmp/backup.snap

(simulator) A real snapshot is a gzip archive; the lab's is a text file with the same role.

Three things decide whether that file saves you:

The seal, operationally

# shown, not run in the lab: rekey and generate-root need a quorum of key holders
$ vault operator rekey -init -key-shares=5 -key-threshold=3
$ vault operator rekey          # each key holder, one share at a time
$ vault operator generate-root -init
$ vault operator generate-root  # each key holder; then -decode with the OTP

Changing it safely: upgrades

Vault ships a minor version about every four months (2.0 in April 2026, 2.1 in September 2026) and patch releases in between. An upgrade is: read the release notes' important changes page, take a snapshot, replace the binary, restart - for HA, the standbys first and the active node last (it hands leadership over). vault status and vault operator members show the version per node.

What Vault 2.0 changed that bites old scripts and configs:

A rolling upgrade of a three-node raft cluster, step by step:

  1. vault operator raft autopilot state - healthy, failure tolerance at least 1, every voter caught up. Do not start an upgrade on a cluster that is already degraded.
  2. vault operator raft snapshot save and copy it off the box.
  3. One standby at a time: install the new package, systemctl restart vault, unseal it (or let auto-unseal do it), and wait until vault status on it shows the new version and its Raft Applied Index has caught up with the leader's.
  4. Last, the active node: vault operator step-down hands leadership to an upgraded standby, then upgrade the old leader like the others.
  5. Check the dependants, not just Vault: logins, a secret read, a dynamic credential.

Never skip the snapshot because "it is only a patch release", and never upgrade the leader first: a mixed-version cluster is supported in one direction only, newer standbys behind an older leader until it steps down.

Licensing, and OpenBao

Since August 2023 (Vault 1.15) Vault is under the Business Source License (BSL 1.1): free to use, including in production, but not to offer as a competing service. HashiCorp is part of IBM since 2025. OpenBao is the community fork of the last MPL-licensed Vault, run by the Linux Foundation (OpenSSF), with the CLI bao and the same API for the parts they share - version 2.7.1 in October 2026. Commands in this chapter work the same with bao in place of vault for the core features (KV, policies, tokens, auth methods, database, PKI, transit); enterprise features (namespaces in Vault, replication, Sentinel) differ. When a job ad says "Vault", knowing that OpenBao exists and why (licence, governance) is part of the answer.

In an interview: "Vault is tier 0: I monitor sys/health and the seal status, run two audit devices because Vault blocks when auditing fails, take raft snapshots and test restores, keep unseal key shares and snapshots apart, and upgrade standbys first after reading the important-changes notes."

What you can now do

Why it helps

Once applications depend on Vault, a sealed or blocked Vault takes everything down at once, and the page you get says "checkout failing", not "Vault sealed". Knowing that vault status exits 2 when sealed, that /v1/sys/health answers 503, that a full disk under the only audit device makes Vault refuse requests, and how to answer "who read this secret?" with jq and sys/audit-hash, turns those incidents into short ones. Backups, key-holder ceremonies, auto-unseal and upgrades are what interviewers ask a platform engineer who claims to "run Vault", and the licence change and OpenBao come up whenever a team chooses a secrets manager.

Commands in this lesson

vault curl mkdir tail jq

FAQ

Why does Vault refuse requests when the audit log cannot be written?

By design: Vault prefers no answer to an unaudited answer. If every enabled audit device fails to log a request - a full disk, a read-only file, a dead socket - the request fails and the server logs that no audit backend succeeded. Run at least two devices of different kinds, alert on the disk, and rotate the file with logrotate plus a reload (SIGHUP) so Vault reopens it.

If the audit log only has HMACs, how can it tell me anything?

Paths, operations, identities, policies, times and remote addresses are in clear text; only values (secrets, tokens, accessors) are HMAC'd. To check a specific value - a leaked key, a token's accessor - ask Vault for its HMAC with vault write sys/audit-hash/<device> input=..., which uses the device's own key, and search the log for that string.

Is a raft snapshot enough as a backup?

It contains everything, encrypted by the barrier, so it is only restorable with the unseal keys (or the auto-unseal key) of the cluster it came from - keep both, separately. It also only counts once you have restored it somewhere: schedule a test restore into a scratch Vault. A restore replaces everything, including tokens and leases, with the snapshot's state.

What does auto-unseal change?

A seal stanza (cloud KMS, HSM, or another Vault's transit) lets a restarted server decrypt its root key itself, so reboots no longer need key holders. vault operator init then gives recovery keys instead of unseal keys: they authorise rekey and generate-root but cannot unseal. The risk moves to whoever controls the KMS key.

What is OpenBao, and should I care?

The Linux Foundation fork of the last MPL-licensed Vault, created after HashiCorp moved Vault to the Business Source License in 2023. Its CLI is bao and its core API matches Vault's for KV, policies, tokens, auth methods and the main engines. You care when a company avoids BSL software or wants community governance; the operating skills are the same.

In an interview Mid

You are on call for a Vault cluster. What do you monitor, and how would you find out who read a particular secret?

I watch the seal status and health: /v1/sys/health (200 active, 429 standby, 503 sealed, 501 not initialised) and vault status, which exits 2 when sealed; raft peers and autopilot's failure tolerance; certificate and token expiry; disk under the audit devices, because Vault blocks requests when no audit device can write.

For "who read it": the audit log has two JSON lines per request. I filter the response lines for the API path - secret/data/... for KV v2 - with jq and read auth.display_name, policies, time and remote address. To prove which value was returned, I compute its HMAC with vault write sys/audit-hash/file input=... and grep the log for it. Backups are raft snapshots with tested restores.

Also asked: How would you set up unsealing so that a reboot at 3 am does not page anyone? · How do you upgrade a three-node Vault cluster without downtime? · What is the difference between rekey and generate-root?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.