OnCallReady

SecretsSRELinux · 6 min read

"Vault is sealed" after a restart: unseal keys, the threshold and auto-unseal

After a reboot every app that needs a secret fails with 503 Vault is sealed. Why Vault restarts sealed, how to unseal with key shares, and auto-unseal.

The page says "orders-sync failing, no orders copied for 25 min". It does not say Vault. The service's journal does:

terminal
$ journalctl -u orders-sync --no-pager -n 6
Sep 22 20:00:03 oncall-lab orders-sync[17387]: 2026-09-22T20:00:03.000Z INFO orders-sync 1.8.0 starting
Sep 22 20:00:03 oncall-lab orders-sync[17387]: 2026-09-22T20:00:03.000Z ERROR vault: GET /v1/database/creds/orders-sync: 503 Vault is sealed
Sep 22 20:00:03 oncall-lab orders-sync[17387]: 2026-09-22T20:00:03.000Z FATAL cannot load database credentials, exiting
Sep 22 20:00:03 oncall-lab systemd[1]: Started orders-sync.service - orders-sync - copies new orders to the reporting database.
Sep 22 20:00:03 oncall-lab systemd[1]: orders-sync.service: Main process exited, code=exited, status=1/FAILURE
Sep 22 20:00:03 oncall-lab systemd[1]: orders-sync.service: Failed with result 'exit-code'.

The night's patch run rebooted the Vault server. And systemd says Vault is fine:

terminal
$ systemctl status vault --no-pager | head -4
● vault.service - "HashiCorp Vault - A tool for managing secrets"
     Loaded: loaded (/usr/lib/systemd/system/vault.service; enabled; preset: enabled)
     Active: active (running) since Tue 2026-09-22 20:00:03 UTC; 1s ago
 Invocation: 37dbbd492b5406b52476812f87620099

Both are true. The process is running, and it cannot serve a single secret.

What is happening: the key lives only in memory

Vault encrypts everything it writes to storage. The key that decrypts it (the root key, called the master key in older docs) is itself encrypted, and with the default Shamir seal nothing on disk can decrypt it. At vault operator init the unseal key was split into shares (here 5), and any threshold of them (here 3) rebuilds it. Fewer than the threshold reveal nothing at all.

Unsealing means: people bring their shares, Vault rebuilds the key in memory, and starts serving. Restart the process (a reboot, an OOM kill, a pod rescheduled to another node, a package upgrade) and that memory is gone. Every restart comes back sealed, by design. A sealed Vault answers almost every request with a 503, and every app that fetches a secret at start fails at its next restart. That is why the page names the app and not Vault. It is the same pattern as an expired cluster certificate: the dependency fails, and every service that needs it gets the blame.

Diagnosis

1. Ask Vault, not systemd

terminal
$ vault status
Key                     Value
---                     -----
Seal Type               shamir
Initialized             true
Sealed                  true
Total Shares            5
Threshold               3
Unseal Progress         0/3
Unseal Nonce            n/a
Version                 2.1.1
Build Date              2026-09-16T10:41:32Z
Storage Type            raft
Removed From Cluster    false
HA Enabled              true
$ vault status > /dev/null; echo "exit $?"
exit 2

Sealed true, threshold 3 of 5. The exit code is scriptable: 0 unsealed, 2 sealed, 1 error (cannot reach Vault at all). Monitoring should use the health endpoint, which needs no token:

terminal
$ curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:8200/v1/sys/health
503

sys/health returns 200 for an active node, 429 for an unsealed standby, 501 if not initialised and 503 when sealed.

2. Confirm why it is sealed

terminal
$ journalctl -u vault --no-pager | tail -5
Sep 22 20:00:03 oncall-lab vault[17386]: 2026-09-22T20:00:03.000Z [INFO]  core: Initializing version history cache for core
Sep 22 20:00:03 oncall-lab vault[17386]: 2026-09-22T20:00:03.000Z [INFO]  events: Starting event system
Sep 22 20:00:03 oncall-lab vault[17386]: 2026-09-22T20:00:03.000Z [INFO]  core.cluster-listener.tcp: starting listener: listener_address=0.0.0.0:8201
Sep 22 20:00:03 oncall-lab vault[17386]: 2026-09-22T20:00:03.000Z [INFO]  core: vault is sealed
Sep 22 20:00:03 oncall-lab systemd[1]: Started vault.service - "HashiCorp Vault - A tool for managing secrets".

A clean start that ends in core: vault is sealed. Nobody sealed it on purpose and nothing is broken: it restarted. (If the log shows storage errors instead, unsealing will not help. Fix the storage first.)

The fix: three key holders, one share each

Each key holder runs vault operator unseal with their own share, from their own machine. The shares should never be stored together, and never on the Vault server. Each call moves the progress on:

terminal
$ vault operator unseal $(awk '/Unseal Key/ {print $4}' ~/oncall-lab/labs/vault/keyholders/keyholder-1.txt)
Key                     Value
---                     -----
Seal Type               shamir
Initialized             true
Sealed                  true
Total Shares            5
Threshold               3
Unseal Progress         1/3
Unseal Nonce            b8aeee5f-3623-9147-fc81-f8bbc6f32bec
...

Here is the trap that makes this incident last an hour instead of five minutes. The second share was accepted (2/3), the third one was fine too, and yet:

terminal
$ vault operator unseal $(awk '/Unseal Key/ {print $4}' ~/oncall-lab/labs/vault/keyholders/keyholder-3.txt)
Error unsealing: Error making API request.

URL: PUT https://127.0.0.1:8200/v1/sys/unseal
Code: 400. Errors:

* cipher: message authentication failed
$ vault status | grep Progress
Unseal Progress         0/3

Vault cannot check a share on its own. It only knows the combination is wrong when the threshold is reached, and then the progress resets. The share that broke it is not the last one typed. In this lab, key holder 2 still had an envelope from an old vault operator init. Try other combinations to find the bad share, and leave it out:

terminal
$ for n in 1 3 4; do vault operator unseal $(awk '/Unseal Key/ {print $4}' ~/oncall-lab/labs/vault/keyholders/keyholder-$n.txt) > /dev/null; done; vault status | grep -E 'Sealed|HA Mode'
Sealed                  false
HA Mode                 active
$ curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:8200/v1/sys/health
200

Then check the thing users care about. orders-sync is in a restart loop, so it picks Vault up on its next try:

terminal
$ journalctl -u orders-sync --no-pager -n 3
Sep 22 20:00:13 oncall-lab orders-sync[17389]: 2026-09-22T20:00:13.000Z INFO vault: got database credentials v-token-or-orders-s-IKWf7OBQapWtTxl3B20H-1790107213 (lease 1h)
Sep 22 20:00:13 oncall-lab orders-sync[17389]: 2026-09-22T20:00:13.000Z INFO sync: copied 3 new orders (as v-token-or-orders-s-IKWf7OBQapWtTxl3B20H-1790107213)
Sep 22 20:00:13 oncall-lab systemd[1]: Started orders-sync.service - orders-sync - copies new orders to the reporting database.

An app that gave up for good (a systemd unit that hit its start limit, a pod in CrashLoopBackOff with a long back-off) needs a restart from you. With several Vault nodes, unseal every node: a sealed standby cannot take over when the active one goes down.

Prevention: auto-unseal, and an alert that fires first

Auto-unseal. Replace the Shamir seal with a key you already trust to be available: a cloud KMS key or another Vault's transit engine. A seal stanza in the server config does it:

hcl
seal "awskms" {
  region     = "eu-west-1"
  kms_key_id = "alias/vault-unseal"
}

(azurekeyvault, gcpckms and transit work the same way.) A restarted server asks the KMS to decrypt its root key and unseals itself in seconds. The key holders do not go away: they get recovery keys, needed for vault operator generate-root and rekeying, not for every boot. Moving an existing cluster from Shamir to KMS is a one-time seal migration (vault operator unseal -migrate). The open-source fork OpenBao has the same seals and the same commands (bao).

Alert on the seal, not on the apps. Page on sys/health returning 503 (or on vault_core_unsealed being 0). Then a sealed Vault pages you within a minute, before twenty services page their teams.

Test the key ceremony. After every init or rekey, every holder unseals once, so a stale envelope shows up on a quiet Tuesday and not during a reboot. Patch runs should reboot Vault nodes one at a time, with someone ready to unseal, or only once auto-unseal is in.

Run it as an incident. Half the company is down while three people get found. Declare it, give one person the key holders, and keep the updates going: see how to run an incident as incident commander. The next Vault outage is often quieter, a token that expired at 3 a.m..

Practise it

The lab is Incident: everything that needs a secret is down after the patch reboot, in the Vault & Secrets Management chapter: a sealed Raft node, five key holders, one stale share, and a service that must be copying orders again before you are done. The drill Unseal a Vault someone started unsealing covers the nonce and the half-done unseal.

OnCallReady is free, with no ads and no tracking. RSS · All posts