OnCallReady

Lesson 32.3 · Vault & Secrets Management · 31 min read

How Vault works: the barrier, seal and unseal, storage, HA

In plain words

Think of a bank vault whose door has a lock that needs three of five managers, each with their own piece of the combination. At night the door closes on its own. In the morning the money is all still inside, but nobody can touch it until three managers turn up and enter their pieces.

Vault works the same way. Everything it stores is encrypted behind a barrier. The key that opens it is not on disk; it is split into unseal key shares with Shamir's secret sharing. A freshly started server is sealed: it has the encrypted data and nothing else, and answers almost every request with "Vault is sealed". vault operator init happens once and prints the shares and a root token; vault operator unseal with enough shares opens it. Auto-unseal hands the job to a key service instead of people.

Why this lesson exists

The dev server of the last lesson forgets everything when it stops and starts already unlocked. A production Vault does neither: it keeps its data on disk, encrypted, and every time the process starts it comes up sealed - running, listening, and refusing to answer anything but "I am sealed" until enough people unlock it. "Vault is sealed" is the most common Vault page an SRE gets, usually right after a reboot. This lesson is how the server works on the inside, so that message makes sense and the fix is routine.

What you need to know already: systemd units and journalctl -u (2.1, 2.30); hardening directives and capabilities (2.26); TLS, CAs and "certificate signed by unknown authority" (9.15); ss -tlnp (9.5); the dev server and VAULT_ADDR (the previous lesson).

The words you need first

The barrier and the seal

When Vault is initialised it generates two keys: the encryption key that protects your data, and a root key that protects the encryption key. The root key is split with Shamir's algorithm into shares - by default 5 shares with a threshold of 3 - printed once and then forgotten by Vault.

Every start goes like this:

  1. The process starts, opens the storage, starts the listener. It is sealed: the data is there, but encrypted, and Vault cannot read its own configuration (mounts, policies, tokens are all stored behind the barrier).
  2. Operators submit unseal key shares, one at a time, from wherever they are.
  3. When the threshold is reached, Vault rebuilds the root key in memory, decrypts the keyring, and reads its mount table. It is unsealed and starts serving.
  4. A restart, a crash, a reboot or vault operator seal throws the root key away: back to step 1.

The point of the split: no single person can unseal Vault alone, and nobody who steals the disk (or a backup) can read it without three of the five key holders.

Auto-unseal

Waiting for three humans after every reboot does not scale, so most production clusters use auto-unseal: the root key is encrypted by a key in a cloud KMS (AWS KMS, GCP KMS, a cloud key vault), an HSM, or another Vault's transit engine, configured with a seal block in the config file. On start Vault asks the KMS to decrypt it and unseals itself. vault operator init then prints recovery keys instead of unseal keys - needed for special operations like generating a new root token, not for unsealing. The trade-off: whoever controls that KMS key controls your Vault.

(simulator) The lab server uses Shamir keys, so you unseal by hand - which is exactly the situation you will be in when auto-unseal itself breaks.

The server on this box

The package installs a systemd unit. Read it - every line has a reason:

$ systemctl cat vault
# /usr/lib/systemd/system/vault.service
[Unit]
Description="HashiCorp Vault - A tool for managing secrets"
Documentation=https://developer.hashicorp.com/vault/docs
Requires=network-online.target
After=network-online.target
ConditionFileNotEmpty=/etc/vault.d/vault.hcl
StartLimitIntervalSec=60
StartLimitBurst=3

[Service]
Type=notify
EnvironmentFile=/etc/vault.d/vault.env
User=vault
Group=vault
ProtectSystem=full
ProtectHome=read-only
PrivateTmp=yes
PrivateDevices=yes
SecureBits=keep-caps
AmbientCapabilities=CAP_IPC_LOCK
CapabilityBoundingSet=CAP_SYSLOG CAP_IPC_LOCK
NoNewPrivileges=yes
ExecStart=/usr/bin/vault server -config=/etc/vault.d/vault.hcl
ExecReload=/bin/kill --signal HUP $MAINPID
KillMode=process
KillSignal=SIGINT
Restart=on-failure
RestartSec=5
TimeoutStopSec=30
LimitNOFILE=65536
LimitMEMLOCK=infinity
LimitCORE=0

[Install]
WantedBy=multi-user.target

The configuration file

The package ships /etc/vault.d/vault.hcl with file storage and a self-signed certificate generated at install time. Talk to that server and the CLI refuses:

# an illustration (no ▶): the package's default config, before the lab replaced it
$ vault status
Error checking seal status: Get "https://127.0.0.1:8200/v1/sys/seal-status": tls: failed to verify certificate: x509: certificate signed by unknown authority

The CLI is a Go program: it verifies the server's certificate against the system trust store (9.15), and a self-signed certificate is in no trust store. The fixes, best first: a certificate from your company CA; VAULT_CACERT=/path/to/ca.pem (or -ca-cert) pointing at the CA that signed it; and for a throwaway test only, VAULT_SKIP_VERIFY=true / -tls-skip-verify, which also turns off protection against anyone in the middle.

The lab's config is what a real single-node server looks like:

$ cat /etc/vault.d/vault.hcl
# /etc/vault.d/vault.hcl - oncall-lab (single node, integrated storage)
ui            = true
cluster_name  = "oncall-lab"
api_addr      = "https://127.0.0.1:8200"
cluster_addr  = "https://127.0.0.1:8201"
disable_mlock = true

storage "raft" {
  path    = "/opt/vault/data"
  node_id = "oncall-lab"
}

listener "tcp" {
  address       = "0.0.0.0:8200"
  tls_cert_file = "/opt/vault/tls/vault.crt"
  tls_key_file  = "/opt/vault/tls/vault.key"
}

A server that is up, and useless

$ sudo ss -tlnp | grep vault
LISTEN 0      4096         0.0.0.0:8200      0.0.0.0:*   users:(("vault",pid=17387,fd=3))
LISTEN 0      4096         0.0.0.0:8201      0.0.0.0:*   users:(("vault",pid=17387,fd=3))

The process is running and listening. Now ask it:

$ vault status
Key                     Value
---                     -----
Seal Type               shamir
Initialized             false
Sealed                  true
Total Shares            0
Threshold               0
Unseal Progress         0/0
Unseal Nonce            n/a
Version                 2.1.1
Build Date              2026-09-16T10:41:32Z
Storage Type            raft
Removed From Cluster    false
HA Enabled              true

Initialized false: the storage is empty, nobody has run vault operator init yet. The journal says the same, in the server's own log format (a timestamp, a level, the subsystem, the message):

$ journalctl -u vault --no-pager | tail -2
Sep 22 20:00:03 oncall-lab vault[17387]: 2026-09-22T20:00:03.400Z [INFO]  core: seal configuration missing, not initialized
Sep 22 20:00:03 oncall-lab systemd[1]: Started vault.service - "HashiCorp Vault - A tool for managing secrets".

Everything except the status endpoints answers with HTTP 503:

$ vault kv list secret/
Error making API request.

URL: GET https://127.0.0.1:8200/v1/sys/internal/ui/mounts/secret
Code: 503. Errors:

* Vault is not initialized

vault status exit codes

vault status exits 0 when unsealed, 2 when sealed (or uninitialised), 1 when it cannot reach the server at all. Scripts and monitoring use exactly this.

Initialising: once, ever

$ vault operator init -key-shares=5 -key-threshold=3
Unseal Key 1: DaNKMXtNzXA/EgFCecK1uR1lDyvYUFofic7I5D5T09Y=
Unseal Key 2: kxVcIR+eaVoqV6Q97REgLrPb0ab5uHKOM0U+Oy+80e0=
Unseal Key 3: /aXuI9+7BHs103P24c1JakQvINNHTZafw/ViDZkLFDs=
Unseal Key 4: gQsgcyhnJSQ4dEhWbllY5TRpeSi+kZ9k0lTpsPSgWYg=
Unseal Key 5: ABtkFk+1lTJRr81iyDgxApTFfN+lgAhsqz4Zy0NVgPg=

Initial Root Token: hvs.nXPe6dAEtUcLuIcpiR9kl4Ms

Vault initialized with 5 key shares and a key threshold of 3. Please securely
distribute the key shares printed above. When the Vault is re-sealed,
restarted, or stopped, you must supply at least 3 of these keys to unseal it
before it can start servicing requests.

Vault does not store the generated root key. Without at least 3 keys to
reconstruct the root key, Vault will remain permanently sealed!

It is possible to generate new unseal keys, provided you have a quorum of
existing unseal keys shares. See "vault operator rekey" for more information.

This output is the most sensitive text Vault ever prints, and it prints it once:

The server is now initialised - and sealed:

$ vault status | grep -E "Initialized|Sealed|Progress"
Initialized             true
Sealed                  true
Unseal Progress         0/3

Unsealing

Each key holder runs vault operator unseal (without an argument it prompts, so the key stays out of shell history):

$ vault operator unseal DaNKMXtNzXA/EgFCecK1uR1lDyvYUFofic7I5D5T09Y= | grep -E "Sealed|Progress|Nonce"
Sealed                  true
Unseal Progress         1/3
Unseal Nonce            fa70d7a7-053c-cc2a-7aba-b202d4c903d7

The nonce identifies this unseal attempt: everyone contributing to it sees the same one. Two things can go wrong:

After the third good share:

$ vault operator unseal ABtkFk+1lTJRr81iyDgxApTFfN+lgAhsqz4Zy0NVgPg=
Key                     Value
---                     -----
Seal Type               shamir
Initialized             true
Sealed                  false
Total Shares            5
Threshold               3
Version                 2.1.1
Build Date              2026-09-16T10:41:32Z
Storage Type            raft
Cluster Name            oncall-lab
Cluster ID              be487137-86c1-d64f-9365-7ec58638536e
Removed From Cluster    false
HA Enabled              true
HA Cluster              https://127.0.0.1:8201
HA Mode                 active
Active Since            2026-09-22T20:00:05.010474190Z
Raft Committed Index    52
Raft Applied Index      52

HA Mode active: this node serves requests. In a 3-node cluster the other two would say standby and forward requests to the active one. Raft Committed / Applied Index

Health checks and the seal as an emergency brake

Load balancers and monitoring do not run the CLI; they call sys/health, which needs no token and answers with a status code that says what state the node is in:

codemeaning
200initialised, unsealed, active
429unsealed standby (healthy, but not the one to send writes to)
472disaster-recovery secondary (Enterprise)
473performance standby (Enterprise)
501not initialised
503sealed
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
200

A sealed Vault is safe: nothing can be read from it. That makes sealing the emergency brake if you suspect a compromise - vault operator seal (it needs a token with sudo on sys/seal, so not everyone can pull it). Everything stops: every application that reads secrets fails until the key holders unseal again.

$ vault operator seal
Success! Vault is sealed.
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
503

What changed in Vault 2.0 for operators

Vault 2.0 (April 2026) tightened a few things you meet in this lesson's area: sys/rekey and sys/generate-root now require a valid token in addition to the key shares (unless enable_unauthenticated_access is set in the config); a config file or policy with the same attribute twice no longer loads; and requests with uncleaned paths (//, /./, /../) are rejected.

In an interview: "Vault restarted and every app is failing - why, and what do you do?" - a restarted Vault is sealed: it has the encrypted data but not the root key. Confirm with vault status (Sealed true, exit 2) or sys/health (503), get three of the five key holders to run vault operator unseal, and long term move to auto-unseal with a cloud KMS or HSM so a reboot does not need humans.

What you can now do

Why it helps

"Vault restarted and every app is failing" is the classic Vault outage, and it is almost always the seal: a restart, a crash or an OOM kill leaves the server sealed until someone unseals it. Knowing that vault status exits 2 when sealed, that sys/health answers 503, where the unit and config live (vault.service, /etc/vault.d/vault.hcl), and why a TLS mismatch looks the way it does, makes that a five-minute incident instead of an hour.

The init ceremony also matters for security reviews: who holds the key shares, where the root token went, and why auto-unseal moves the risk to whoever controls the KMS key.

Commands in this lesson

systemctl vault cat ss journalctl curl

FAQ

Why can Vault not just store its key next to the data?

Then anyone who copied the disk or a backup would have both, and the encryption would protect nothing. Keeping the root key only in memory, rebuilt from shares held by different people (or protected by an external KMS), means a stolen disk, snapshot or VM image is useless on its own.

What exactly does init print, and what do I do with it?

The unseal key shares (five by default) and the initial root token, once, never again. Each share goes to a different person or safe; the root token is used to set up auth methods and an admin policy and is then revoked. Losing more shares than the threshold allows means the data can never be decrypted again.

What is the difference between the unseal key and the root token?

The unseal key shares decrypt the data at the storage level: they open the server. The root token is an API credential with every permission: it authorises requests once the server is open. Unsealing never needs a token, and a token can never unseal.

Why does vault status exit with 2?

The Vault CLI uses 0 for success, 1 for an error and 2 for "it worked, but the answer is not what you hoped": for vault status that is a sealed server. Scripts and health checks rely on that difference between "sealed" and "unreachable".

Is the HTTP/HTTPS error I got a Vault bug?

No, it is the client and the listener disagreeing. "Client sent an HTTP request to an HTTPS server" means VAULT_ADDR says http but the listener has TLS; "x509: certificate signed by unknown authority" means the client does not trust the CA, fixed with VAULT_CACERT rather than skipping verification.

In an interview Mid

Vault restarted and every application that uses it is failing. Why, and what do you do?

A restarted Vault comes up sealed: it has the encrypted data but not the root key in memory, so every request except status and health returns 503 "Vault is sealed". I confirm it with vault status (Sealed true, exit code 2) or sys/health returning 503, then get enough key holders - three of five here - to run vault operator unseal with their shares until Sealed is false.

Afterwards I check why it restarted (the unit's journal, OOM, a host reboot) and, long term, move to auto-unseal with a cloud KMS or HSM so a reboot at night does not need people.

Also asked: What is Shamir secret sharing and why does Vault use it? · What would you do with the root token after initialising Vault? · How do recovery keys differ from unseal keys?

Practise this lesson in the terminal Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.