Why this lesson exists
The dev server of the last lesson forgets everything when it stops and starts already unlocked. A production Vault does neither: it keeps its data on disk, encrypted, and every time the process starts it comes up sealed - running, listening, and refusing to answer anything but "I am sealed" until enough people unlock it. "Vault is sealed" is the most common Vault page an SRE gets, usually right after a reboot. This lesson is how the server works on the inside, so that message makes sense and the fix is routine.
What you need to know already: systemd units and journalctl -u (2.1, 2.30); hardening directives and capabilities (2.26); TLS, CAs and "certificate signed by unknown authority" (9.15); ss -tlnp (9.5); the dev server and VAULT_ADDR (the previous lesson).
The words you need first
- Barrier - Vault's encryption layer. Everything Vault writes to storage goes through it and is encrypted (AES-256-GCM); the storage only ever sees ciphertext.
- Encryption key / keyring - the key the barrier encrypts data with. It is stored in the storage too - encrypted by the root key.
- Root key (older docs: master key) - the key that decrypts the keyring. Vault never writes it anywhere in plain text.
- Sealed - Vault does not have the root key in memory, so it cannot decrypt anything. Unsealing = giving it what it needs to rebuild the root key.
- Shamir's secret sharing - a way to split a secret into N key shares so that any T of them (the threshold) rebuild it, and fewer than T reveal nothing at all.
- Storage backend - where the encrypted data lives: integrated storage (Raft, Vault's own replicated database, the recommended choice), or file, Consul and others.
- HA (high availability) - several Vault servers sharing storage: one active node serves requests, the others are standby and take over if it dies.
The barrier and the seal
When Vault is initialised it generates two keys: the encryption key that protects your data, and a root key that protects the encryption key. The root key is split with Shamir's algorithm into shares - by default 5 shares with a threshold of 3 - printed once and then forgotten by Vault.
Every start goes like this:
- The process starts, opens the storage, starts the listener. It is sealed: the data is there, but encrypted, and Vault cannot read its own configuration (mounts, policies, tokens are all stored behind the barrier).
- Operators submit unseal key shares, one at a time, from wherever they are.
- When the threshold is reached, Vault rebuilds the root key in memory, decrypts the keyring, and reads its mount table. It is unsealed and starts serving.
- A restart, a crash, a reboot or
vault operator sealthrows the root key away: back to step 1.
The point of the split: no single person can unseal Vault alone, and nobody who steals the disk (or a backup) can read it without three of the five key holders.
Auto-unseal
Waiting for three humans after every reboot does not scale, so most production clusters use auto-unseal: the root key is encrypted by a key in a cloud KMS (AWS KMS, GCP KMS, a cloud key vault), an HSM, or another Vault's transit engine, configured with a seal block in the config file. On start Vault asks the KMS to decrypt it and unseals itself. vault operator init then prints recovery keys instead of unseal keys - needed for special operations like generating a new root token, not for unsealing. The trade-off: whoever controls that KMS key controls your Vault.
(simulator) The lab server uses Shamir keys, so you unseal by hand - which is exactly the situation you will be in when auto-unseal itself breaks.
The server on this box
The package installs a systemd unit. Read it - every line has a reason:
$ systemctl cat vault
# /usr/lib/systemd/system/vault.service
[Unit]
Description="HashiCorp Vault - A tool for managing secrets"
Documentation=https://developer.hashicorp.com/vault/docs
Requires=network-online.target
After=network-online.target
ConditionFileNotEmpty=/etc/vault.d/vault.hcl
StartLimitIntervalSec=60
StartLimitBurst=3
[Service]
Type=notify
EnvironmentFile=/etc/vault.d/vault.env
User=vault
Group=vault
ProtectSystem=full
ProtectHome=read-only
PrivateTmp=yes
PrivateDevices=yes
SecureBits=keep-caps
AmbientCapabilities=CAP_IPC_LOCK
CapabilityBoundingSet=CAP_SYSLOG CAP_IPC_LOCK
NoNewPrivileges=yes
ExecStart=/usr/bin/vault server -config=/etc/vault.d/vault.hcl
ExecReload=/bin/kill --signal HUP $MAINPID
KillMode=process
KillSignal=SIGINT
Restart=on-failure
RestartSec=5
TimeoutStopSec=30
LimitNOFILE=65536
LimitMEMLOCK=infinity
LimitCORE=0
[Install]
WantedBy=multi-user.target
- User=vault - the server never runs as root. Everything it touches (storage directory, certificates, audit log files) must be readable or writable by the
vaultuser. - CAP_IPC_LOCK + LimitMEMLOCK=infinity - lets Vault call
mlock()so its memory (which holds the root key while unsealed) is never swapped to disk. LimitCORE=0 - no core dumps, which would contain the same memory. - Type=notify - systemd waits until Vault reports it is ready (listening). Ready is not unsealed:
systemctl statussaysactive (running)for a sealed server. - KillSignal=SIGINT, TimeoutStopSec=30 - Vault shuts down cleanly on SIGINT.
- ExecReload sends SIGHUP: Vault reloads its TLS certificates and log level (not storage or listener addresses) without restarting - and without sealing.
- ConditionFileNotEmpty - no config file, no start.
The configuration file
The package ships /etc/vault.d/vault.hcl with file storage and a self-signed certificate generated at install time. Talk to that server and the CLI refuses:
# an illustration (no ▶): the package's default config, before the lab replaced it
$ vault status
Error checking seal status: Get "https://127.0.0.1:8200/v1/sys/seal-status": tls: failed to verify certificate: x509: certificate signed by unknown authority
The CLI is a Go program: it verifies the server's certificate against the system trust store (9.15), and a self-signed certificate is in no trust store. The fixes, best first: a certificate from your company CA; VAULT_CACERT=/path/to/ca.pem (or -ca-cert) pointing at the CA that signed it; and for a throwaway test only, VAULT_SKIP_VERIFY=true / -tls-skip-verify, which also turns off protection against anyone in the middle.
The lab's config is what a real single-node server looks like:
$ cat /etc/vault.d/vault.hcl
# /etc/vault.d/vault.hcl - oncall-lab (single node, integrated storage)
ui = true
cluster_name = "oncall-lab"
api_addr = "https://127.0.0.1:8200"
cluster_addr = "https://127.0.0.1:8201"
disable_mlock = true
storage "raft" {
path = "/opt/vault/data"
node_id = "oncall-lab"
}
listener "tcp" {
address = "0.0.0.0:8200"
tls_cert_file = "/opt/vault/tls/vault.crt"
tls_key_file = "/opt/vault/tls/vault.key"
}
- storage "raft" - integrated storage in
/opt/vault/data(mode 700, owned byvault). Raft is a consensus protocol: with 3 or 5 nodes each write is accepted when a majority has it, and a node can fail without losing data. Here there is one node. - listener "tcp" - the API on port 8200 on every interface, with TLS. The certificate comes from the LabCorp issuing CA, which this box trusts - so the default
VAULT_ADDR(https://127.0.0.1:8200) just works. - api_addr / cluster_addr - how other nodes and clients reach this node; the cluster port (8201) carries node-to-node traffic and request forwarding.
- disable_mlock = true - HashiCorp's recommendation for integrated storage, because mlock of Raft's memory-mapped database files can exhaust memory. You then disable swap on the host instead, so secrets in memory still never reach disk.
A server that is up, and useless
$ sudo ss -tlnp | grep vault
LISTEN 0 4096 0.0.0.0:8200 0.0.0.0:* users:(("vault",pid=17387,fd=3))
LISTEN 0 4096 0.0.0.0:8201 0.0.0.0:* users:(("vault",pid=17387,fd=3))
The process is running and listening. Now ask it:
$ vault status
Key Value
--- -----
Seal Type shamir
Initialized false
Sealed true
Total Shares 0
Threshold 0
Unseal Progress 0/0
Unseal Nonce n/a
Version 2.1.1
Build Date 2026-09-16T10:41:32Z
Storage Type raft
Removed From Cluster false
HA Enabled true
Initialized false: the storage is empty, nobody has run vault operator init yet. The journal says the same, in the server's own log format (a timestamp, a level, the subsystem, the message):
$ journalctl -u vault --no-pager | tail -2
Sep 22 20:00:03 oncall-lab vault[17387]: 2026-09-22T20:00:03.400Z [INFO] core: seal configuration missing, not initialized
Sep 22 20:00:03 oncall-lab systemd[1]: Started vault.service - "HashiCorp Vault - A tool for managing secrets".
Everything except the status endpoints answers with HTTP 503:
$ vault kv list secret/
Error making API request.
URL: GET https://127.0.0.1:8200/v1/sys/internal/ui/mounts/secret
Code: 503. Errors:
* Vault is not initialized
vault status exit codes
vault status exits 0 when unsealed, 2 when sealed (or uninitialised), 1 when it cannot reach the server at all. Scripts and monitoring use exactly this.
Initialising: once, ever
$ vault operator init -key-shares=5 -key-threshold=3
Unseal Key 1: DaNKMXtNzXA/EgFCecK1uR1lDyvYUFofic7I5D5T09Y=
Unseal Key 2: kxVcIR+eaVoqV6Q97REgLrPb0ab5uHKOM0U+Oy+80e0=
Unseal Key 3: /aXuI9+7BHs103P24c1JakQvINNHTZafw/ViDZkLFDs=
Unseal Key 4: gQsgcyhnJSQ4dEhWbllY5TRpeSi+kZ9k0lTpsPSgWYg=
Unseal Key 5: ABtkFk+1lTJRr81iyDgxApTFfN+lgAhsqz4Zy0NVgPg=
Initial Root Token: hvs.nXPe6dAEtUcLuIcpiR9kl4Ms
Vault initialized with 5 key shares and a key threshold of 3. Please securely
distribute the key shares printed above. When the Vault is re-sealed,
restarted, or stopped, you must supply at least 3 of these keys to unseal it
before it can start servicing requests.
Vault does not store the generated root key. Without at least 3 keys to
reconstruct the root key, Vault will remain permanently sealed!
It is possible to generate new unseal keys, provided you have a quorum of
existing unseal keys shares. See "vault operator rekey" for more information.
This output is the most sensitive text Vault ever prints, and it prints it once:
- The five unseal keys go to five different people (or, with
-pgp-keys, Vault encrypts each share to one person's PGP key so the terminal never shows them in clear). Lose more than two and the data is gone for good - "permanently sealed" means it. - The initial root token has every permission. Use it to set up the first admin access, then revoke it. If you ever need root again, a quorum of key holders can generate a new one (
vault operator generate-root).
The server is now initialised - and sealed:
$ vault status | grep -E "Initialized|Sealed|Progress"
Initialized true
Sealed true
Unseal Progress 0/3
Unsealing
Each key holder runs vault operator unseal (without an argument it prompts, so the key stays out of shell history):
$ vault operator unseal DaNKMXtNzXA/EgFCecK1uR1lDyvYUFofic7I5D5T09Y= | grep -E "Sealed|Progress|Nonce"
Sealed true
Unseal Progress 1/3
Unseal Nonce fa70d7a7-053c-cc2a-7aba-b202d4c903d7
The nonce identifies this unseal attempt: everyone contributing to it sees the same one. Two things can go wrong:
- A string that is not a key at all:
* 'key' must be a valid hex or base64 string. - A well-formed key from another init (an old envelope, another cluster): Vault cannot tell until the threshold is reached, then fails with
* cipher: message authentication failedand the progress starts again from 0.vault operator unseal -resetthrows away a half-finished attempt on purpose.
After the third good share:
$ vault operator unseal ABtkFk+1lTJRr81iyDgxApTFfN+lgAhsqz4Zy0NVgPg=
Key Value
--- -----
Seal Type shamir
Initialized true
Sealed false
Total Shares 5
Threshold 3
Version 2.1.1
Build Date 2026-09-16T10:41:32Z
Storage Type raft
Cluster Name oncall-lab
Cluster ID be487137-86c1-d64f-9365-7ec58638536e
Removed From Cluster false
HA Enabled true
HA Cluster https://127.0.0.1:8201
HA Mode active
Active Since 2026-09-22T20:00:05.010474190Z
Raft Committed Index 52
Raft Applied Index 52
HA Mode active: this node serves requests. In a 3-node cluster the other two would say standby and forward requests to the active one. Raft Committed / Applied Index
- the position in Raft's log; on a healthy cluster every node's numbers move together.
Health checks and the seal as an emergency brake
Load balancers and monitoring do not run the CLI; they call sys/health, which needs no token and answers with a status code that says what state the node is in:
| code | meaning |
|---|---|
| 200 | initialised, unsealed, active |
| 429 | unsealed standby (healthy, but not the one to send writes to) |
| 472 | disaster-recovery secondary (Enterprise) |
| 473 | performance standby (Enterprise) |
| 501 | not initialised |
| 503 | sealed |
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
200
A sealed Vault is safe: nothing can be read from it. That makes sealing the emergency brake if you suspect a compromise - vault operator seal (it needs a token with sudo on sys/seal, so not everyone can pull it). Everything stops: every application that reads secrets fails until the key holders unseal again.
$ vault operator seal
Success! Vault is sealed.
$ curl -s -o /dev/null -w "%{http_code}\n" https://127.0.0.1:8200/v1/sys/health
503
What changed in Vault 2.0 for operators
Vault 2.0 (April 2026) tightened a few things you meet in this lesson's area: sys/rekey and sys/generate-root now require a valid token in addition to the key shares (unless enable_unauthenticated_access is set in the config); a config file or policy with the same attribute twice no longer loads; and requests with uncleaned paths (//, /./, /../) are rejected.
In an interview: "Vault restarted and every app is failing - why, and what do you do?" - a restarted Vault is sealed: it has the encrypted data but not the root key. Confirm with vault status (Sealed true, exit 2) or sys/health (503), get three of the five key holders to run vault operator unseal, and long term move to auto-unseal with a cloud KMS or HSM so a reboot does not need humans.
What you can now do
- Explain the barrier, the root key, Shamir shares and why a restart means sealed.
- Read the systemd unit and
vault.hcl: user, mlock, storage, listener, TLS. - Fix "certificate signed by unknown authority" the right way.
- Initialise a server, distribute the output, unseal it, and read
vault status. - Use
sys/healthcodes and know when to seal on purpose.