Reporting is empty since the night. orders-sync.service has been failing since exactly 03:00 UTC, and its journal tells a story that looks like it should have worked:
$ journalctl -u orders-sync --no-pager | grep -E "renew|FATAL|ERROR" | head -8
Sep 21 04:00:09 oncall-lab orders-sync[4417]: 2026-09-21T04:00:09.000Z INFO vault: token renewed, ttl=1h
Sep 21 10:00:09 oncall-lab orders-sync[4417]: 2026-09-21T10:00:09.000Z INFO vault: token renewed, ttl=1h
Sep 21 16:00:09 oncall-lab orders-sync[4417]: 2026-09-21T16:00:09.000Z INFO vault: token renewed, ttl=1h
Sep 21 22:00:09 oncall-lab orders-sync[4417]: 2026-09-21T22:00:09.000Z INFO vault: token renewed, ttl=1h
Sep 22 02:45:09 oncall-lab orders-sync[4417]: 2026-09-22T02:45:09.000Z INFO vault: token renewed, ttl=15m (capped by the token max TTL)
Sep 22 03:00:30 oncall-lab orders-sync[4417]: 2026-09-22T03:00:30.000Z ERROR vault: token renew failed: 403 permission denied
Sep 22 03:00:31 oncall-lab orders-sync[4417]: 2026-09-22T03:00:31.000Z ERROR sync: failed to connect to `user=v-approle-orders-s-Hk2xQ9bT4mWfL8nR3pZy-1790046000 database=orders`: 10.0.3.44:5432 (pg.lab): failed SASL auth: FATAL: password authentication failed for user "v-approle-orders-s-Hk2xQ9bT4mWfL8nR3pZy-1790046000" (SQLSTATE 28P01)
Sep 22 03:00:31 oncall-lab orders-sync[4417]: 2026-09-22T03:00:31.000Z FATAL database unavailable, exitingSomeone set it up a while ago: logged in once, put the token in /etc/orders-sync/token, and made the app renew it "so it never expires". The line at 02:45 says otherwise: capped by the token max TTL.
What is happening: renewal has a ceiling
Every Vault token has two clocks:
- TTL: how long until it expires if nobody renews it. Renewing pushes this forward.
- Max TTL: a hard limit counted from when the token was created. No renewal goes past it.
The effective max is the smallest limit that applies: the token's own explicit_max_ttl, the auth role's token_max_ttl, the auth mount's max_lease_ttl, and Vault's system default of 768h (32 days). Renewing near the end gets a shortened TTL and a warning. Here is the same thing on a test token with a 1-hour explicit max:
$ vault token renew -increment=2h $(cat short.tok)
WARNING! The following warnings were returned from Vault:
* TTL of "2h" exceeded the effective max_ttl of "59m59s"; TTL value is capped
accordingly
Key Value
--- -----
token hvs.CAESIMLdtBo8Fr51uEFf3eB5oVERaC2yE4ayBSMtxVGh4KHGh2cy4FD7oDpn6HLFBw5KzosD21fFq3fQBdiHD
token_accessor aO44RNdeEM7A856BQNPpCzmO
token_duration 59m59s
...That warning is the only notice a client gets. An app that only renews treats it as success and dies at the max.
The second half of the outage is worse. The database credentials the app had (v-approle-orders-s-...) came from the database secrets engine, and dynamic secrets are leases owned by the token that created them. When the token expired, Vault revoked its leases and dropped that database user. So the app did not just lose Vault: at 03:00 its database user stopped existing, and the next database login failed.
Diagnosis
1. Is the token dead?
$ vault token lookup $(sudo cat /etc/orders-sync/token)
Error looking up token: Error making API request.
URL: POST https://127.0.0.1:8200/v1/auth/token/lookup
Code: 403. Errors:
* bad tokenbad token = expired or revoked. Compare with permission denied on a secret path, which means the token is alive and the policy says no: that one is a policy and path problem.
2. Where did the max come from?
The token came from the AppRole orders-sync:
$ vault read auth/approle/role/orders-sync
Key Value
--- -----
bind_secret_id true
...
token_max_ttl 24h
token_num_uses 0
token_period 0s
token_policies [orders-sync]
token_ttl 1h
token_type defaulttoken_ttl 1h, token_max_ttl 24h. A token from this role lives at most 24 hours from login. The app logged in at 03:00 the day before. This was a timer set the day it was deployed, not a random outage. Raising token_max_ttl would not help an existing token either: the max is fixed when the token is issued.
3. What did the app hold?
$ sudo cat /etc/orders-sync/orders-sync.conf
# orders-sync configuration
vault_addr=https://127.0.0.1:8200
vault_token_file=/etc/orders-sync/token
vault_creds_path=database/creds/orders-sync
batch_size=500A long-lived token in a file, read by the app. Even with a perfect renewal loop, the app would still need to log in again before the max, so it would need the AppRole secret ID too. That is a lot of Vault logic to put in every service.
The fix: let Vault Agent hold the token
Vault Agent runs next to the app. It logs in with auto-auth, renews, logs in again before the max, and renders secrets into files with templates. The app reads a file and never sees a Vault token.
auto_auth {
method "approle" {
config = {
role_id_file_path = "/etc/vault-agent/role-id"
secret_id_file_path = "/etc/vault-agent/secret-id"
remove_secret_id_file_after_reading = false
}
}
sink "file" {
config = {
path = "/run/vault-agent/orders-sync.token"
mode = 0600
}
}
}
template {
source = "/etc/vault-agent/orders-sync-db.env.tpl"
destination = "/etc/orders-sync/db.env"
perms = "0640"
error_on_missing_key = true
}The template asks for the dynamic credentials:
{{ with secret "database/creds/orders-sync" -}}
DB_USER={{ .Data.username }}
DB_PASSWORD={{ .Data.password }}
{{- end }}Run it as its own unit (ExecStart=/usr/bin/vault agent -config=/etc/vault-agent/orders-sync.hcl), enable it, and watch it work:
$ journalctl -u vault-agent --no-pager -n 12
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.template.server: starting template server
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.sink.server: starting sink server
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.auth.handler: authenticating
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.auth.handler: authentication successful, sending token to sinks
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.sink.file: token written: path=/run/vault-agent/orders-sync.token
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.auth.handler: starting renewal process
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent.template.server: template server received new token
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent: (runner) creating new runner (dry: false, once: false)
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent: (runner) creating watcher
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent: (runner) starting
Sep 22 20:00:05 oncall-lab vault[17396]: 2026-09-22T20:00:05.800Z [INFO] agent: (runner) rendered "/etc/vault-agent/orders-sync-db.env.tpl" => "/etc/orders-sync/db.env"
Sep 22 20:00:05 oncall-lab systemd[1]: Started vault-agent.service - Vault Agent for orders-sync.Point the app at db_env_file=/etc/orders-sync/db.env, delete the old token file, restart it:
$ journalctl -u orders-sync -n 3 --no-pager
Sep 22 20:00:08 oncall-lab orders-sync[17400]: 2026-09-22T20:00:08.800Z INFO orders-sync 1.8.0 starting
Sep 22 20:00:08 oncall-lab orders-sync[17400]: 2026-09-22T20:00:08.800Z INFO sync: copied 3 new orders (as v-approle-orders-s-mNPY5OEzQmVwTahAyp6j-1790107205)
Sep 22 20:00:08 oncall-lab systemd[1]: Started orders-sync.service - orders-sync - copies new orders to the reporting database.One thing to know before the next 3 a.m.: when the agent renders new credentials (its lease is about to end), the file changes. The app must re-read it or be restarted. The template stanza's exec (or a command) does that. Otherwise you have moved the outage from the token to the database password.
Keeping it from coming back
- Read the max, not the TTL.
vault token lookupshowsexpire_timeandexplicit_max_ttl. For roles,vault read auth/<method>/role/<name>showstoken_max_ttl. Put both in the service's runbook. - Clients log in, they do not hoard. Agent (or a client library with a lifetime watcher that re-authenticates) for VMs, the Agent Injector or the CSI provider on Kubernetes. A periodic token (
-period=24h) does live for ever if renewed, which is exactly why it is dangerous when it leaks. Keep those for the few things that cannot log in. - Alert on the warning. Count
capped/ renewal failures in the app's logs, and alert on lease revocations for a role you care about. - Expect expiry. It is the same lesson as kubeadm certificates that expire after a year: a timer set at creation, invisible until it fires. And when Vault itself is down after a reboot, it is a different problem: Vault is sealed.
Practise it
The incident Incident: orders-sync died at 03:00 in the Vault & Secrets Management chapter is this exact setup. You find the max TTL, then replace the hand-made token with a Vault Agent unit, a template and an app that holds no token. The mission Token lifetimes, renewal and the token tree and the drill When does this token really die? train reading TTLs until you can say the expiry time at a glance.