After every docker compose up -d, the app container keeps coming back. "It works if we restart it by hand a minute later... sometimes."
$ docker ps --filter name=shop-app
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
aaec3f38640f shop-app "java -jar app.jar" 24 seconds ago Restarting (1) 2 seconds ago shop-app-1Look twice and the STATUS column sometimes says Up 1 second instead. Both are the same loop caught at different moments.
What is going on
A container lives exactly as long as its main process, PID 1 inside it. When that process exits, the container stops, and its exit code is PID 1's exit status. A restart policy decides what Docker does next:
--restart no (default) never restart
--restart on-failure[:N] only after a non-zero exit, at most N times
--restart always after any exit, and when the daemon starts again
--restart unless-stopped like always, but a manual docker stop is rememberedWith always, an app that fails at startup is started again, fails again, and so on. Docker waits between attempts: 100 ms, then double each time, capped at one minute, and reset once the container has stayed up for 10 seconds. That is the crash loop. Restarting (1) means "the last run exited with code 1, and I am waiting to start it again".
The restart policy is not the bug. It just turns a loud failure into something that looks like flakiness.
The diagnosis path
1. How did it end, and how often?
$ docker inspect -f '{{.RestartCount}} {{.State.ExitCode}} {{.State.OOMKilled}}' shop-app-1
6 1 falseSix restarts, last exit code 1, not killed for memory. The exit code tells you which kind of problem to look for:
| code | meaning |
|---|---|
| 1 | the application failed - read its logs |
| 125 | docker run itself failed: bad flag, name taken, port in use |
| 126 | the command exists but cannot be executed |
| 127 | the command was not found in the image |
| 137 | SIGKILL (128+9): the OOM killer, docker kill, or a stop that timed out |
| 143 | SIGTERM (128+15): a normal docker stop |
128+N means "killed by signal N". Codes 125-127 mean the app never really started: look at the image and the command, not the app's logs. 137 with OOMKilled: true is memory, a different story: exit code 137 and OOMKilled.
2. Read the logs of every attempt
You cannot docker exec into a crash-looping container; it is never up long enough:
$ docker exec shop-app-1 env
Error response from daemon: container aaec3f38640f67c1c78f3f7e9aa5ba86a57afbeb1e9c1f12814a5293c99d68b1 is not runningBut docker logs keeps the output of every run, so the reason is there:
$ docker logs --tail 6 shop-app-1
2026-09-22T20:00:12.510+00:00 INFO 1 --- [shop] [ main] com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Starting...
2026-09-22T20:00:12.770+00:00 ERROR 1 --- [shop] [ main] com.zaxxer.hikari.pool.HikariPool : HikariPool-1 - Exception during pool initialization.
org.postgresql.util.PSQLException: Connection to localhost:5432 refused. Check that the hostname and port are correct and that the postmaster is accepting TCP/IP connections.
Caused by: java.net.ConnectException: Connection refused
...
2026-09-22T20:00:12.770+00:00 ERROR 1 --- [shop] [ main] o.s.boot.SpringApplication : Application run failed3. See the rhythm with docker events
$ docker events --since 1m --until now --filter container=shop-app-1
2026-09-22T20:00:03.000557000+00:00 container start aaec3f38640f... (image=shop-app, name=shop-app-1)
2026-09-22T20:00:05.330008270+00:00 container die aaec3f38640f... (exitCode=1, image=shop-app, name=shop-app-1)
2026-09-22T20:00:05.430800170+00:00 container start aaec3f38640f... (image=shop-app, name=shop-app-1)
2026-09-22T20:00:07.760251440+00:00 container die aaec3f38640f... (exitCode=1, image=shop-app, name=shop-app-1)
...Start, die about two seconds later with exit 1, start again. It fails during startup every time, so timing alone will not fix it.
4. Check the configuration it started with
$ docker inspect -f '{{range .Config.Env}}{{println .}}{{end}}' shop-app-1
...
DB_HOST=localhost
DB_PORT=5432localhost inside a container is that container, not the host and not the database container. Nothing listens on 5432 in the app's own network namespace, so every attempt is refused. Containers on the same Compose network reach each other by service name:
$ docker run --rm --network shop_default nicolaka/netshoot getent hosts db
172.18.0.2 dbThe second bug, hiding behind the first
Fix DB_HOST: db and the app usually starts. Not always: on a fresh volume, Postgres takes a few seconds to initialise, and the app can still race it and lose. That is the "works if we restart it a minute later" part.
depends_on: [db] only orders the container starts. It does not wait for the database to accept connections. Make the order real with a healthcheck and a condition:
services:
db:
image: postgres:16
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 2s
retries: 15
app:
build: .
environment:
DB_HOST: db
DB_PORT: "5432"
depends_on:
db:
condition: service_healthy
restart: alwaysThen prove it from a clean start:
$ docker compose down
[+] Running 3/3
✔ Container shop-db-1 Removed 0.8s
✔ Container shop-app-1 Removed 0.8s
✔ Network shop_default Removed 0.1s
$ docker compose up -d
[+] Running 3/3
✔ Network shop_default Created 0.1s
✔ Container shop-db-1 Healthy 1.5s
✔ Container shop-app-1 Started 1.6s
$ docker inspect -f '{{.RestartCount}}' shop-app-1
0RestartCount 0 after a fresh down + up is the proof. "It is running now" after six restarts is not.
Keeping it from coming back
- Make the app retry its dependencies. A service that dies on the first refused connection turns every database restart into an outage. Retry with back-off, and fail readiness meanwhile.
- Alert on restarts, not only on "down". A container with a climbing
RestartCountis up most of the time and broken all of the time. - Never use
localhostfor another container. Use the service name on a shared network. - In Kubernetes the same loop is called CrashLoopBackOff and is read the same way: exit code,
--previouslogs, events. And make sure PID 1 handles signals, or every stop becomes a 137 after the timeout: see zombies and PID 1.
Practise it
Incident: the app restarts every few seconds after every deploy (11.35) is this Compose project with both bugs in place. Restart policies, back-off, and what survives a daemon restart (11.8) lets you watch the delays grow.