Docker Runtime & Networking: interview questions
The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 11 of the course.
A container keeps restarting. How do you find out why? Junior
A container lives exactly as long as its PID 1, so the question is what PID 1 did:
docker ps -a- statusRestarting (1)orExited (137): the last exit code.docker logs --tail 50 NAME 2>&1- what it said before dying (logs survive restarts;2>&1so stderr reaches grep).docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.RestartCount}}' NAME.docker events --since 30m --filter container=NAME- the timeline:oomright beforedieis a memory limit.
Then decode the exit code: 1 the app failed (read the logs - often a dependency like the database not reachable or not ready yet), 126/127 the command is not executable / not found, 137 SIGKILL (OOMKilled true = memory limit), 139 segfault, 143 SIGTERM. You cannot exec into a crash loop; to look inside, run the image with --entrypoint sh.
Also asked: What is the difference between a container and a virtual machine? · How do two containers talk to each other? · How do you persist data from a container?
What does exit code 137 mean for a container, and how do you find out what caused it? Junior
137 = 128 + 9: PID 1 was killed by SIGKILL. Two possible senders:
- A person or Docker:
docker kill, ordocker stopgiving up after the 10-second grace period because PID 1 ignored SIGTERM. - The kernel's OOM killer: the container went over its memory limit.
Tell them apart:
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' api
OOMKilled: true means memory. Then docker logs api for what it printed before.
The rest of the table: 0 finished normally, 1 the app failed (logs), 126 not executable, 127 command not found (the container never left Created), 139 segfault (128 + 11), 143 SIGTERM (128 + 15). And remember docker ps hides stopped containers: docker ps -a.
Also asked: A container is not in docker ps. Where did it go? · Why does a container started from the ubuntu image exit immediately? · What is the difference between docker create, docker start and docker run?
Learn it: 11.1 Running containers, and reading how they ended
How do you look at the logs of a container, including one that already stopped? Junior
docker logs NAME prints what PID 1 wrote to stdout and stderr - and it works on exited containers too, so it is the first command after docker ps -a.
--tail 50the last lines,-ffollow,--since 10ma time window,-ttimestamps.- Grep with
docker logs web 2>&1 | grep ...- Docker keeps the two streams apart, and a pipe only carries stdout, so error lines (often on stderr) skip grep without2>&1. - An app that logs to a file inside the container has empty
docker logs; images like nginx link their log files to/dev/stdoutand/dev/stderrfor that reason.
The rest of the toolkit: docker inspect -f for settings and state, docker events for the timeline, docker stats for usage, docker exec into a running one, docker cp to copy a file out.
Also asked: What is docker inspect good for, and which fields do you check most? · How do you get a shell in a running container? · How would you reconstruct why a container restarted overnight?
Learn it: 11.5 The debugging toolkit: logs, exec, inspect, events, stats
What happens when you run docker stop? Junior
Docker sends the stop signal (SIGTERM, or the image's STOPSIGNAL) to the container's PID 1, waits the grace period (10 s, -t to change it), then sends SIGKILL.
What it means depends on PID 1:
- It handles SIGTERM - it shuts down cleanly and exits 0 or 143, in milliseconds.
- It has no handler - the kernel ignores default signal actions for PID 1 of a PID namespace, so nothing happens for 10 seconds, then SIGKILL: exit 137, requests dropped. A shell-form ENTRYPOINT (sh as PID 1) does the same.
Fixes: a real SIGTERM handler, exec form / exec "$@", or --init (tini as PID 1 forwards signals; your app becomes PID 2 where default actions apply).
docker kill sends SIGKILL immediately.
Also asked: What is the difference between the restart policies always and unless-stopped? · What is the difference between restarting and recreating a container? · Why is it a problem when a shell script is PID 1 in a container?
Learn it: 11.7 Stopping, killing and restarting
How do you limit memory and CPU for a container, and what happens when it exceeds them? Junior
They are the same cgroup files as a systemd unit's MemoryMax= / CPUQuota=:
-m 512mwritesmemory.max. Go over it and the kernel's cgroup OOM killer kills a process in the container: exit 137,OOMKilled: true.--cpus 1.5writescpu.max(150ms of CPU per 100ms period). Use it up and the process is throttled - paused until the next period - not killed.nr_throttledincpu.statshows it; for a latency-sensitive service that means small unexplained pauses.--pids-limit 200caps processes and threads.
docker stats shows usage against the limit; docker update changes it live (the real fix goes into the run command). Trap: free inside the container shows the host's memory; the truth is memory.max. For Java: the JVM takes 25% of the limit as heap by default; size the limit at heap x 1.3-1.5, or set -XX:MaxRAMPercentage=75.
Also asked: How would you size the memory limit for a Java service in a container? · What is the difference between a JVM OutOfMemoryError and a cgroup OOM kill? · What is CPU throttling?
How do two containers talk to each other in Docker? Junior
Put them on the same user-defined network and use the container name:
docker network create appnet
docker run -d --name db --network appnet -e POSTGRES_PASSWORD=x postgres:16
docker run -d --name api --network appnet api:1 # connects to db:5432
A user-defined bridge runs Docker's embedded DNS at 127.0.0.11, so names resolve. The default bridge does not (bad address 'db'), and two different networks cannot reach each other at all (timeout).
Three classic mistakes:
- Using the published host port between containers - use the container port (
db:5432, not 15432). localhostinside a container means that container, not the host or its neighbours.- An app that binds
127.0.0.1inside its container is unreachable from anywhere else - bind0.0.0.0.
Container IPs change on restart; the name follows, a hard-coded IP does not.
Also asked: What does -p 8080:80 do, and what does -p 127.0.0.1:8080:80 change? · Why does a published port give "connection reset" when the app is up? · What is host networking, and when would you use it?
Learn it: 11.15 Networks: bridges, DNS, published ports, and localhost
Service A cannot connect to service B. What do you check? Junior
Four steps, in order, from A's network namespace (not from the host, which has different DNS and sees published ports):
docker run --rm -it --network container:A nicolaka/netshoot
- Name -
dig B. NXDOMAIN from127.0.0.11= B is on another network, A is on the default bridge, or the name is wrong. Compare networks withdocker inspect -f '{{json .NetworkSettings.Networks}}' A B; fix withdocker network connect. - Route - same network? A timeout means dropped packets: isolated networks or a firewall.
- Port -
nc -zv B 8080.Connection refused= reached, nothing listening there: checkss -tlnpin B's namespace (bound to127.0.0.1? a different port?). - Application -
curl -sv http://B:8080/health: status line and response.
Use the container port, never the published one. The error text usually names the layer.
Also asked: What do connection refused and connection timed out tell you? · How do you debug networking in a container that has no shell or tools? · Why should you test from the caller's network namespace and not from the host?
Learn it: 11.20 Debugging container networking, layer by layer
How do you persist data from a container? Junior
Anything written to the container's writable layer is destroyed with the container. Put data in a mount:
- Named volume -
-v pgdata:/var/lib/postgresql/data. Docker-managed, under/var/lib/docker/volumes/, survivesdocker rm. An empty volume is initialised from the image, ownership included. The default for databases. - Bind mount -
-v /srv/batch-out:/out, a host directory. It hides whatever the image had at that path, and keeps host ownership. - tmpfs - in memory, gone on stop; for scratch files.
The permission trap: without user namespaces, uid 1000 inside is uid 1000 outside. Compare ls -lnd /host/path with docker exec C id, then chown the host directory to the container's uid or run with -u "$(id -u):$(id -g)" - never chmod 777.
And a typo in the volume name gives you a new, empty volume: "the database is empty after the redeploy".
Also asked: A non-root container cannot write to a bind-mounted directory. How do you fix it? · What is the difference between a named volume and a bind mount? · How do you back up a Docker volume?
Learn it: 11.22 Volumes, bind mounts, tmpfs, and the UID problem
A Docker host's root disk is full. How do you find and fix the cause? Junior
The usual culprit is a container's log. The default json-file logging driver writes everything PID 1 prints to /var/lib/docker/containers/<id>/<id>-json.log, and by default never rotates it. docker system df does not count log files, which is why it "looks fine".
df -hto confirm, thensudo sh -c 'du -sh /var/lib/docker/containers/*' | sort -h | tail -3- the directory name is the container ID.- Set rotation as the default in
/etc/docker/daemon.json("log-opts": {"max-size": "10m", "max-file": "3"}), validate it withjq .(bad JSON stops dockerd),sudo systemctl restart docker. - Recreate the chatty container - existing containers keep the log settings they were created with.
Do not rm the log file: dockerd keeps it open, so the space is not freed. Then docker image prune / container prune for the rest.
Also asked: Where does docker logs get its data from? · What does docker system df show, and what does it miss? · What is a logging driver?
Learn it: 11.26 Logging drivers, and the log file that fills the disk
How would you harden a docker run command for a typical web service? Junior
Each flag removes something an attacker could use:
docker run -d --name api \
--user 10001:10001 \
--read-only --tmpfs /tmp \
--cap-drop ALL \
--security-opt no-new-privileges \
-m 512m --cpus 1 --pids-limit 200 \
-p 127.0.0.1:8080:8080 \
api:1
--user- not root; a non-root process has no effective capabilities.--read-only- nobody can drop a binary or change the app; tmpfs only where it must write.--cap-drop ALL- root in a default container keeps 14 capabilities; most services need none (add one back with--cap-add).no-new-privileges- no escalation through setuid binaries.- Limits, so it cannot starve the host; the port published on localhost only.
Never --privileged and never mount /var/run/docker.sock into an app - both mean root on the host.
Also asked: Why should containers not run as root? · What are Linux capabilities? · Why is mounting the Docker socket into a container dangerous?
Learn it: 11.29 Container security: capabilities, read-only, and the socket
What is Docker Compose, and how do you make a service start only after its database is ready? Junior
Compose describes several containers as one project in compose.yaml: each service is what you would pass to docker run (image, environment, volumes, ports, restart, healthcheck). docker compose up -d creates a project network (service names resolve by DNS, so the app reaches db:5432), the volumes and the containers. It runs on one host: for local development, tests and small deployments.
depends_on: [db] only orders the start. To wait for readiness, give the database a healthcheck and use a condition:
depends_on:
db:
condition: service_healthy
with healthcheck: test: ["CMD-SHELL", "pg_isready -U postgres"] on db. Still not enough on its own: the app should retry its connections, because the database can restart later when Compose is not ordering anything.
Also asked: What is the difference between the .env file and env_file in Compose? · What does docker compose down -v do? · Why does docker compose restart not apply changes from the compose file?
Learn it: 11.31 Docker Compose: several containers as one project
Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.