OnCallReady

Docker Runtime & Networking: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 11 of the course.

A container keeps restarting. How do you find out why? Junior

A container lives exactly as long as its PID 1, so the question is what PID 1 did:

  1. docker ps -a - status Restarting (1) or Exited (137): the last exit code.
  2. docker logs --tail 50 NAME 2>&1 - what it said before dying (logs survive restarts; 2>&1 so stderr reaches grep).
  3. docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}} {{.RestartCount}}' NAME.
  4. docker events --since 30m --filter container=NAME - the timeline: oom right before die is a memory limit.

Then decode the exit code: 1 the app failed (read the logs - often a dependency like the database not reachable or not ready yet), 126/127 the command is not executable / not found, 137 SIGKILL (OOMKilled true = memory limit), 139 segfault, 143 SIGTERM. You cannot exec into a crash loop; to look inside, run the image with --entrypoint sh.

Also asked: What is the difference between a container and a virtual machine? · How do two containers talk to each other? · How do you persist data from a container?

What does exit code 137 mean for a container, and how do you find out what caused it? Junior

137 = 128 + 9: PID 1 was killed by SIGKILL. Two possible senders:

Tell them apart:

docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' api

OOMKilled: true means memory. Then docker logs api for what it printed before.

The rest of the table: 0 finished normally, 1 the app failed (logs), 126 not executable, 127 command not found (the container never left Created), 139 segfault (128 + 11), 143 SIGTERM (128 + 15). And remember docker ps hides stopped containers: docker ps -a.

Also asked: A container is not in docker ps. Where did it go? · Why does a container started from the ubuntu image exit immediately? · What is the difference between docker create, docker start and docker run?

Learn it: 11.1 Running containers, and reading how they ended

How do you look at the logs of a container, including one that already stopped? Junior

docker logs NAME prints what PID 1 wrote to stdout and stderr - and it works on exited containers too, so it is the first command after docker ps -a.

The rest of the toolkit: docker inspect -f for settings and state, docker events for the timeline, docker stats for usage, docker exec into a running one, docker cp to copy a file out.

Also asked: What is docker inspect good for, and which fields do you check most? · How do you get a shell in a running container? · How would you reconstruct why a container restarted overnight?

Learn it: 11.5 The debugging toolkit: logs, exec, inspect, events, stats

What happens when you run docker stop? Junior

Docker sends the stop signal (SIGTERM, or the image's STOPSIGNAL) to the container's PID 1, waits the grace period (10 s, -t to change it), then sends SIGKILL.

What it means depends on PID 1:

Fixes: a real SIGTERM handler, exec form / exec "$@", or --init (tini as PID 1 forwards signals; your app becomes PID 2 where default actions apply).

docker kill sends SIGKILL immediately.

Also asked: What is the difference between the restart policies always and unless-stopped? · What is the difference between restarting and recreating a container? · Why is it a problem when a shell script is PID 1 in a container?

Learn it: 11.7 Stopping, killing and restarting

How do you limit memory and CPU for a container, and what happens when it exceeds them? Junior

They are the same cgroup files as a systemd unit's MemoryMax= / CPUQuota=:

docker stats shows usage against the limit; docker update changes it live (the real fix goes into the run command). Trap: free inside the container shows the host's memory; the truth is memory.max. For Java: the JVM takes 25% of the limit as heap by default; size the limit at heap x 1.3-1.5, or set -XX:MaxRAMPercentage=75.

Also asked: How would you size the memory limit for a Java service in a container? · What is the difference between a JVM OutOfMemoryError and a cgroup OOM kill? · What is CPU throttling?

Learn it: 11.10 Memory, CPU and PIDs: the cgroups again

How do two containers talk to each other in Docker? Junior

Put them on the same user-defined network and use the container name:

docker network create appnet
docker run -d --name db --network appnet -e POSTGRES_PASSWORD=x postgres:16
docker run -d --name api --network appnet api:1      # connects to db:5432

A user-defined bridge runs Docker's embedded DNS at 127.0.0.11, so names resolve. The default bridge does not (bad address 'db'), and two different networks cannot reach each other at all (timeout).

Three classic mistakes:

Container IPs change on restart; the name follows, a hard-coded IP does not.

Also asked: What does -p 8080:80 do, and what does -p 127.0.0.1:8080:80 change? · Why does a published port give "connection reset" when the app is up? · What is host networking, and when would you use it?

Learn it: 11.15 Networks: bridges, DNS, published ports, and localhost

Service A cannot connect to service B. What do you check? Junior

Four steps, in order, from A's network namespace (not from the host, which has different DNS and sees published ports):

docker run --rm -it --network container:A nicolaka/netshoot
  1. Name - dig B. NXDOMAIN from 127.0.0.11 = B is on another network, A is on the default bridge, or the name is wrong. Compare networks with docker inspect -f '{{json .NetworkSettings.Networks}}' A B; fix with docker network connect.
  2. Route - same network? A timeout means dropped packets: isolated networks or a firewall.
  3. Port - nc -zv B 8080. Connection refused = reached, nothing listening there: check ss -tlnp in B's namespace (bound to 127.0.0.1? a different port?).
  4. Application - curl -sv http://B:8080/health: status line and response.

Use the container port, never the published one. The error text usually names the layer.

Also asked: What do connection refused and connection timed out tell you? · How do you debug networking in a container that has no shell or tools? · Why should you test from the caller's network namespace and not from the host?

Learn it: 11.20 Debugging container networking, layer by layer

How do you persist data from a container? Junior

Anything written to the container's writable layer is destroyed with the container. Put data in a mount:

The permission trap: without user namespaces, uid 1000 inside is uid 1000 outside. Compare ls -lnd /host/path with docker exec C id, then chown the host directory to the container's uid or run with -u "$(id -u):$(id -g)" - never chmod 777.

And a typo in the volume name gives you a new, empty volume: "the database is empty after the redeploy".

Also asked: A non-root container cannot write to a bind-mounted directory. How do you fix it? · What is the difference between a named volume and a bind mount? · How do you back up a Docker volume?

Learn it: 11.22 Volumes, bind mounts, tmpfs, and the UID problem

A Docker host's root disk is full. How do you find and fix the cause? Junior

The usual culprit is a container's log. The default json-file logging driver writes everything PID 1 prints to /var/lib/docker/containers/<id>/<id>-json.log, and by default never rotates it. docker system df does not count log files, which is why it "looks fine".

  1. df -h to confirm, then sudo sh -c 'du -sh /var/lib/docker/containers/*' | sort -h | tail -3 - the directory name is the container ID.
  2. Set rotation as the default in /etc/docker/daemon.json ("log-opts": {"max-size": "10m", "max-file": "3"}), validate it with jq . (bad JSON stops dockerd), sudo systemctl restart docker.
  3. Recreate the chatty container - existing containers keep the log settings they were created with.

Do not rm the log file: dockerd keeps it open, so the space is not freed. Then docker image prune / container prune for the rest.

Also asked: Where does docker logs get its data from? · What does docker system df show, and what does it miss? · What is a logging driver?

Learn it: 11.26 Logging drivers, and the log file that fills the disk

How would you harden a docker run command for a typical web service? Junior

Each flag removes something an attacker could use:

docker run -d --name api \
  --user 10001:10001 \
  --read-only --tmpfs /tmp \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  -m 512m --cpus 1 --pids-limit 200 \
  -p 127.0.0.1:8080:8080 \
  api:1

Never --privileged and never mount /var/run/docker.sock into an app - both mean root on the host.

Also asked: Why should containers not run as root? · What are Linux capabilities? · Why is mounting the Docker socket into a container dangerous?

Learn it: 11.29 Container security: capabilities, read-only, and the socket

What is Docker Compose, and how do you make a service start only after its database is ready? Junior

Compose describes several containers as one project in compose.yaml: each service is what you would pass to docker run (image, environment, volumes, ports, restart, healthcheck). docker compose up -d creates a project network (service names resolve by DNS, so the app reaches db:5432), the volumes and the containers. It runs on one host: for local development, tests and small deployments.

depends_on: [db] only orders the start. To wait for readiness, give the database a healthcheck and use a condition:

depends_on:
  db:
    condition: service_healthy

with healthcheck: test: ["CMD-SHELL", "pg_isready -U postgres"] on db. Still not enough on its own: the app should retry its connections, because the database can restart later when Compose is not ordering anything.

Also asked: What is the difference between the .env file and env_file in Compose? · What does docker compose down -v do? · Why does docker compose restart not apply changes from the compose file?

Learn it: 11.31 Docker Compose: several containers as one project

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.