OnCallReady

Memory & OOM: interview questions

The question you are most likely to get for each topic, a model answer, and what else comes up. From chapter 5 of the course.

A service restarts every few hours with exit code 137. What is your diagnosis path? Junior

  1. Decode it: 137 = 128 + 9, the process got SIGKILL. With a memory limit that is usually the OOM killer; it can also be systemd's SIGKILL after a stop timeout, or someone's kill -9.
  2. systemd's view: systemctl status svc - Result: oom-kill? journalctl -u svc | grep -i oom shows "A process of this unit has been killed by the OOM killer".
  3. The kernel's view: journalctl -k | grep -i oom (or sudo dmesg -T). CONSTRAINT_MEMCG with oom_memcg=... = the service hit its own limit, the machine was fine; CONSTRAINT_NONE + global_oom = the whole machine ran out.
  4. The proof: cat /sys/fs/cgroup/system.slice/svc.service/memory.events - a non-zero oom_kill. Compare memory.current / memory.peak with memory.max.
  5. Why: a leak, a load spike, or a limit too close to normal use - the classic is a Java service with MemoryMax equal to -Xmx.

Also asked: Why is almost no free memory on a Linux server usually not a problem? · What is the difference between a host OOM and a cgroup OOM? · How would you size a memory limit for a Java service?

What is the difference between RSS and VSZ? Junior

RSS over-counts shared pages: libc is in RAM once but in every process's RSS, so do not add RSS up - forty copies of a worker can "use" more than the machine has. For a group of processes, sum PSS (from /proc/PID/smaps_rollup), which divides shared pages fairly. For one process, /proc/PID/status splits RSS into RssAnon (heap and stacks, the costly part) and RssFile (droppable), and VmHWM shows its peak.

Also asked: Forty worker processes each show 300 MB RSS on an 8 GB box. How is that possible? · What is the difference between anonymous and file-backed memory? · How do you see how high a process's memory went, after the spike has passed?

Learn it: 5.1 RSS, VSZ and shared

Why is almost no free memory on a Linux server usually not a problem? Junior

Because Linux fills idle RAM with the page cache - copies of files it has read, kept in case they are read again - and throws those pages out the moment a program needs the memory. Unused RAM is wasted RAM.

In free -m: free is memory doing nothing (low is fine), buff/cache is mostly that cache, and available is the kernel's estimate of what a new program could get right now, cache included. That is the number to read and to alert on. 2 GB free with 4.5 GB available is a healthy box.

Signs of a real shortage: low available, and in vmstat 1 5 sustained si/so (swapping in and out - thrashing). Dropping the cache with echo 3 | sudo tee /proc/sys/vm/drop_caches proves it: free shoots up, available barely moves. Never do that as a "fix". One exception: shared (tmpfs) is counted in buff/cache but is not reclaimable.

Also asked: Explain every column of free -m. · What is swappiness, and why would you lower it? · What is memory overcommit, and why does it mean malloc rarely fails?

Learn it: 5.3 free, available, and the page cache

How do you find out whether a process was killed by the OOM killer? Junior

The OOM killer is part of the kernel, so the evidence is in the kernel log:

The kernel picks by size (/proc/PID/oom_score), not by guilt.

Also asked: How does the OOM killer choose which process to kill? · How would you protect sshd from the OOM killer? · Why is oom_score_adj=-1000 on an ordinary service risky?

Learn it: 5.7 The OOM killer

How does the OOM killer choose which process to kill? Junior

It kills the eligible process with the highest badness, the kernel's score:

badness = rss + swapents + pgtables_bytes/4096   (pages)
        + oom_score_adj x totalpages / 1000

So it is mostly size: resident pages plus swapped pages. totalpages is RAM + swap for a host OOM, or the cgroup's limit for a cgroup OOM - where only the processes inside that cgroup are candidates. oom_score_adj is the thumb on the scale: +500 adds half of all memory to the score, -1000 makes a process ineligible (sshd in the report). /proc/PID/oom_score shows the live, normalised version for every process.

The consequence: the biggest process dies, not the guilty one. In the report, read the first line (the invoker - who asked for a page), the Tasks table (who was there, in pages), the oom-kill: line (constraint and task_memcg, which maps the victim to its service) and "Killed process".

Also asked: What does "python3 invoked oom-killer" tell you, and what does it not? · How do you map an OOM kill in the kernel log back to a systemd service? · How do you tell a host OOM from a cgroup OOM in the report?

Learn it: 5.8 Reading an OOM report line by line

What does exit code 137 mean for a service, and how do you confirm the cause? Junior

137 = 128 + 9: the process was killed by SIGKILL. It does not say who sent it. With a memory limit it is usually the cgroup OOM killer - the service went over its MemoryMax while the machine may have had gigabytes free. The other common sender is systemd escalating to SIGKILL after a stop timeout.

To confirm:

Try it safely with sudo systemd-run --scope -p MemoryMax=200M memhog 400M; echo $? - it prints 137.

Also asked: What is the difference between MemoryMax and MemoryHigh? · How do you set and verify a memory limit on a running service? · Why does /sys/fs/cgroup/memory.max not exist on the host?

Learn it: 5.11 cgroup OOM and exit 137

How would you size a memory limit for a Java service? Junior

First, know that the heap is only one part. A JVM's RSS is the heap (what -Xmx limits) plus metaspace, a stack per thread, the code cache, direct buffers, GC bookkeeping and native libraries. A JVM with -Xmx1g sits at 1.3-1.5 GB of RSS, so a limit equal to -Xmx is an OOM kill waiting for the first busy minute.

Then, in order of preference:

  1. Let the heap follow the limit: set MemoryMax= on the unit, drop -Xmx, use -XX:MaxRAMPercentage=75. UseContainerSupport makes the JVM read its cgroup limit as "the RAM".
  2. Fixed heap, derived limit: limit = heap x 1.3-1.5, or better heap + measured non-heap (Native Memory Tracking, jcmd PID VM.native_memory summary) + 10-20%.
  3. Watch the unit's memory.current under real load.

How each failure looks: exit 137 with nothing in the app log = the kernel (limit too small); OutOfMemoryError: Java heap space in the log = the heap itself.

Also asked: A Java service is OOM-killed but its heap never goes above 60%. What is happening? · What does UseContainerSupport do? · Without -Xmx, how much heap does the JVM pick?

Learn it: 5.13 JVM memory: what the heap is not

Practise these answers with flashcards and labs Free, in your browser - a real Ubuntu terminal to try it in, with missions that check your work.