15 min
Beginner to Intermediate
6.x (current LTS)
oom_score and badness
oom_score_adj
Reading kernel OOM logs
cgroup-aware OOM killer
PSI and early OOM daemons
Welcome back to our free Linux kernel development course. In the previous lecture we covered memory overcommit. Now we look at what happens on the other side: the moment physical memory and swap are truly exhausted and the kernel must forcibly reclaim memory by killing a process. This is the Linux OOM killer, one of the most misunderstood subsystems for anyone learning free Linux device drivers or free embedded systems development.
- Comfort with the memory overcommit concepts from the previous lecture
- Basic familiarity with
/proc,dmesg, and process IDs - Root or sudo access on a disposable test VM (never test OOM behavior on a production box)
The Out-Of-Memory killer, usually just called the OOM killer, is a kernel-level last resort. It only activates when the kernel genuinely cannot satisfy a memory request through any other means: no free physical pages, no reclaimable cache, and no more swap space. Rather than let the entire system freeze or crash, the kernel selects one process to sacrifice, frees its memory, and lets the rest of the system keep running.
The kernel does not kill processes at random. Every process is assigned a badness score, visible as oom_score in /proc. The score weighs mainly the resident memory a process is using, adjusted by a bias value you control yourself.
$ cat /proc/<pid>/oom_score
542
Roughly speaking, the process consuming the largest share of memory, relative to total available memory, ends up with the highest score and becomes the most likely candidate. This design intentionally favors killing the single biggest memory consumer rather than many small processes, since that reclaims the most memory in one action.
You can influence, but not fully override, the kernel’s choice using oom_score_adj, a per-process value ranging from -1000 to +1000.
| Value | Effect |
|---|---|
| -1000 | Process is completely exempt from being OOM-killed |
| 0 | Default, no adjustment |
| +1000 | Process is almost guaranteed to be picked first |
# Make a critical service much less likely to be killed
$ echo -900 | sudo tee /proc/$(pgrep my-critical-service)/oom_score_adj
# Make a disposable batch job the preferred victim
$ echo 800 | sudo tee /proc/$(pgrep batch-worker)/oom_score_adj
Systemd services can set this declaratively as well, which is the recommended approach on modern distributions instead of scripting /proc writes:
[Service]
OOMScoreAdjust=-900
Every OOM kill is logged to the kernel ring buffer. On a current 6.x kernel you will find it with:
$ dmesg -T | grep -i "out of memory"
$ journalctl -k | grep -i oom
The log entry reports the invoking process, the trigger allocation order, and the victim’s PID, memory footprint and computed score, which is exactly the data you need when investigating a production incident.
A major improvement that classic kernel textbooks written before kernel 4.19 do not cover is the cgroup-aware OOM killer. On modern container hosts running cgroup v2, you can set memory.oom.group on a cgroup so that when any process inside it triggers an OOM condition, the kernel kills the entire cgroup as a unit rather than a single arbitrary process.
# Enable group-wide OOM kill for a systemd service's cgroup
$ echo 1 | sudo tee /sys/fs/cgroup/system.slice/myapp.service/memory.oom.group
This matters enormously for microservices and container workloads: killing only one process out of a multi-process container often leaves the container in a broken, half-alive state, whereas group-wide killing lets the orchestrator cleanly restart the whole container.
Modern kernels expose Pressure Stall Information (PSI) through /proc/pressure/memory, giving a live percentage of time processes spend stalled waiting for memory. This enables user-space daemons such as systemd-oomd and earlyoom to react to memory pressure before the kernel’s own last-resort OOM killer has to step in, usually producing a much gentler and more predictable outcome.
$ cat /proc/pressure/memory
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
- Kubernetes nodes: rely on cgroup-aware OOM killing so an entire pod’s containers die together rather than leaving orphaned processes.
- Desktop Linux: systemd-oomd protects an interactive desktop session from freezing when a browser tab runs away with memory.
- Embedded boards: engineers often disable overcommit and set strict per-process
oom_score_adjvalues to protect safety-critical drivers and daemons. - Database servers: set a strongly negative
oom_score_adjon the database process while allowing batch or reporting jobs to be sacrificed first.
| Mistake | Why It Hurts | Fix |
|---|---|---|
| Setting oom_score_adj to -1000 on too many processes | Leaves the kernel with too few eligible victims, risking a full system hang | Protect only the truly critical process, not entire categories |
| Assuming the OOM killer always targets the “guilty” process | The victim is chosen by memory footprint, not by who caused the pressure | Always read the full dmesg trace to understand actual memory usage at kill time |
| Ignoring PSI metrics in monitoring | You lose early warning before a hard OOM kill occurs | Add /proc/pressure/memory to your monitoring dashboards |
- Use
OOMScoreAdjustin systemd unit files rather than manual/procwrites for anything persistent. - Enable
memory.oom.groupfor multi-process containers so the whole unit is reclaimed together. - Deploy
systemd-oomdorearlyoomon interactive and latency-sensitive systems to avoid full system stalls before the kernel intervenes. - Always correlate an OOM incident with
dmesg,journalctl -k, and PSI history, not memory usage alone.
Waiting for the kernel’s own OOM killer is the slowest and most disruptive path, since the system may already be thrashing under memory pressure by the time it triggers. Proactive PSI-based tools intervene sooner and generally cause far less disruption to unrelated processes.
Uncontrolled memory-hungry processes, whether malicious or buggy, can be used to trigger denial-of-service conditions by forcing the OOM killer to terminate unrelated critical services. Combining sane oom_score_adj policy with cgroup memory limits (covered in the overcommit lecture) closes this gap.
- The OOM killer is a last-resort kernel mechanism, not an error condition to fear.
- Victim selection is driven by
oom_score, which you can bias withoom_score_adj. - Modern cgroup v2 systems support group-wide OOM kills via
memory.oom.group. - PSI and user-space daemons like
systemd-oomdcan prevent hard kernel OOM kills entirely. - Always read
dmesgorjournalctl -klogs to understand exactly what happened during an OOM event.
The OOM killer completes the picture we started with memory overcommit: overcommit decides what gets promised, and the OOM killer decides what happens when those promises can no longer be kept. Understanding both together gives you the foundation to confidently tune memory behavior on any Linux system, from a tiny embedded board to a large Kubernetes cluster.
It triggers only when the kernel cannot satisfy a memory allocation after exhausting free memory, reclaimable caches, and swap space.
It computes a badness score, mainly based on resident memory usage, and adjusts it using each process’s oom_score_adj value.
Yes, by setting its oom_score_adj to -1000, though this should be reserved for truly critical processes only.
The kernel OOM killer is a last-resort in-kernel mechanism. systemd-oomd is a user-space daemon that watches PSI metrics and intervenes earlier, often avoiding a hard kernel OOM kill altogether.
Run dmesg -T | grep -i "out of memory" or journalctl -k | grep -i oom to see kernel log entries for past OOM events.
It is a cgroup v2 control that forces the kernel to kill every process in a cgroup together during an OOM event, which is essential for multi-process containers.
Yes. Swap gives the kernel extra room to reclaim memory before resorting to killing a process, so systems without swap tend to hit the OOM killer earlier under memory pressure.
Explore more free lectures on Linux kernel memory management and device driver development.

2 Comments