How a failed memory allocation triggers it
Reading and interpreting the OOM score
Tuning oom_score_adj to protect processes
The oom_reaper kernel thread
Modern PSI and cgroup v2 memory controls
Hands-on commands to observe OOM behaviour
The Linux kernel OOM killer (Out-Of-Memory killer) is one of the most misunderstood parts of the kernel’s memory management subsystem. Every embedded systems engineer, kernel module author, and device driver developer eventually runs into a system that mysteriously kills a process under memory pressure — and the reason is almost always the OOM killer doing its job. This tutorial breaks down exactly how it works on current Linux kernels, in plain language, with hands-on commands you can try today. It’s part of our free Linux kernel programming course track, so if you’re new here, feel free to start from the beginning of the series using the navigation links above.
Prerequisites
Before diving into the Linux kernel OOM killer, you should be comfortable with:
- Basic C programming and pointers
- Fundamental Linux command-line usage (shell, redirection, permissions)
- A general idea of what “physical memory” vs “virtual memory” means
- Access to a Linux machine or VM where you can safely run
sudocommands (a disposable VM is strongly recommended for the hands-on section)
If you’re missing any of these, our free embedded systems course and free Linux device drivers course cover them from scratch — check the course index for the right starting point.
What Is the OOM Killer and Why Does the Kernel Need One?
Linux, like most modern operating systems, over-commits memory. When a process calls malloc() or a driver requests pages internally, the kernel usually promises the memory immediately without actually reserving physical RAM for it. Real physical pages are only handed out later, when the process actually touches that memory and a page fault occurs. This strategy — called memory overcommit — lets the system support many processes that each reserve more memory than they will ever really use.
The catch is that overcommitting can eventually catch up with the kernel. If enough processes touch their promised memory at the same time, the kernel may reach a point where it simply has no free physical pages left to hand out, and no way to reclaim any either. At that moment, the kernel has to make a decision: which process should be sacrificed to keep the rest of the system alive? That decision-making logic is the OOM killer.
Without an OOM killer, a memory exhaustion event could freeze the entire machine, since the kernel itself needs memory to keep running critical services. The OOM killer exists as a last line of defense, and it’s also the reason a poorly written fork bomb eventually gets shut down instead of crashing the whole box permanently.
How a Failed Allocation Triggers the OOM Killer
Every time a running program needs more physical memory, the request eventually reaches the kernel’s core page allocator. Under normal conditions the allocator finds free pages, possibly after reclaiming some clean cache pages, and the request succeeds silently. The OOM killer only gets involved in the rare case where the allocator has exhausted every reasonable option — direct reclaim, compaction, and swapping — and still cannot produce the requested pages.
This all runs in process context — meaning the very process that triggered the shortage is often the one running the kernel code that decides who dies. That’s part of what makes OOM killer debugging tricky: the “current” process at the time of the kill is rarely the actual culprit that caused the memory pressure over time.
Understanding the OOM Score
To choose a victim quickly during a memory crisis, the kernel maintains a running OOM score for every process, readable from /proc/<pid>/oom_score. The scale runs from 0 to 1000.
| OOM Score | Meaning |
|---|---|
| 0 | Process is using virtually none of the available memory |
| ~500 | Process is consuming a moderate-to-large share of available memory |
| 1000 | Process is effectively using all memory available to it — prime OOM target |
The kernel doesn’t blindly kill whichever process has the highest raw memory usage — it also applies built-in protections. Root-owned system processes, kernel threads, and processes holding certain hardware devices open are heavily favored to survive, because killing them could destabilize the whole system rather than free it up.
Tuning oom_score_adj to Protect (or Sacrifice) a Process
As an embedded or systems engineer, you’ll often want explicit control over which processes are allowed to die under memory pressure. The kernel exposes a per-process tunable for exactly this:
cat /proc/<pid>/oom_score_adj
The final score the kernel uses is simply:
net_oom_score = oom_score + oom_score_adj
| oom_score_adj Value | Effect |
|---|---|
| -1000 | Process is completely exempt from being OOM-killed |
| 0 (default) | No adjustment; pure usage-based scoring applies |
| 1000 | Process becomes an almost guaranteed OOM victim |
To set this value, you need root access:
echo -1000 | sudo tee /proc/1234/oom_score_adj
A convenient modern tool for reading and setting this in one shot is choom, part of util-linux:
choom -p 1234
choom -p 1234 -n -500
The oom_reaper Kernel Thread
Actually terminating a process isn’t instant — the victim may be holding locks or stuck in uninterruptible states, which can stall the reclaim of its memory. To keep the OOM path fast and reliable, the kernel runs a dedicated helper thread called oom_reaper. Its job is to asynchronously unmap and free the victim’s memory pages as quickly as possible, independent of whether the victim’s own thread manages to exit cleanly. This design has been standard in Linux for a long time now and remains the mechanism current kernels rely on to keep OOM recovery snappy.
Modern Kernel Improvements: PSI and cgroup v2
The classic OOM killer is reactive — it only acts after an allocation has already failed, which can mean the system was sluggish and thrashing for several seconds beforehand. Current kernels give you better tools to react earlier:
- Pressure Stall Information (PSI): exposed at
/proc/pressure/memory, PSI reports how much time tasks spend stalled waiting on memory, letting user-space daemons likesystemd-oomdorearlyoomintervene before the kernel’s last-resort killer has to. - cgroup v2 memory controller: with
memory.maxandmemory.oom.group, you can scope OOM behaviour to a specific container or service group instead of letting it pick any process system-wide — essential for predictable behaviour on embedded boards running multiple containerized workloads.
For embedded systems with tight RAM budgets, combining cgroup v2 limits with a user-space early-OOM daemon is now considered better practice than relying purely on the kernel’s reactive killer.
Hands-On: Observing OOM Behaviour Safely
Try these on a disposable VM, never on production hardware. First, check the current OOM score of a process by PID:
cat /proc/$(pgrep -n bash)/oom_score
Here’s a small original shell snippet to list the ten processes with the highest current OOM score on your system:
for pid in $(ls /proc | grep -E '^[0-9]+$'); do
score=$(cat /proc/$pid/oom_score 2>/dev/null)
name=$(cat /proc/$pid/comm 2>/dev/null)
[ -n "$score" ] && echo "$score $pid $name"
done | sort -n | tail -10
To watch memory pressure building in real time on a modern kernel:
watch -n1 cat /proc/pressure/memory
Real-World Use Cases
- Embedded boards with limited RAM: setting
oom_score_adjto protect a watchdog or safety-critical process from ever being killed. - Containerized services: using cgroup v2’s
memory.oom.groupso an entire misbehaving container is torn down together, rather than one process inside it. - Server workloads: deploying
systemd-oomdwith PSI thresholds to proactively kill runaway processes before user-facing latency spikes.
Common Mistakes and Troubleshooting
- Assuming the killed process caused the leak: the OOM killer’s victim is chosen by score at the moment of crisis, not necessarily the process that gradually caused memory pressure — always check
dmesghistory around the event. - Setting oom_score_adj to -1000 everywhere: this just moves the crisis elsewhere and can prevent the kernel from recovering at all if too many processes are protected.
- Forgetting swap accounting: a process with a huge number of swap entries can still be a strong OOM candidate even if its resident memory looks moderate.
- Ignoring cgroup limits: in containerized environments, a process can be OOM-killed by its cgroup’s own
memory.maxlong before system-wide memory is actually exhausted.
Best Practices
- Reserve
oom_score_adj = -1000only for genuinely critical processes (watchdogs, init, core drivers’ user-space helpers). - Monitor
/proc/pressure/memoryinstead of waiting for the kernel’s last-resort killer on latency-sensitive systems. - Use cgroup v2 memory limits to contain workloads instead of relying solely on global OOM behaviour.
- Always correlate an OOM kill with historical memory trends, not just the single dmesg snapshot.
Performance Considerations
Scanning every process to compute OOM scores has a real cost on systems with thousands of tasks, which is one reason modern kernels favor early intervention via PSI and cgroups over letting the system reach the reactive OOM path at all. The oom_reaper thread itself is designed to be lightweight, but frequent OOM events are still a strong signal that your memory budget, not just your tuning, needs attention.
Security Considerations
Because oom_score_adj can only be lowered by a privileged process once a lower value has been set (a non-root process can only raise it further), it can’t be used by an unprivileged attacker to make itself immune to being killed. Still, on multi-tenant systems it’s good practice to keep cgroup memory limits in place so a single misbehaving workload can’t create enough system-wide pressure to threaten unrelated processes in the first place.
Summary: Key Takeaways
- The OOM killer is the kernel’s last-resort defense when a memory allocation cannot be satisfied by any other means.
- Every process has an OOM score (0–1000) based on memory usage, adjustable per-process with
oom_score_adj. - The
oom_reaperthread reclaims a victim’s memory asynchronously to keep recovery fast. - Modern kernels add PSI and cgroup v2 controls so systems can react before the reactive OOM killer ever needs to run.
Conclusion
Understanding the Linux kernel OOM killer turns a scary, mysterious process termination into a predictable, tunable system behaviour. Whether you’re building embedded devices with a few hundred megabytes of RAM or managing containerized services on a server fleet, knowing how to read OOM scores, tune oom_score_adj, and lean on modern tools like PSI and cgroup v2 will save you hours of confusing debugging. This lesson is part of our ongoing free Linux kernel programming course — keep working through the series to build a complete, practical understanding of kernel memory management.
Frequently Asked Questions
1. What triggers the Linux kernel OOM killer?
It triggers only when the page allocator has exhausted reclaim, compaction, and swap options and still cannot satisfy a memory request.
2. How do I check a process’s OOM score?
Read /proc/<pid>/oom_score, or use the choom -p <pid> command for a more readable view.
3. Can I make a process completely immune to the OOM killer?
Yes, by setting its oom_score_adj to -1000 as root, though this should be reserved for truly critical processes.
4. What is the oom_reaper thread?
A kernel thread that asynchronously frees a victim process’s memory pages so OOM recovery isn’t delayed by a stuck victim.
5. Is the classic OOM killer still relevant on modern kernels?
Yes, it remains the final safety net, though PSI and cgroup v2 now let systems react to memory pressure earlier and more predictably.
6. Why did the OOM killer kill a process that wasn’t using much memory?
Built-in heuristics can protect high-usage but critical processes (root-owned, kernel threads, device-holding tasks), shifting the score-based decision toward a different candidate.
7. What’s the difference between oom_score and oom_score_adj?
oom_score is the kernel-computed, real-time usage-based score; oom_score_adj is a manual bias you add on top of it.
8. How does cgroup v2 change OOM behaviour for containers?
With memory.oom.group enabled, the kernel can kill every process in a container’s cgroup together instead of a single process, giving cleaner container-level recovery.
Explore more free lessons on Linux kernel memory management, device drivers, and embedded systems at EmbeddedPathashala.

2 Comments