What Is oom_reaper in the Linux Kernel? – Linux Device Drivers Coaching in Hyderabad

Linux Kernel OOM Killer Explained: Complete Tutorial
A beginner-friendly, updated-for-modern-kernels guide to how Linux handles out-of-memory situations
Difficulty: Beginner–Intermediate
Read Time: 14 min
Kernel Version: 6.x

What You Will Learn
Why the Linux kernel OOM killer exists
How a failed memory allocation triggers it
Reading and interpreting the OOM score
Tuning oom_score_adj to protect processes
The oom_reaper kernel thread
Modern PSI and cgroup v2 memory controls
Hands-on commands to observe OOM behaviour

The Linux kernel OOM killer (Out-Of-Memory killer) is one of the most misunderstood parts of the kernel’s memory management subsystem. Every embedded systems engineer, kernel module author, and device driver developer eventually runs into a system that mysteriously kills a process under memory pressure — and the reason is almost always the OOM killer doing its job. This tutorial breaks down exactly how it works on current Linux kernels, in plain language, with hands-on commands you can try today. It’s part of our free Linux kernel programming course track, so if you’re new here, feel free to start from the beginning of the series using the navigation links above.

Prerequisites

Before diving into the Linux kernel OOM killer, you should be comfortable with:

  • Basic C programming and pointers
  • Fundamental Linux command-line usage (shell, redirection, permissions)
  • A general idea of what “physical memory” vs “virtual memory” means
  • Access to a Linux machine or VM where you can safely run sudo commands (a disposable VM is strongly recommended for the hands-on section)

If you’re missing any of these, our free embedded systems course and free Linux device drivers course cover them from scratch — check the course index for the right starting point.

What Is the OOM Killer and Why Does the Kernel Need One?

Linux, like most modern operating systems, over-commits memory. When a process calls malloc() or a driver requests pages internally, the kernel usually promises the memory immediately without actually reserving physical RAM for it. Real physical pages are only handed out later, when the process actually touches that memory and a page fault occurs. This strategy — called memory overcommit — lets the system support many processes that each reserve more memory than they will ever really use.

The catch is that overcommitting can eventually catch up with the kernel. If enough processes touch their promised memory at the same time, the kernel may reach a point where it simply has no free physical pages left to hand out, and no way to reclaim any either. At that moment, the kernel has to make a decision: which process should be sacrificed to keep the rest of the system alive? That decision-making logic is the OOM killer.

Without an OOM killer, a memory exhaustion event could freeze the entire machine, since the kernel itself needs memory to keep running critical services. The OOM killer exists as a last line of defense, and it’s also the reason a poorly written fork bomb eventually gets shut down instead of crashing the whole box permanently.

How a Failed Allocation Triggers the OOM Killer

Every time a running program needs more physical memory, the request eventually reaches the kernel’s core page allocator. Under normal conditions the allocator finds free pages, possibly after reclaiming some clean cache pages, and the request succeeds silently. The OOM killer only gets involved in the rare case where the allocator has exhausted every reasonable option — direct reclaim, compaction, and swapping — and still cannot produce the requested pages.

Simplified OOM Trigger Flow
Process touches new memory
→
Page allocator tries to find free pages
→
Reclaim & swap attempted, still not enough
→
OOM killer invoked
→
Victim process selected & reaped

This all runs in process context — meaning the very process that triggered the shortage is often the one running the kernel code that decides who dies. That’s part of what makes OOM killer debugging tricky: the “current” process at the time of the kill is rarely the actual culprit that caused the memory pressure over time.

Understanding the OOM Score

To choose a victim quickly during a memory crisis, the kernel maintains a running OOM score for every process, readable from /proc/<pid>/oom_score. The scale runs from 0 to 1000.

OOM Score Meaning
0 Process is using virtually none of the available memory
~500 Process is consuming a moderate-to-large share of available memory
1000 Process is effectively using all memory available to it — prime OOM target

The kernel doesn’t blindly kill whichever process has the highest raw memory usage — it also applies built-in protections. Root-owned system processes, kernel threads, and processes holding certain hardware devices open are heavily favored to survive, because killing them could destabilize the whole system rather than free it up.

Tuning oom_score_adj to Protect (or Sacrifice) a Process

As an embedded or systems engineer, you’ll often want explicit control over which processes are allowed to die under memory pressure. The kernel exposes a per-process tunable for exactly this:

cat /proc/<pid>/oom_score_adj

The final score the kernel uses is simply:

net_oom_score = oom_score + oom_score_adj
oom_score_adj Value Effect
-1000 Process is completely exempt from being OOM-killed
0 (default) No adjustment; pure usage-based scoring applies
1000 Process becomes an almost guaranteed OOM victim

To set this value, you need root access:

echo -1000 | sudo tee /proc/1234/oom_score_adj

A convenient modern tool for reading and setting this in one shot is choom, part of util-linux:

choom -p 1234
choom -p 1234 -n -500

The oom_reaper Kernel Thread

Actually terminating a process isn’t instant — the victim may be holding locks or stuck in uninterruptible states, which can stall the reclaim of its memory. To keep the OOM path fast and reliable, the kernel runs a dedicated helper thread called oom_reaper. Its job is to asynchronously unmap and free the victim’s memory pages as quickly as possible, independent of whether the victim’s own thread manages to exit cleanly. This design has been standard in Linux for a long time now and remains the mechanism current kernels rely on to keep OOM recovery snappy.

Modern Kernel Improvements: PSI and cgroup v2

The classic OOM killer is reactive — it only acts after an allocation has already failed, which can mean the system was sluggish and thrashing for several seconds beforehand. Current kernels give you better tools to react earlier:

  • Pressure Stall Information (PSI): exposed at /proc/pressure/memory, PSI reports how much time tasks spend stalled waiting on memory, letting user-space daemons like systemd-oomd or earlyoom intervene before the kernel’s last-resort killer has to.
  • cgroup v2 memory controller: with memory.max and memory.oom.group, you can scope OOM behaviour to a specific container or service group instead of letting it pick any process system-wide — essential for predictable behaviour on embedded boards running multiple containerized workloads.

For embedded systems with tight RAM budgets, combining cgroup v2 limits with a user-space early-OOM daemon is now considered better practice than relying purely on the kernel’s reactive killer.

Hands-On: Observing OOM Behaviour Safely

Try these on a disposable VM, never on production hardware. First, check the current OOM score of a process by PID:

cat /proc/$(pgrep -n bash)/oom_score

Here’s a small original shell snippet to list the ten processes with the highest current OOM score on your system:

for pid in $(ls /proc | grep -E '^[0-9]+$'); do
  score=$(cat /proc/$pid/oom_score 2>/dev/null)
  name=$(cat /proc/$pid/comm 2>/dev/null)
  [ -n "$score" ] && echo "$score $pid $name"
done | sort -n | tail -10

To watch memory pressure building in real time on a modern kernel:

watch -n1 cat /proc/pressure/memory

Real-World Use Cases

  • Embedded boards with limited RAM: setting oom_score_adj to protect a watchdog or safety-critical process from ever being killed.
  • Containerized services: using cgroup v2’s memory.oom.group so an entire misbehaving container is torn down together, rather than one process inside it.
  • Server workloads: deploying systemd-oomd with PSI thresholds to proactively kill runaway processes before user-facing latency spikes.

Common Mistakes and Troubleshooting

  • Assuming the killed process caused the leak: the OOM killer’s victim is chosen by score at the moment of crisis, not necessarily the process that gradually caused memory pressure — always check dmesg history around the event.
  • Setting oom_score_adj to -1000 everywhere: this just moves the crisis elsewhere and can prevent the kernel from recovering at all if too many processes are protected.
  • Forgetting swap accounting: a process with a huge number of swap entries can still be a strong OOM candidate even if its resident memory looks moderate.
  • Ignoring cgroup limits: in containerized environments, a process can be OOM-killed by its cgroup’s own memory.max long before system-wide memory is actually exhausted.

Best Practices

  • Reserve oom_score_adj = -1000 only for genuinely critical processes (watchdogs, init, core drivers’ user-space helpers).
  • Monitor /proc/pressure/memory instead of waiting for the kernel’s last-resort killer on latency-sensitive systems.
  • Use cgroup v2 memory limits to contain workloads instead of relying solely on global OOM behaviour.
  • Always correlate an OOM kill with historical memory trends, not just the single dmesg snapshot.

Performance Considerations

Scanning every process to compute OOM scores has a real cost on systems with thousands of tasks, which is one reason modern kernels favor early intervention via PSI and cgroups over letting the system reach the reactive OOM path at all. The oom_reaper thread itself is designed to be lightweight, but frequent OOM events are still a strong signal that your memory budget, not just your tuning, needs attention.

Security Considerations

Because oom_score_adj can only be lowered by a privileged process once a lower value has been set (a non-root process can only raise it further), it can’t be used by an unprivileged attacker to make itself immune to being killed. Still, on multi-tenant systems it’s good practice to keep cgroup memory limits in place so a single misbehaving workload can’t create enough system-wide pressure to threaten unrelated processes in the first place.

Summary: Key Takeaways

  • The OOM killer is the kernel’s last-resort defense when a memory allocation cannot be satisfied by any other means.
  • Every process has an OOM score (0–1000) based on memory usage, adjustable per-process with oom_score_adj.
  • The oom_reaper thread reclaims a victim’s memory asynchronously to keep recovery fast.
  • Modern kernels add PSI and cgroup v2 controls so systems can react before the reactive OOM killer ever needs to run.

Conclusion

Understanding the Linux kernel OOM killer turns a scary, mysterious process termination into a predictable, tunable system behaviour. Whether you’re building embedded devices with a few hundred megabytes of RAM or managing containerized services on a server fleet, knowing how to read OOM scores, tune oom_score_adj, and lean on modern tools like PSI and cgroup v2 will save you hours of confusing debugging. This lesson is part of our ongoing free Linux kernel programming course — keep working through the series to build a complete, practical understanding of kernel memory management.

Frequently Asked Questions

1. What triggers the Linux kernel OOM killer?

It triggers only when the page allocator has exhausted reclaim, compaction, and swap options and still cannot satisfy a memory request.

2. How do I check a process’s OOM score?

Read /proc/<pid>/oom_score, or use the choom -p <pid> command for a more readable view.

3. Can I make a process completely immune to the OOM killer?

Yes, by setting its oom_score_adj to -1000 as root, though this should be reserved for truly critical processes.

4. What is the oom_reaper thread?

A kernel thread that asynchronously frees a victim process’s memory pages so OOM recovery isn’t delayed by a stuck victim.

5. Is the classic OOM killer still relevant on modern kernels?

Yes, it remains the final safety net, though PSI and cgroup v2 now let systems react to memory pressure earlier and more predictably.

6. Why did the OOM killer kill a process that wasn’t using much memory?

Built-in heuristics can protect high-usage but critical processes (root-owned, kernel threads, device-holding tasks), shifting the score-based decision toward a different candidate.

7. What’s the difference between oom_score and oom_score_adj?

oom_score is the kernel-computed, real-time usage-based score; oom_score_adj is a manual bias you add on top of it.

8. How does cgroup v2 change OOM behaviour for containers?

With memory.oom.group enabled, the kernel can kill every process in a container’s cgroup together instead of a single process, giving cleaner container-level recovery.

Continue the Free Linux Kernel Programming Course

Explore more free lessons on Linux kernel memory management, device drivers, and embedded systems at EmbeddedPathashala.

Visit EmbeddedPathashala

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *