What Is oom_score in Linux? – Linux Device Drivers Coaching in Hyderabad

Linux OOM Killer Explained (Free Linux Device Drivers Course)
How the kernel’s Out-Of-Memory killer works on modern 6.x kernels, from oom_score to cgroup-aware OOM handling
Reading Time
15 min
Level
Beginner to Intermediate
Kernel Version
6.x (current LTS)
What You Will Learn
Why the OOM killer exists
oom_score and badness
oom_score_adj
Reading kernel OOM logs
cgroup-aware OOM killer
PSI and early OOM daemons

Welcome back to our free Linux kernel development course. In the previous lecture we covered memory overcommit. Now we look at what happens on the other side: the moment physical memory and swap are truly exhausted and the kernel must forcibly reclaim memory by killing a process. This is the Linux OOM killer, one of the most misunderstood subsystems for anyone learning free Linux device drivers or free embedded systems development.

Prerequisites
  • Comfort with the memory overcommit concepts from the previous lecture
  • Basic familiarity with /proc, dmesg, and process IDs
  • Root or sudo access on a disposable test VM (never test OOM behavior on a production box)
What Is the OOM Killer?

The Out-Of-Memory killer, usually just called the OOM killer, is a kernel-level last resort. It only activates when the kernel genuinely cannot satisfy a memory request through any other means: no free physical pages, no reclaimable cache, and no more swap space. Rather than let the entire system freeze or crash, the kernel selects one process to sacrifice, frees its memory, and lets the rest of the system keep running.

Path to an OOM kill
Allocation request
→
Reclaim caches, swap out pages
→
Still no memory available
→
OOM killer selects a victim
How the Kernel Picks a Victim: oom_score and badness

The kernel does not kill processes at random. Every process is assigned a badness score, visible as oom_score in /proc. The score weighs mainly the resident memory a process is using, adjusted by a bias value you control yourself.

$ cat /proc/<pid>/oom_score
542

Roughly speaking, the process consuming the largest share of memory, relative to total available memory, ends up with the highest score and becomes the most likely candidate. This design intentionally favors killing the single biggest memory consumer rather than many small processes, since that reclaims the most memory in one action.

Tuning Victim Priority with oom_score_adj

You can influence, but not fully override, the kernel’s choice using oom_score_adj, a per-process value ranging from -1000 to +1000.

Value Effect
-1000 Process is completely exempt from being OOM-killed
0 Default, no adjustment
+1000 Process is almost guaranteed to be picked first
# Make a critical service much less likely to be killed
$ echo -900 | sudo tee /proc/$(pgrep my-critical-service)/oom_score_adj

# Make a disposable batch job the preferred victim
$ echo 800 | sudo tee /proc/$(pgrep batch-worker)/oom_score_adj

Systemd services can set this declaratively as well, which is the recommended approach on modern distributions instead of scripting /proc writes:

[Service]
OOMScoreAdjust=-900
Reading OOM Killer Logs

Every OOM kill is logged to the kernel ring buffer. On a current 6.x kernel you will find it with:

$ dmesg -T | grep -i "out of memory"
$ journalctl -k | grep -i oom

The log entry reports the invoking process, the trigger allocation order, and the victim’s PID, memory footprint and computed score, which is exactly the data you need when investigating a production incident.

The Cgroup-Aware OOM Killer

A major improvement that classic kernel textbooks written before kernel 4.19 do not cover is the cgroup-aware OOM killer. On modern container hosts running cgroup v2, you can set memory.oom.group on a cgroup so that when any process inside it triggers an OOM condition, the kernel kills the entire cgroup as a unit rather than a single arbitrary process.

# Enable group-wide OOM kill for a systemd service's cgroup
$ echo 1 | sudo tee /sys/fs/cgroup/system.slice/myapp.service/memory.oom.group

This matters enormously for microservices and container workloads: killing only one process out of a multi-process container often leaves the container in a broken, half-alive state, whereas group-wide killing lets the orchestrator cleanly restart the whole container.

PSI and Proactive OOM Prevention

Modern kernels expose Pressure Stall Information (PSI) through /proc/pressure/memory, giving a live percentage of time processes spend stalled waiting for memory. This enables user-space daemons such as systemd-oomd and earlyoom to react to memory pressure before the kernel’s own last-resort OOM killer has to step in, usually producing a much gentler and more predictable outcome.

$ cat /proc/pressure/memory
some avg10=0.00 avg60=0.00 avg300=0.00 total=0
full avg10=0.00 avg60=0.00 avg300=0.00 total=0
Layers of memory pressure response on a modern kernel
PSI signals rising pressure
→
systemd-oomd / earlyoom intervene early
→
Kernel OOM killer as final fallback
Real-World Use Cases
  • Kubernetes nodes: rely on cgroup-aware OOM killing so an entire pod’s containers die together rather than leaving orphaned processes.
  • Desktop Linux: systemd-oomd protects an interactive desktop session from freezing when a browser tab runs away with memory.
  • Embedded boards: engineers often disable overcommit and set strict per-process oom_score_adj values to protect safety-critical drivers and daemons.
  • Database servers: set a strongly negative oom_score_adj on the database process while allowing batch or reporting jobs to be sacrificed first.
Common Mistakes and Troubleshooting
Mistake Why It Hurts Fix
Setting oom_score_adj to -1000 on too many processes Leaves the kernel with too few eligible victims, risking a full system hang Protect only the truly critical process, not entire categories
Assuming the OOM killer always targets the “guilty” process The victim is chosen by memory footprint, not by who caused the pressure Always read the full dmesg trace to understand actual memory usage at kill time
Ignoring PSI metrics in monitoring You lose early warning before a hard OOM kill occurs Add /proc/pressure/memory to your monitoring dashboards
Best Practices
  • Use OOMScoreAdjust in systemd unit files rather than manual /proc writes for anything persistent.
  • Enable memory.oom.group for multi-process containers so the whole unit is reclaimed together.
  • Deploy systemd-oomd or earlyoom on interactive and latency-sensitive systems to avoid full system stalls before the kernel intervenes.
  • Always correlate an OOM incident with dmesg, journalctl -k, and PSI history, not memory usage alone.
Performance Considerations

Waiting for the kernel’s own OOM killer is the slowest and most disruptive path, since the system may already be thrashing under memory pressure by the time it triggers. Proactive PSI-based tools intervene sooner and generally cause far less disruption to unrelated processes.

Security Considerations

Uncontrolled memory-hungry processes, whether malicious or buggy, can be used to trigger denial-of-service conditions by forcing the OOM killer to terminate unrelated critical services. Combining sane oom_score_adj policy with cgroup memory limits (covered in the overcommit lecture) closes this gap.

Summary / Key Takeaways
  • The OOM killer is a last-resort kernel mechanism, not an error condition to fear.
  • Victim selection is driven by oom_score, which you can bias with oom_score_adj.
  • Modern cgroup v2 systems support group-wide OOM kills via memory.oom.group.
  • PSI and user-space daemons like systemd-oomd can prevent hard kernel OOM kills entirely.
  • Always read dmesg or journalctl -k logs to understand exactly what happened during an OOM event.
Conclusion

The OOM killer completes the picture we started with memory overcommit: overcommit decides what gets promised, and the OOM killer decides what happens when those promises can no longer be kept. Understanding both together gives you the foundation to confidently tune memory behavior on any Linux system, from a tiny embedded board to a large Kubernetes cluster.

Frequently Asked Questions
1. What triggers the Linux OOM killer?

It triggers only when the kernel cannot satisfy a memory allocation after exhausting free memory, reclaimable caches, and swap space.

2. How does the kernel choose which process to kill?

It computes a badness score, mainly based on resident memory usage, and adjusts it using each process’s oom_score_adj value.

3. Can I make a process immune to the OOM killer?

Yes, by setting its oom_score_adj to -1000, though this should be reserved for truly critical processes only.

4. What is the difference between the kernel OOM killer and systemd-oomd?

The kernel OOM killer is a last-resort in-kernel mechanism. systemd-oomd is a user-space daemon that watches PSI metrics and intervenes earlier, often avoiding a hard kernel OOM kill altogether.

5. How do I check if the OOM killer has run on my system?

Run dmesg -T | grep -i "out of memory" or journalctl -k | grep -i oom to see kernel log entries for past OOM events.

6. What is memory.oom.group used for?

It is a cgroup v2 control that forces the kernel to kill every process in a cgroup together during an OOM event, which is essential for multi-process containers.

7. Does disabling swap make the OOM killer trigger sooner?

Yes. Swap gives the kernel extra room to reclaim memory before resorting to killing a process, so systems without swap tend to hit the OOM killer earlier under memory pressure.

Continue Your Free Linux Device Drivers Course

Explore more free lectures on Linux kernel memory management and device driver development.

Previous Lecture
Next lecture

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *