How Does Linux Manage Memory with Cgroups? – Free Linux Device Drivers Course in Hyderabad

Linux Cgroups Memory Management and Slab Allocator Caveats
Free Linux Kernel Programming Course · Free Linux Device Drivers Course
18 min read
Intermediate
Cgroups v2 Ready

Linux cgroups memory management
memory.max
memory.high
slab allocator internal fragmentation
Free Linux Kernel Development Course

Understanding Linux cgroups memory management is essential for anyone writing kernel code, building embedded systems, or administering production servers. Cgroups let the kernel divide memory fairly between groups of processes, while the slab allocator underneath every kernel allocation has its own set of sizing caveats that every driver and kernel developer should be aware of.

This lesson continues our free Linux kernel development course and free Linux device drivers course at EmbeddedPathashala. We will explain, in simple language, how memory control groups regulate memory usage on modern systems, and then walk through why the underlying page and slab allocators can waste more memory than beginners expect.

What You Will Learn
  • What cgroups are and why the memory controller matters
  • The difference between memory.min, memory.low, memory.high, and memory.max
  • How to set up and test a memory cgroup on a modern Linux system
  • Why the page allocator only hands out power-of-2 sized chunks
  • How the slab allocator’s design causes internal fragmentation for small allocations
  • How to reason about memory waste when choosing allocation sizes in your own code
Prerequisites
  • Basic Linux command-line skills
  • Root or sudo access to a Linux machine for the hands-on cgroup example
  • Having read Part 1 of this series on device driver memory allocation is recommended but not required

What Are Cgroups and Why Memory Control Matters

Control groups, or cgroups, are a Linux kernel feature that organises processes into hierarchical groups and then applies resource limits, accounting, and prioritisation to each group as a whole. Among the several resource controllers available, the memory controller (memcg) is one of the most widely used, because uncontrolled memory usage by one application can starve every other application — and even the kernel itself — of memory.

Most modern distributions now default to cgroups v2, which uses a single unified hierarchy instead of the separate, sometimes inconsistent hierarchies of the older cgroups v1. Systemd, container runtimes such as containerd and Docker, and Kubernetes all rely on cgroups v2 memory accounting today, which makes this topic directly relevant even if you never touch the pseudo-files by hand.

Cgroup Memory Hierarchy
Root Cgroup
All system memory
app.slice
memory.max set
container-1
memory.high set

memory.min, memory.low, memory.high, memory.max Explained

Inside every cgroup’s directory under /sys/fs/cgroup/, the memory controller exposes a small set of pseudo-files that let you shape how memory pressure is handled. Understanding the difference between them is the key to using cgroup memory limits correctly.

Pseudo-file Type of Protection/Limit What Happens When Crossed
memory.min Hard protection (guarantee) This much memory is never reclaimed from the group, even under system-wide pressure
memory.low Best-effort protection The kernel tries not to reclaim below this line, but may if there is no better option
memory.high Soft usage ceiling The group is throttled and reclaimed aggressively once usage passes this point
memory.max Hard usage ceiling The kernel’s OOM killer is invoked inside the group if usage tries to exceed this value

A simple way to remember the relationship: memory.min and memory.low protect a group from having memory taken away, while memory.high and memory.max stop a group from taking too much in the first place.

Hands-On: Creating a Memory-Limited Cgroup

Here is a simple, original walkthrough for creating a cgroup and giving it a hard memory ceiling on a system using cgroups v2. Run these as root.

# Create a new cgroup called "demo"
mkdir /sys/fs/cgroup/demo

# Set a 100 MB hard memory limit for this group
echo 100M > /sys/fs/cgroup/demo/memory.max

# Move the current shell into the cgroup
echo $$ > /sys/fs/cgroup/demo/cgroup.procs

# Check current memory usage of the group
cat /sys/fs/cgroup/demo/memory.current

Once your shell is inside the demo group, any process you launch from that shell inherits the same 100 MB ceiling. If a process inside the group tries to allocate beyond that ceiling and the kernel cannot reclaim enough memory elsewhere in the group, the kernel’s out-of-memory killer terminates a process inside that specific group rather than affecting the rest of the system.

Tip: Always test cgroup memory limits on a non-critical machine first. Setting memory.max too low on a group running important services can trigger unexpected OOM kills.

Why Small Allocations Can Waste Memory: Slab Allocator Caveats

To understand the slab allocator’s caveats, it helps to recall how kernel memory allocation is layered. At the bottom sits the page (buddy system) allocator, which only ever hands out memory in power-of-2 numbers of pages. The power that 2 is raised to is called the order, and on both x86_64 and ARM systems it typically ranges from order 0 (a single page) up to order 10 (1024 pages).

This works beautifully when you need a large, page-aligned buffer. It works poorly when you need something much smaller than a page — and small allocation requests, well under 4096 bytes, are extremely common in kernel code. This mismatch is exactly why the slab allocator exists: it is layered on top of the page allocator specifically to serve these small, frequent allocation requests efficiently, using object caches and small generic memory caches instead of handing out a whole page for a tiny structure.

Allocator Layers
Driver / Kernel Code
kmalloc(), kzalloc()
Slab Allocator
object caches for small sizes
Page (Buddy) Allocator
power-of-2 page chunks

What Both Allocators Guarantee

It is worth being precise about what each layer promises, because both the page allocator and the slab allocator built on top of it guarantee physically contiguous memory. In addition, both guarantee cacheline-aligned memory, meaning the start of the returned buffer is aligned to a CPU cache line boundary. This alignment matters for performance-sensitive code and for hardware that requires aligned DMA buffers.

Property Page Allocator Slab Allocator
Physically contiguous memory Yes Yes
Cacheline-aligned memory Yes Yes
Granularity Power-of-2 pages only Small, application-sized objects
Best suited for Large, page-sized or larger buffers Frequent small allocations under one page

Internal Fragmentation: Why Requested Size Is Not Delivered Size

Because the page allocator only deals in power-of-2 page counts, a request for even a single byte more than one page forces the allocator to hand out an entire additional page — doubling the allocation to the next power of 2. This wasted space is called internal fragmentation, and it can be significant when many such allocations happen across a running system.

The slab allocator reduces this waste for small objects by grouping many same-sized objects inside a shared page, rather than rounding every tiny request up to a full page. However, slab caches themselves are still built from fixed-size object “buckets” internally, so a request that falls just above one bucket size boundary can still be rounded up to the next bucket, wasting some space — just far less than it would if the page allocator handled every small request directly.

Rule of thumb: The closer your requested allocation size is to a power of 2 (or to a common slab bucket size), the less memory is wasted to internal fragmentation. When you control the layout of a structure, it is often worth checking its actual sizeof() against nearby bucket boundaries.

Real-World Use Case

Consider a network driver that allocates a small per-packet metadata structure on every incoming packet, potentially thousands of times per second. If that structure’s size sits just above a slab bucket boundary, every single allocation wastes the difference — and at high packet rates, that small per-object waste adds up to a measurable amount of memory pressure system-wide. Trimming or padding the structure so it fits neatly inside a bucket is a small change that can meaningfully reduce memory overhead in high-throughput drivers.

Common Mistakes and Troubleshooting
Mistake Why It’s a Problem Fix
Setting memory.max without testing Can trigger unexpected in-group OOM kills Test limits on non-critical workloads first, watch memory.events
Confusing memory.low with a hard guarantee memory.low is best-effort, not absolute Use memory.min when a true hard guarantee is required
Assuming small kmalloc() requests cost exactly the requested bytes Slab bucket rounding can allocate more than requested Check actual slab bucket sizes via /proc/slabinfo
Requesting slightly-over-a-page buffers repeatedly Forces allocation of double the pages due to power-of-2 rounding Keep large buffers at or just under a page-size boundary where possible

Best Practices
  • Use memory.min for guarantees that must never be violated, and memory.low for softer, best-effort protection.
  • Set memory.high as an early-warning throttle before relying on the harder memory.max ceiling.
  • Monitor memory.current and memory.events when tuning cgroup limits in production.
  • When designing frequently-allocated kernel structures, check their size against common slab bucket boundaries to minimise waste.

Performance Considerations

Cgroup memory accounting adds a small amount of bookkeeping overhead to every allocation and free inside a group, which is generally acceptable given the isolation benefits it provides. On the slab side, choosing allocation sizes that align well with existing slab buckets avoids unnecessary rounding, which in turn reduces overall memory pressure and can indirectly improve cache behaviour across the system.

Security Considerations

Memory cgroups are also a security boundary in multi-tenant systems such as container hosts: a properly configured memory.max prevents one tenant’s workload from starving memory away from every other tenant on the same machine. Getting the limits wrong in either direction — too loose, or too strict — can turn into either a resource-exhaustion risk or an availability problem, so limits should be reviewed as part of your normal security hardening process.

Summary and Key Takeaways
  • Cgroups let the kernel divide and protect memory between groups of processes; cgroups v2 is the modern, unified approach used by systemd and most container runtimes.
  • memory.min and memory.low protect memory from being taken away; memory.high and memory.max cap how much a group can use.
  • The page allocator only hands out power-of-2 page chunks, which can waste memory for slightly-over-a-page requests.
  • The slab allocator reduces (but does not eliminate) this waste for small, frequent allocations by using fixed-size object buckets.
  • Both the page allocator and slab allocator guarantee physically contiguous and cacheline-aligned memory.

Conclusion

Memory management in Linux is a story of layers: cgroups decide how much memory a group of processes may use overall, while the page and slab allocators decide how that memory is actually carved up and handed out internally. Understanding both layers — from setting a sensible memory.max on a container, to sizing your kernel structures with slab bucket boundaries in mind — will make you a noticeably more effective Linux kernel and device driver developer. This wraps up this section of our free Linux kernel development course; continue to the next lesson to keep building your kernel programming skills.

Frequently Asked Questions

Q1. What is the difference between cgroups v1 and cgroups v2?
Cgroups v2 uses a single unified hierarchy for all resource controllers, replacing the separate, sometimes inconsistent hierarchies used in cgroups v1, and is the default on most modern distributions today.

Q2. What happens if a process exceeds memory.max?
The kernel attempts to reclaim memory within the group, and if it cannot reclaim enough, it invokes the out-of-memory killer inside that specific cgroup.

Q3. Is memory.low a guaranteed protection?
No, it is best-effort. For an absolute guarantee, use memory.min instead.

Q4. Why does the page allocator only give power-of-2 pages?
This buddy-system design makes tracking and merging free memory blocks simple and fast, at the cost of some wasted space for requests that do not fit a power-of-2 size exactly.

Q5. Does the slab allocator eliminate memory waste completely?
No. It substantially reduces waste for small allocations compared to using the page allocator directly, but rounding to fixed bucket sizes can still waste a small amount of memory.

Q6. Are slab allocations physically contiguous?
Yes. Both the page allocator and the slab allocator built on top of it guarantee physically contiguous, cacheline-aligned memory.

Q7. Can memory cgroups help secure a shared server?
Yes. Properly configured memory limits prevent one workload from starving memory away from others on the same machine, which is an important part of multi-tenant security.

Continue the Free Linux Kernel Development Course

This lesson is part of EmbeddedPathashala’s free Linux device drivers course and free embedded systems course.

Browse All Lessons

 

Leave a Reply

Your email address will not be published. Required fields are marked *