memory.max
memory.high
slab allocator internal fragmentation
Free Linux Kernel Development Course
Understanding Linux cgroups memory management is essential for anyone writing kernel code, building embedded systems, or administering production servers. Cgroups let the kernel divide memory fairly between groups of processes, while the slab allocator underneath every kernel allocation has its own set of sizing caveats that every driver and kernel developer should be aware of.
This lesson continues our free Linux kernel development course and free Linux device drivers course at EmbeddedPathashala. We will explain, in simple language, how memory control groups regulate memory usage on modern systems, and then walk through why the underlying page and slab allocators can waste more memory than beginners expect.
- What cgroups are and why the memory controller matters
- The difference between
memory.min,memory.low,memory.high, andmemory.max - How to set up and test a memory cgroup on a modern Linux system
- Why the page allocator only hands out power-of-2 sized chunks
- How the slab allocator’s design causes internal fragmentation for small allocations
- How to reason about memory waste when choosing allocation sizes in your own code
- Basic Linux command-line skills
- Root or sudo access to a Linux machine for the hands-on cgroup example
- Having read Part 1 of this series on device driver memory allocation is recommended but not required
Control groups, or cgroups, are a Linux kernel feature that organises processes into hierarchical groups and then applies resource limits, accounting, and prioritisation to each group as a whole. Among the several resource controllers available, the memory controller (memcg) is one of the most widely used, because uncontrolled memory usage by one application can starve every other application — and even the kernel itself — of memory.
Most modern distributions now default to cgroups v2, which uses a single unified hierarchy instead of the separate, sometimes inconsistent hierarchies of the older cgroups v1. Systemd, container runtimes such as containerd and Docker, and Kubernetes all rely on cgroups v2 memory accounting today, which makes this topic directly relevant even if you never touch the pseudo-files by hand.
Inside every cgroup’s directory under /sys/fs/cgroup/, the memory controller exposes a small set of pseudo-files that let you shape how memory pressure is handled. Understanding the difference between them is the key to using cgroup memory limits correctly.
| Pseudo-file | Type of Protection/Limit | What Happens When Crossed |
|---|---|---|
memory.min |
Hard protection (guarantee) | This much memory is never reclaimed from the group, even under system-wide pressure |
memory.low |
Best-effort protection | The kernel tries not to reclaim below this line, but may if there is no better option |
memory.high |
Soft usage ceiling | The group is throttled and reclaimed aggressively once usage passes this point |
memory.max |
Hard usage ceiling | The kernel’s OOM killer is invoked inside the group if usage tries to exceed this value |
A simple way to remember the relationship: memory.min and memory.low protect a group from having memory taken away, while memory.high and memory.max stop a group from taking too much in the first place.
Here is a simple, original walkthrough for creating a cgroup and giving it a hard memory ceiling on a system using cgroups v2. Run these as root.
# Create a new cgroup called "demo"
mkdir /sys/fs/cgroup/demo
# Set a 100 MB hard memory limit for this group
echo 100M > /sys/fs/cgroup/demo/memory.max
# Move the current shell into the cgroup
echo $$ > /sys/fs/cgroup/demo/cgroup.procs
# Check current memory usage of the group
cat /sys/fs/cgroup/demo/memory.current
Once your shell is inside the demo group, any process you launch from that shell inherits the same 100 MB ceiling. If a process inside the group tries to allocate beyond that ceiling and the kernel cannot reclaim enough memory elsewhere in the group, the kernel’s out-of-memory killer terminates a process inside that specific group rather than affecting the rest of the system.
memory.max too low on a group running important services can trigger unexpected OOM kills.To understand the slab allocator’s caveats, it helps to recall how kernel memory allocation is layered. At the bottom sits the page (buddy system) allocator, which only ever hands out memory in power-of-2 numbers of pages. The power that 2 is raised to is called the order, and on both x86_64 and ARM systems it typically ranges from order 0 (a single page) up to order 10 (1024 pages).
This works beautifully when you need a large, page-aligned buffer. It works poorly when you need something much smaller than a page — and small allocation requests, well under 4096 bytes, are extremely common in kernel code. This mismatch is exactly why the slab allocator exists: it is layered on top of the page allocator specifically to serve these small, frequent allocation requests efficiently, using object caches and small generic memory caches instead of handing out a whole page for a tiny structure.
It is worth being precise about what each layer promises, because both the page allocator and the slab allocator built on top of it guarantee physically contiguous memory. In addition, both guarantee cacheline-aligned memory, meaning the start of the returned buffer is aligned to a CPU cache line boundary. This alignment matters for performance-sensitive code and for hardware that requires aligned DMA buffers.
| Property | Page Allocator | Slab Allocator |
|---|---|---|
| Physically contiguous memory | Yes | Yes |
| Cacheline-aligned memory | Yes | Yes |
| Granularity | Power-of-2 pages only | Small, application-sized objects |
| Best suited for | Large, page-sized or larger buffers | Frequent small allocations under one page |
Because the page allocator only deals in power-of-2 page counts, a request for even a single byte more than one page forces the allocator to hand out an entire additional page — doubling the allocation to the next power of 2. This wasted space is called internal fragmentation, and it can be significant when many such allocations happen across a running system.
The slab allocator reduces this waste for small objects by grouping many same-sized objects inside a shared page, rather than rounding every tiny request up to a full page. However, slab caches themselves are still built from fixed-size object “buckets” internally, so a request that falls just above one bucket size boundary can still be rounded up to the next bucket, wasting some space — just far less than it would if the page allocator handled every small request directly.
sizeof() against nearby bucket boundaries.Consider a network driver that allocates a small per-packet metadata structure on every incoming packet, potentially thousands of times per second. If that structure’s size sits just above a slab bucket boundary, every single allocation wastes the difference — and at high packet rates, that small per-object waste adds up to a measurable amount of memory pressure system-wide. Trimming or padding the structure so it fits neatly inside a bucket is a small change that can meaningfully reduce memory overhead in high-throughput drivers.
| Mistake | Why It’s a Problem | Fix |
|---|---|---|
Setting memory.max without testing |
Can trigger unexpected in-group OOM kills | Test limits on non-critical workloads first, watch memory.events |
Confusing memory.low with a hard guarantee |
memory.low is best-effort, not absolute |
Use memory.min when a true hard guarantee is required |
| Assuming small kmalloc() requests cost exactly the requested bytes | Slab bucket rounding can allocate more than requested | Check actual slab bucket sizes via /proc/slabinfo |
| Requesting slightly-over-a-page buffers repeatedly | Forces allocation of double the pages due to power-of-2 rounding | Keep large buffers at or just under a page-size boundary where possible |
- Use
memory.minfor guarantees that must never be violated, andmemory.lowfor softer, best-effort protection. - Set
memory.highas an early-warning throttle before relying on the hardermemory.maxceiling. - Monitor
memory.currentandmemory.eventswhen tuning cgroup limits in production. - When designing frequently-allocated kernel structures, check their size against common slab bucket boundaries to minimise waste.
Cgroup memory accounting adds a small amount of bookkeeping overhead to every allocation and free inside a group, which is generally acceptable given the isolation benefits it provides. On the slab side, choosing allocation sizes that align well with existing slab buckets avoids unnecessary rounding, which in turn reduces overall memory pressure and can indirectly improve cache behaviour across the system.
Memory cgroups are also a security boundary in multi-tenant systems such as container hosts: a properly configured memory.max prevents one tenant’s workload from starving memory away from every other tenant on the same machine. Getting the limits wrong in either direction — too loose, or too strict — can turn into either a resource-exhaustion risk or an availability problem, so limits should be reviewed as part of your normal security hardening process.
- Cgroups let the kernel divide and protect memory between groups of processes; cgroups v2 is the modern, unified approach used by systemd and most container runtimes.
memory.minandmemory.lowprotect memory from being taken away;memory.highandmemory.maxcap how much a group can use.- The page allocator only hands out power-of-2 page chunks, which can waste memory for slightly-over-a-page requests.
- The slab allocator reduces (but does not eliminate) this waste for small, frequent allocations by using fixed-size object buckets.
- Both the page allocator and slab allocator guarantee physically contiguous and cacheline-aligned memory.
Memory management in Linux is a story of layers: cgroups decide how much memory a group of processes may use overall, while the page and slab allocators decide how that memory is actually carved up and handed out internally. Understanding both layers — from setting a sensible memory.max on a container, to sizing your kernel structures with slab bucket boundaries in mind — will make you a noticeably more effective Linux kernel and device driver developer. This wraps up this section of our free Linux kernel development course; continue to the next lesson to keep building your kernel programming skills.
Q1. What is the difference between cgroups v1 and cgroups v2?
Cgroups v2 uses a single unified hierarchy for all resource controllers, replacing the separate, sometimes inconsistent hierarchies used in cgroups v1, and is the default on most modern distributions today.
Q2. What happens if a process exceeds memory.max?
The kernel attempts to reclaim memory within the group, and if it cannot reclaim enough, it invokes the out-of-memory killer inside that specific cgroup.
Q3. Is memory.low a guaranteed protection?
No, it is best-effort. For an absolute guarantee, use memory.min instead.
Q4. Why does the page allocator only give power-of-2 pages?
This buddy-system design makes tracking and merging free memory blocks simple and fast, at the cost of some wasted space for requests that do not fit a power-of-2 size exactly.
Q5. Does the slab allocator eliminate memory waste completely?
No. It substantially reduces waste for small allocations compared to using the page allocator directly, but rounding to fixed bucket sizes can still waste a small amount of memory.
Q6. Are slab allocations physically contiguous?
Yes. Both the page allocator and the slab allocator built on top of it guarantee physically contiguous, cacheline-aligned memory.
Q7. Can memory cgroups help secure a shared server?
Yes. Properly configured memory limits prevent one workload from starving memory away from others on the same machine, which is an important part of multi-tenant security.
This lesson is part of EmbeddedPathashala’s free Linux device drivers course and free embedded systems course.
