If you are enrolled in our free Linux kernel development course, you already know that kernel memory is not handed out the way user-space malloc() hands out heap memory. Every byte a kernel module or driver requests through kmalloc() memory allocation has to come from a fixed set of pre-sized buckets, and that design choice has a side effect almost nobody talks about until they profile it: wasted RAM.
In this lecture, part of our free linux kernel development course and free linux device drivers course, you will learn exactly why a driver that asks for 1.5 MB of memory can end up costing the system 2 MB or more, why crossing certain size thresholds spikes that waste to almost 100%, and what the modern kernel gives you to detect and avoid it.
What You Will Learn
Why large allocations waste memory
ksize() and kmalloc_size_roundup()
alloc_pages_exact() / free_pages_exact()
Slab vs page allocator boundary
Debugging allocation waste
A built kernel source tree
Loadable kernel module basics
Familiarity with printk()
If you have not yet covered how to write and load a basic kernel module, go back one lecture in this free linux kernel development course before continuing here — you will need to build and insert a small module to try the experiments below.
Why Linux Kernel Memory Allocation Is Not Like malloc()
In user space, malloc() can, in principle, hand you almost any number of bytes you ask for, because the C library’s heap manager is free to carve up virtual memory however it likes. The kernel cannot afford that luxury. It runs in a privileged, resource-constrained environment where fragmentation, locking overhead, and cache-line behaviour matter enormously, so kernel memory allocation is built around two cooperating layers instead of one general-purpose heap.
Serves small, fixed-size “buckets”: 8, 16, 32, 64, 128 … up to a few KB
Hands out whole pages, always in power-of-two counts: 1, 2, 4, 8, 16 pages…
Every kernel memory allocation ultimately bottoms out at the page allocator, also called the buddy system, because physical RAM itself is managed in fixed-size pages (commonly 4 KB on most architectures). The slab layer sits on top of the page allocator purely as an optimisation: most kernel allocations are small and short-lived, so handing out whole pages for a 40-byte structure would be absurdly wasteful. The slab layer instead reserves a handful of pages up front and slices them into same-sized objects, which is where kmalloc() buckets come from.
How kmalloc() Memory Allocation Buckets Actually Work
When your driver calls kmalloc(size, GFP_KERNEL), the kernel does not allocate exactly size bytes. It rounds your request up to the nearest available bucket. On a typical 64-bit build you will find general-purpose buckets around these sizes:
| Requested Size | Bucket Actually Used | Approximate Waste |
|---|---|---|
| 20 bytes | kmalloc-32 | ~37% |
| 100 bytes | kmalloc-128 | ~22% |
| 1000 bytes | kmalloc-1k | ~2% |
| 5000 bytes | kmalloc-8k | ~38% |
Notice the pattern: waste is small right after you cross into a new bucket and grows steadily until you hit the next one, where it resets. This “sawtooth” behaviour is completely predictable once you understand it, and it is one of the most important things to internalise in this free linux kernel development course.
Where kmalloc() Hands Off to the Page Allocator
General-purpose kmalloc() buckets stop at a fixed ceiling — historically around 8 KB, though the exact cutoff depends on architecture and kernel configuration. Ask for anything larger and the slab layer cannot help you anymore; the request is served directly by the page allocator instead. This is the point where kernel memory allocation waste can get genuinely ugly, because the page allocator can only give you memory in power-of-two page counts.
Requested
~2.1 MB
Allocated
4 MB (next order)
Wasted
~1.9 MB
The moment a request crosses just past a page-order boundary (for example, just over 2 MB when the buddy system’s next available block is 4 MB), the kernel must round all the way up to that next power-of-two block. The waste in that single allocation can approach 90% of the memory actually handed out. This is exactly the kind of allocation pattern that quietly bloats memory usage on embedded boards and low-RAM devices, and it is a favourite interview question for embedded Linux roles.
Measuring Real Allocation Size: ksize() and kmalloc_size_roundup()
You cannot manage what you do not measure. The kernel gives driver authors two complementary tools to find out how much memory an allocation truly consumes.
ksize() — Check After Allocation
ksize() reports the true usable size of a pointer that kmalloc() already returned. It has existed for a long time and is still useful for auditing, but modern kernels discourage using it to “opportunistically” grow into memory you did not originally ask for, since compiler-based buffer-overflow hardening (things like __alloc_size hints and FORTIFY_SOURCE) work best when the requested size and the used size agree from the start.
#include <linux/slab.h>
void *buf = kmalloc(100, GFP_KERNEL);
if (buf) {
pr_info("requested=100 actual_usable=%zu\n", ksize(buf));
kfree(buf);
}
kmalloc_size_roundup() — Check Before Allocation
Modern kernels added kmalloc_size_roundup() precisely to fix the awkwardness of discovering your real bucket size only after allocating. It answers “how many bytes would I actually get if I asked for N?” without allocating anything, which lets you size buffers correctly up front instead of relying on ksize() afterwards.
#include <linux/slab.h>
size_t want = 100;
size_t real_bucket = kmalloc_size_roundup(want);
pr_info("asking for %zu will actually give %zu\n", want, real_bucket);
void *buf = kmalloc(real_bucket, GFP_KERNEL);
Older approach
Allocate first → call ksize() afterwards → opportunistically use the extra space → compiler bounds-checking gets confused
Modern approach
Call kmalloc_size_roundup() first → allocate exactly that size → compiler and runtime checks stay accurate
Avoiding Page Allocator Waste with alloc_pages_exact()
When a driver genuinely needs more memory than the slab ceiling allows, and that memory does not need to be a clean power-of-two size, the page allocator’s default API is the wrong tool. The kernel instead provides alloc_pages_exact() and its counterpart free_pages_exact(), which internally allocate the necessary power-of-two block but then free back the unused tail pages, leaving you with memory sized much closer to what you actually asked for.
#include <linux/gfp.h>
size_t need = 2200100; /* just over 2 MB */
void *mem = alloc_pages_exact(need, GFP_KERNEL);
if (mem) {
/* use mem ... */
free_pages_exact(mem, need);
}
This does not eliminate rounding entirely — you still get memory in whole-page units — but it avoids the extreme power-of-two jumps that make naive large allocations so costly.
Real-World Use Case: Why This Matters on Embedded Boards
Picture a network driver on a 512 MB embedded board that allocates several receive buffers slightly larger than 2 MB each during initialisation, using plain kmalloc()-style page allocation without checking actual usage. On a desktop with 32 GB of RAM nobody would notice. On that embedded board, a handful of such allocations can eat tens of megabytes more than intended, indirectly triggering the OOM killer under memory pressure elsewhere in the system. This is precisely the class of bug this free linux device drivers course trains you to spot before it ships.
Common Mistakes and Troubleshooting Tips
| Mistake | Why It Hurts | Fix |
|---|---|---|
| Allocating just over a page-order boundary | Waste can approach 90%+ | Use alloc_pages_exact() or reshape the request |
| Relying on ksize() to grow buffers after the fact | Confuses compiler bounds hardening | Use kmalloc_size_roundup() before allocating |
| Assuming kmalloc() gives exactly what you asked for | Silent memory bloat under load | Check /proc/slabinfo and bucket sizes during design |
| Never profiling allocation size vs actual usage | Waste hides until memory pressure hits | Add temporary ksize() logging during development |
Best Practices for Kernel Memory Allocation
- Know your bucket sizes before designing a data structure that gets allocated frequently.
- Prefer
kmalloc_size_roundup()over after-the-factksize()queries in new code. - For allocations that regularly cross the slab ceiling, evaluate
alloc_pages_exact()orvmalloc()depending on whether physical contiguity is required. - Use
/proc/slabinfoand/proc/meminfoto sanity-check real-world memory pressure during driver bring-up. - Batch small, frequent allocations into a dedicated
kmem_cacheinstead of repeated generic kmalloc() calls where it makes sense.
Performance and Security Considerations
Beyond raw memory waste, bucket rounding interacts with security hardening. Because objects of the same bucket size share a cache, precise sizing reduces the “spray” surface that certain heap-exploitation techniques rely on, and it keeps compiler-assisted bounds checking (FORTIFY_SOURCE, KASAN, UBSAN) accurate. Getting allocation sizing right is therefore not just a memory-efficiency exercise — it is part of writing defensively secure kernel code, a theme you will see repeated throughout this free linux kernel development course.
Summary and Key Takeaways
- Kernel memory allocation happens in two layers: the slab allocator for small, fixed-size buckets, and the page allocator for whole pages.
- kmalloc() always rounds your request up to the nearest bucket, and waste follows a predictable sawtooth pattern.
- Crossing a page-order boundary can spike waste close to 90–100% for a single large allocation.
kmalloc_size_roundup()is the modern way to know your real allocation size before you allocate, replacing many olderksize()patterns.alloc_pages_exact()/free_pages_exact()reduce waste for large allocations that don’t need to be power-of-two sized.
Why does kmalloc() waste memory instead of giving exact sizes?
Because the slab allocator only maintains a limited set of pre-sized buckets for performance and fragmentation reasons. Serving arbitrary exact sizes would require far more bookkeeping and would fragment memory badly over time.
What is the maximum size I can request from kmalloc()?
It depends on architecture and kernel configuration, but general-purpose kmalloc() buckets typically stop in the low tens of kilobytes range. Anything beyond that limit falls through to the page allocator or should use vmalloc()/kvmalloc() instead.
Is ksize() deprecated?
ksize() still exists and is documented as a reporting-only function, but the kernel now steers new code toward kmalloc_size_roundup() for deciding allocation sizes ahead of time, reserving ksize() mainly for auditing after the fact.
Does this memory wastage affect application performance?
It mainly affects total system memory pressure rather than the speed of the allocation itself. On memory-constrained embedded systems, however, that pressure can trigger swapping or the OOM killer, which very much does affect performance.
Should I always use alloc_pages_exact() for large buffers?
Only when the buffer size does not naturally align to a power-of-two page count and physical contiguity is required. If contiguity is not required, vmalloc() or kvmalloc() are usually simpler and often preferable.
Where can I see real kmalloc bucket sizes on my own machine?
Run cat /proc/slabinfo on a Linux system and look for entries named kmalloc-8, kmalloc-16, kmalloc-32, and so on — these list every active general-purpose bucket along with how many objects are currently in use.
Is this lecture part of a complete free Linux kernel development course?
Yes. This lecture is one part of an ongoing free linux kernel development course and free linux device drivers course covering memory management, IPC, sockets, and driver bring-up from the ground up.
Conclusion
Kernel memory allocation looks simple from the outside — call kmalloc(), get a pointer back — but underneath it hides a bucket-and-page system with real, measurable waste built into its design. Understanding where that waste comes from, and reaching for tools like kmalloc_size_roundup() and alloc_pages_exact() when it matters, is exactly the kind of practical skill that separates driver code that merely works from driver code that is production-ready. Keep following this free linux kernel development course and free embedded systems course as we continue into the different slab allocator implementations in the next lecture.
Next up: SLAB vs SLUB vs SLOB — how the kernel’s slab allocator implementations evolved, and which one your kernel is running today.
Further reading (official sources):
