« Previous Lecture | Next Lecture »
SLUB-only
Kernel Module Demo
EmbeddedPathashala
In the previous lesson of this free Linux kernel development course, we learned about the page allocator (Buddy System Allocator) — the engine that hands out physical memory in whole-page chunks. But most kernel code, including the majority of Linux device drivers, needs much smaller allocations: a 64-byte structure here, a 200-byte buffer there. Allocating a full 4 KB page for a tiny structure would waste enormous amounts of memory. That is exactly the problem the slab allocator solves, and in modern Linux this role is filled by the SLUB allocator.
SLUB allocator
kmalloc kzalloc
free linux kernel development course
free linux device drivers course
kmem_cache
What You Will Learn
- Why the kernel needs a slab allocator on top of the page allocator
- How SLUB organizes objects inside slabs
- The difference between kmalloc, kzalloc, and a custom kmem_cache
- kmalloc size limits and when to switch to vmalloc or the page allocator instead
- Creating and using your own object cache with kmem_cache_create
- Common pitfalls, debugging tools, and best practices
Prerequisites
This lesson assumes you have already completed the page allocator lesson in this free kernel programming course, and that you are comfortable writing and loading a basic kernel module.
Why the Slab Allocator Exists
Imagine a device driver that needs to allocate a small 48-byte structure thousands of times per second as packets arrive. Going to the page allocator every single time, and getting back a full 4 KB page for a 48-byte object, would be both slow and incredibly wasteful. The slab allocator solves this by requesting a handful of pages from the page allocator up front, then slicing each page into many same-sized “objects” that can be handed out and reclaimed almost instantly, with the freed objects kept in a per-CPU cache for the next request.
How SLUB Organizes Memory
SLUB groups objects of the same size into a structure called a kmem_cache. Each cache owns one or more “slabs,” and each slab is simply one or more contiguous pages obtained from the page allocator, carved into equally sized object slots. The kernel maintains a set of generic caches for common power-of-two sizes (8, 16, 32, 64 bytes and upward) which is exactly what powers the everyday kmalloc() call. Drivers that repeatedly allocate the same custom structure can also create their own dedicated cache for even better cache-line locality and performance.
Using kmalloc and kzalloc in a Kernel Module
For most everyday allocations, you do not create a custom cache; you simply call kmalloc() or its zero-initializing variant kzalloc(). Below is a minimal, original example module.
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/slab.h>
struct my_packet_info {
int id;
char label[32];
};
static struct my_packet_info *info;
static int __init slab_demo_init(void)
{
info = kzalloc(sizeof(*info), GFP_KERNEL);
if (!info) {
pr_err("slab_demo: kzalloc failed\n");
return -ENOMEM;
}
info->id = 42;
pr_info("slab_demo: allocated object id=%d\n", info->id);
return 0;
}
static void __exit slab_demo_exit(void)
{
kfree(info);
pr_info("slab_demo: freed object\n");
}
module_init(slab_demo_init);
module_exit(slab_demo_exit);
MODULE_LICENSE("GPL");
kzalloc() is simply kmalloc() followed by an automatic zero-fill, and it is the recommended default unless you have a measured performance reason to skip zeroing.
kmalloc Size Limits
kmalloc is designed for relatively small, physically contiguous allocations. There is an upper practical limit, commonly capped well under 8 MB by default kernel configuration, beyond which requests are likely to fail because finding that much contiguous physical memory becomes increasingly difficult as the system runs. For large buffers that do not need to be physically contiguous, vmalloc() is the appropriate tool instead, since it only needs to be virtually contiguous.
| Allocator | Best For | Physically Contiguous? |
|---|---|---|
| kmalloc / kzalloc | Small, frequent allocations (a few bytes up to a few KB) | Yes |
| kmem_cache (custom) | Repeated allocation of one specific structure type | Yes |
| vmalloc | Large buffers where physical contiguity is not required | No |
| page allocator directly | Whole-page or multi-page DMA buffers | Yes |
Creating a Custom kmem_cache
When a driver allocates and frees the exact same structure type extremely frequently, a dedicated cache reduces overhead and improves cache locality compared to the generic kmalloc caches.
#include <linux/slab.h>
static struct kmem_cache *my_cache;
/* during init */
my_cache = kmem_cache_create("my_packet_cache",
sizeof(struct my_packet_info),
0, SLAB_HWCACHE_ALIGN, NULL);
/* allocate an object */
struct my_packet_info *p = kmem_cache_alloc(my_cache, GFP_KERNEL);
/* free an object */
kmem_cache_free(my_cache, p);
/* during exit, after all objects are freed */
kmem_cache_destroy(my_cache);
Real-World Use Cases
- Networking stack’s per-packet metadata structures (sk_buff related caches)
- Filesystem inode and dentry caches
- Device drivers allocating fixed-size descriptor or request structures repeatedly
- Any short-lived object allocated and freed at high frequency
Common Mistakes and Troubleshooting
- Calling kmalloc for very large sizes: requests beyond the practical limit will fail; switch to vmalloc or the page allocator.
- Forgetting kfree on every code path: leads to kernel memory leaks that accumulate over the system’s uptime.
- Destroying a kmem_cache while objects are still allocated from it: always free all outstanding objects before calling kmem_cache_destroy.
- Using kmalloc for huge DMA buffers: prefer dedicated DMA allocation APIs designed for that purpose instead.
Best Practices
- Default to
kzalloc()overkmalloc()to avoid uninitialized memory bugs. - Create a dedicated kmem_cache only when allocation frequency or structure size justifies the extra setup.
- Always pair every allocation with exactly one matching free.
- Use kernel debugging tools such as slab usage statistics under /proc and kmemleak during development to catch leaks early.
Performance and Security Considerations
SLUB’s per-CPU object caching makes allocation and free operations extremely fast in the common case, avoiding lock contention on multi-core systems. From a security standpoint, object caches that mix allocations of similar sizes are a known target for heap-spray and use-after-free style kernel exploits, which is why hardened kernel configurations enable slab-related hardening features. Always zero sensitive structures with kzalloc() before they could ever be exposed to user space.
Summary / Key Takeaways
- The slab allocator (SLUB) sits above the page allocator and efficiently serves small, frequent allocations.
kmalloc()/kzalloc()are the everyday APIs backed by generic SLUB caches.- kmalloc has a practical upper size limit; use vmalloc for large, non-contiguous buffers.
- Custom kmem_cache objects improve performance for high-frequency, fixed-size allocations.
- Always match every allocation with a free, and destroy caches only after all objects are released.
Conclusion
Together with the page allocator from the previous lesson, the slab allocator completes the picture of how the Linux kernel manages memory internally. Almost every kernel module and device driver you will ever write relies on kmalloc, kzalloc, or a custom kmem_cache somewhere in its code. Mastering both layers gives you the confidence to write efficient, leak-free, production-quality kernel code. Continue to the next free lecture in this Linux kernel and device drivers course to keep building your skills.
Frequently Asked Questions
1. What is the slab allocator in the Linux kernel?
It is the kernel subsystem that slices pages obtained from the page allocator into smaller, fixed-size objects for fast, low-overhead allocation.
2. What is SLUB?
SLUB is the modern implementation of the slab allocator concept used by current mainline Linux kernels.
3. What is the difference between kmalloc and kzalloc?
kzalloc behaves exactly like kmalloc but additionally zero-fills the allocated memory before returning it.
4. When should I use vmalloc instead of kmalloc?
Use vmalloc when you need a large buffer that does not need to be physically contiguous, since kmalloc has practical size limits tied to contiguous memory availability.
5. When should I create my own kmem_cache?
When your driver allocates and frees the exact same structure type at high frequency and would benefit from dedicated, cache-aligned storage.
6. Is kfree required for every kmalloc call?
Yes, every successful allocation must eventually be freed with kfree (or kmem_cache_free for cache-based objects) to avoid kernel memory leaks.
7. Can I destroy a kmem_cache while objects are still allocated?
No, all objects must be freed back to the cache before calling kmem_cache_destroy.
8. Does this lesson cover the latest kernel behavior?
Yes, this lesson reflects current mainline SLUB-based kernels rather than older textbook-era allocator implementations.
This lesson is part of EmbeddedPathashala’s 100% free Linux Kernel Programming and Linux Device Drivers course.
Explore Free Kernel Course
Join Free Embedded Systems Course
