Allocation Layers
APIs Covered
Free Tutorial
Every Linux kernel module or device driver you write eventually needs memory. But unlike user-space programming, where malloc() is basically the only tool you reach for, kernel-space gives you a whole family of allocation APIs — each tuned for a different job. Picking the wrong one can silently waste memory, break on embedded boards with limited RAM, or even crash your system.
This free Linux kernel development course lecture is part of EmbeddedPathashala’s ongoing free Linux device drivers course and free embedded systems course series. In this guide, we break down the kernel’s memory allocation architecture from the ground up and give you a practical decision framework you can apply immediately.
Slab allocator internals
kmalloc vs kzalloc
vmalloc vs kmalloc
kvmalloc explained
devm_ managed APIs
Custom slab caches
DMA-safe allocation
Kernel module basics (insmod/rmmod)
Understanding of virtual memory concepts
A Linux VM or board for testing
Why the Kernel Doesn’t Just Use malloc()
User-space malloc() can be lazy about memory — it can swap pages out, fail gracefully, and let the OS clean up when your process exits. The kernel doesn’t have that luxury. A kernel allocation might run in an interrupt handler where sleeping is illegal, might need to guarantee physically contiguous pages for a DMA-capable device, or might need to avoid fragmenting precious low memory on a resource-constrained embedded board. That’s why the Linux kernel memory allocation subsystem is layered, with each layer solving a specific problem.
The Three-Layer Memory Allocation Architecture
Modern Linux kernels (6.x series and later) still follow the same fundamental three-layer design, though internal implementations have evolved — SLAB was deprecated in favor of SLUB as the default allocator, and kvmalloc() has become the recommended default for many drivers instead of manually choosing between kmalloc and vmalloc.
Object caching on top of pages. APIs: kmalloc(), kzalloc(), kmem_cache_alloc()
Virtually contiguous, physically scattered memory. APIs: vmalloc(), vzalloc()
The foundation. Every allocation ultimately traces back here. APIs: alloc_pages(), __get_free_pages()
Think of it like a warehouse. The page allocator hands out memory in fixed-size crates (pages, typically 4 KB). If your driver only needs a small box of data — say, 64 bytes for a struct — asking for a whole crate wastes space. That’s where the slab allocator steps in: it takes crates from the page allocator and slices them into small, reusable boxes matching common object sizes, drastically cutting internal fragmentation. The vmalloc region solves a different problem entirely — when you need a large chunk of memory that doesn’t have to be physically contiguous, it stitches together scattered physical pages into one virtually contiguous block.
Kernel Memory Allocation APIs: The Complete List
1. kmalloc() and kzalloc() — Your Default Choice
kmalloc() is the workhorse for small-to-medium, physically contiguous allocations, backed by the slab/SLUB allocator. kzalloc() is identical but zeroes the memory before returning it — almost always the safer choice to avoid leaking stale kernel data.
struct my_device_data *data;
data = kzalloc(sizeof(*data), GFP_KERNEL);
if (!data)
return -ENOMEM;
/* ... use data ... */
kfree(data);
GFP_KERNEL when it’s safe to sleep (typical driver probe/init paths), and GFP_ATOMIC only inside interrupt context or when holding a spinlock.2. devm_kmalloc() and devm_kzalloc() — Resource-Managed Allocation
For device drivers, the devm_ family ties memory lifetime to the device itself. When the device is removed or probe fails, the kernel automatically frees this memory — no manual kfree() needed in your remove() path.
struct my_device_data *data;
data = devm_kzalloc(&pdev->dev, sizeof(*data), GFP_KERNEL);
if (!data)
return -ENOMEM;
3. kmem_cache_create() — Custom Slab Caches
If your driver repeatedly allocates and frees objects of the exact same size and type, a dedicated slab cache is far more efficient than generic kmalloc, since it avoids repeated size-class lookups and keeps related objects cache-friendly.
static struct kmem_cache *my_cache;
my_cache = kmem_cache_create("my_driver_objs",
sizeof(struct my_obj),
0, SLAB_HWCACHE_ALIGN, NULL);
struct my_obj *obj = kmem_cache_alloc(my_cache, GFP_KERNEL);
/* ... */
kmem_cache_free(my_cache, obj);
kmem_cache_destroy(my_cache);
4. vmalloc() and vzalloc() — Large, Non-Contiguous Buffers
When you need a large software-only buffer (no DMA involved) and physical contiguity doesn’t matter, vmalloc() lets the kernel stitch together scattered pages. It’s less TLB-friendly than kmalloc, so it’s reserved for cases where kmalloc would fail due to fragmentation.
void *buf = vzalloc(4 * 1024 * 1024); /* 4 MB software buffer */
if (!buf)
return -ENOMEM;
vfree(buf);
5. kvmalloc() — The Modern Recommended Default
In current kernels, kvmalloc() is the preferred way to request memory of unpredictable size. It tries kmalloc() first (fast, contiguous), and transparently falls back to vmalloc() if the allocation is too large or physical memory is too fragmented.
void *buf = kvzalloc(size, GFP_KERNEL);
if (!buf)
return -ENOMEM;
kvfree(buf);
6. alloc_pages() and __get_free_pages() — Raw Page Allocator Access
These are the lowest-level APIs exposed to driver authors, used when you need physically contiguous memory that’s larger than what the slab allocator comfortably provides, and DMA constraints don’t require the dedicated DMA API.
unsigned long addr = __get_free_pages(GFP_KERNEL, 2); /* 2^2 = 4 pages */
if (!addr)
return -ENOMEM;
free_pages(addr, 2);
Decision Framework: Which API Should You Use?
Here’s a simplified way to reason about your choice, expressed as an inline flow rather than an image:
Comparison Table
| API | Backing Layer | Physically Contiguous | Typical Use Case |
|---|---|---|---|
| kmalloc / kzalloc | Slab / SLUB | Yes | Small generic structures |
| devm_kzalloc | Slab / SLUB | Yes | Driver-lifetime allocations |
| kmem_cache_alloc | Custom slab cache | Yes | High-frequency same-size objects |
| vmalloc / vzalloc | vmalloc region | No | Large software-only buffers |
| kvmalloc / kvzalloc | Slab, falls back to vmalloc | Usually | Unpredictable/variable sizes |
| alloc_pages / __get_free_pages | Page allocator | Yes | Large contiguous or DMA buffers |
Real-World Use Cases
- Network drivers often use kvmalloc() for descriptor rings whose size depends on hardware queue configuration.
- Character device drivers commonly use devm_kzalloc() for their per-device private data structure in probe().
- Filesystem and block layer code frequently create dedicated kmem_cache instances (like inode caches) since they allocate the same-sized object millions of times.
- DMA-capable hardware drivers avoid vmalloc/kmalloc entirely for buffers handed to hardware, instead using the DMA-mapping API (dma_alloc_coherent()) to guarantee physically contiguous, cache-coherent memory.
Common Mistakes and Troubleshooting
Best Practices
- Prefer
kzalloc()overkmalloc()unless you have a measured performance reason to skip zeroing. - Prefer
devm_variants inside driver probe() to simplify error handling and prevent leaks. - Use
kvzalloc()when allocation size is variable or could exceed a page, instead of manually branching between kmalloc and vmalloc. - Always check the return value of every allocation — the kernel will not throw an exception, it simply returns NULL.
Performance Considerations
Slab-backed allocations (kmalloc family) are generally faster and more cache-friendly than vmalloc, because vmalloc requires extra page-table manipulation to create a virtually contiguous mapping. On embedded systems with limited RAM, excessive vmalloc use can also fragment the dedicated vmalloc address space, which is often smaller than you’d expect on 32-bit platforms.
Security Considerations
Always zero memory before exposing it to user space (kzalloc, not kmalloc) to avoid leaking stale kernel data through uninitialized fields. Be mindful of integer overflow when computing allocation sizes from user-controlled input — prefer helpers like kmalloc_array(), which check for multiplication overflow internally, over manual multiplication.
Summary / Key Takeaways
- The kernel’s memory architecture has three layers: page allocator, slab/SLUB allocator, and vmalloc region.
- kzalloc() is the right default for small, physically contiguous, driver-lifetime-agnostic allocations.
- devm_kzalloc() is preferred inside driver probe() for automatic cleanup.
- kvzalloc() is the modern go-to for variable or large allocations where DMA is not involved.
- Never use vmalloc-backed memory for DMA buffers.
Conclusion
Choosing the right kernel memory allocation API isn’t about memorizing a list — it’s about understanding what each layer optimizes for: speed and cache-friendliness (slab), flexibility for large buffers (vmalloc), and raw physical guarantees (page allocator). Once that mental model clicks, the “which API do I use?” question answers itself almost every time. In the next lecture, we’ll build on this foundation and look at kernel virtual memory operations in more depth.
Frequently Asked Questions
kmalloc() returns physically contiguous memory backed by the slab allocator and is faster, while vmalloc() returns virtually contiguous but physically scattered memory, useful for large software-only buffers.
Use kzalloc() whenever you don’t have a strong performance reason to skip zeroing — it prevents accidental leaks of stale kernel memory and is the safer default.
The original SLAB allocator has been removed in favor of SLUB as the default slab implementation in recent mainline kernels, though the kmalloc/kzalloc API surface remains unchanged for driver authors.
No. vmalloc memory is not physically contiguous, so it cannot be safely used for DMA. Use the DMA-mapping API instead.
kvmalloc() is a convenience wrapper that tries kmalloc first and automatically falls back to vmalloc if the requested size is too large or physical memory is fragmented, making it the recommended default for variable-size allocations.
devm_kzalloc() ties the allocation’s lifetime to the device, so the kernel automatically frees it on driver removal or probe failure, reducing the risk of memory leaks.
Kernel allocation functions return NULL on failure rather than throwing an exception, so every allocation must be checked and handled with an appropriate error path, typically returning -ENOMEM.
Create a custom slab cache when your driver or subsystem repeatedly allocates and frees objects of the exact same size and type at high frequency, since dedicated caches are faster and reduce fragmentation.
This lecture is part of EmbeddedPathashala’s free Linux kernel programming and device drivers course.
