Linux Kernel Memory Regions — vmalloc, ioremap and Lowmem Explained
What You Will Learn
- How the lowmem region works and why physical contiguity matters for DMA
- What vmalloc does internally — page allocation, page table mapping, and TLB shootdown
- How ioremap maps device registers into kernel space so drivers can access hardware
- The slab/slub allocator and why it exists on top of the page allocator
- Practical decision guide: which allocator to use in which situation
- How to use devm_ variants for automatic resource cleanup in device drivers
In the previous lecture we looked at the big picture — how the Linux kernel divides virtual address space into distinct regions. Now we go deeper. Each region has a specific allocator and a specific set of rules. Knowing which allocator to use, when, and why is one of the most practical skills in Linux kernel programming and Linux device driver development.
The Foundation: The Page Allocator (Buddy System)
Everything in the Linux kernel memory subsystem is ultimately built on top of one thing: the page allocator, also called the buddy allocator. This allocator manages physical memory in units of pages (4 KB on most architectures). When any kernel subsystem needs physical memory, it ultimately calls into the page allocator.
The buddy allocator gets its name from how it works. Physical pages are organized into free lists. Each list holds blocks of pages that are powers of two in size — 1 page, 2 pages, 4 pages, 8 pages, up to 210 (1024) pages. When you request memory, it finds the smallest block that satisfies the request and splits it if needed. When you free memory, it tries to merge the freed block with its “buddy” (the adjacent block of the same size) to form a larger block. Over time this keeps large physically contiguous regions available.
4 pages
2 pages
2 pages
1 page
1 page
2 pages
You call the page allocator directly using functions like alloc_pages(), __get_free_pages(), and free_pages(). But most kernel code does not call these directly — it uses higher-level allocators built on top of the page allocator.
The Slab/SLUB Allocator — Efficient Small Allocations
The page allocator works in page-sized chunks (4 KB minimum). But most kernel data structures are far smaller — a struct task_struct is a few kilobytes, a network socket structure might be a few hundred bytes. Allocating an entire 4 KB page for a 200-byte structure wastes over 95% of the memory.
The slab allocator (modernized as SLUB in Linux 2.6.23 and later) solves this. It pre-allocates pages from the buddy allocator and then carves them up into fixed-size slots. Each type of frequently-used kernel object gets its own slab cache. When you need a new object, SLUB takes the next free slot from the cache. When you free it, SLUB puts it back — ready for the next allocation without going back to the page allocator.
/* Creating a slab cache for your driver's private data structure */
struct my_driver_data {
int id;
unsigned long flags;
void *hw_ptr;
};
static struct kmem_cache *my_cache;
static int __init mydriver_init(void)
{
/* Create a cache of my_driver_data objects */
my_cache = kmem_cache_create(
"my_driver_data", /* name shown in /proc/slabinfo */
sizeof(struct my_driver_data), /* object size */
0, /* alignment (0 = natural) */
SLAB_HWCACHE_ALIGN, /* flags */
NULL /* constructor (optional) */
);
if (!my_cache)
return -ENOMEM;
return 0;
}
static void __exit mydriver_exit(void)
{
kmem_cache_destroy(my_cache);
}
/* Allocating and freeing objects */
struct my_driver_data *obj = kmem_cache_alloc(my_cache, GFP_KERNEL);
if (!obj)
return -ENOMEM;
/* ... use the object ... */
kmem_cache_free(my_cache, obj);
You can see all active slab caches and their current usage at /proc/slabinfo.
kmalloc() — The General Purpose Kernel Allocator
kmalloc() is the kernel equivalent of user-space malloc(). Under the hood it uses a set of pre-defined SLUB caches for common sizes (8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096, 8192 bytes, and so on). When you call kmalloc(size, flags), the kernel picks the smallest cache slot that fits your size and returns a pointer.
The crucial property of kmalloc() is that the returned memory is physically contiguous. This is required for DMA operations — if your driver tells a hardware DMA engine to copy data to or from a kernel buffer, the DMA engine uses physical addresses and expects the buffer to be a single contiguous block of physical memory.
#include <linux/slab.h>
/* Basic allocation */
void *buf = kmalloc(1024, GFP_KERNEL);
if (!buf)
return -ENOMEM;
/* Always free what you allocate */
kfree(buf);
/* kzalloc — same as kmalloc but zeroes the memory */
void *zbuf = kzalloc(1024, GFP_KERNEL);
/* krealloc — resize an existing kmalloc buffer */
buf = krealloc(buf, 2048, GFP_KERNEL);
GFP Flags — Telling the Allocator What You Need
The second argument to kmalloc() is a set of flags that tell the allocator about your context and requirements. Getting these wrong causes bugs that are hard to diagnose.
| Flag | When to Use | Can Sleep? |
|---|---|---|
GFP_KERNEL |
Normal kernel context (process context, can sleep) | Yes |
GFP_ATOMIC |
Interrupt handlers, spinlock held, any atomic context | No |
GFP_DMA |
DMA operations requiring memory in the low 16 MB (legacy x86 ISA DMA) | Yes |
GFP_DMA32 |
DMA to devices that can only address 32-bit physical addresses | Yes |
GFP_NOWAIT |
Similar to GFP_ATOMIC — do not sleep, do not reclaim | No |
GFP_ZERO |
Add to any flag to zero-fill the allocated memory | — |
GFP_KERNEL. If you are in interrupt context, a softirq, or holding a spinlock, use GFP_ATOMIC. Using GFP_KERNEL in atomic context causes the kernel to warn or deadlock.vmalloc() — Large, Virtually-Contiguous Allocations
When you need more memory than kmalloc() can give you in a single call (the per-allocation limit for kmalloc is around 4 MB, though this is configurable) — or when physical contiguity is not required — use vmalloc().
vmalloc() allocates individual physical pages (via the page allocator) and then maps them as a single contiguous virtual range in the vmalloc region of the kernel VAS. To the caller, the memory appears contiguous. In physical memory, the pages may be scattered everywhere.
Page A
Page B
Page C
Page D
page
tables
Page A
Page B
Page C
Page D
#include <linux/vmalloc.h>
/* Allocate 2 MB — fine for vmalloc, too large for kmalloc */
void *buf = vmalloc(2 * 1024 * 1024);
if (!buf)
return -ENOMEM;
/* Use the buffer — it looks contiguous in virtual space */
memset(buf, 0, 2 * 1024 * 1024);
/* Free when done */
vfree(buf);
The Cost of vmalloc
vmalloc is not free. For each call, the kernel must:
- Allocate individual pages from the buddy allocator
- Create new page table entries mapping those pages into the vmalloc region
- Perform a TLB shootdown — sending IPIs to all CPU cores to flush their TLBs so they pick up the new mappings
This makes vmalloc() considerably slower than kmalloc(). Use it only when you actually need large allocations. Do not use it in hot paths or in frequently-called code.
ioremap() — Mapping Device Registers Into Kernel Space
Every device driver writer needs to understand ioremap(). Hardware devices expose their control registers and status registers through Memory-Mapped I/O (MMIO). On most modern systems, these hardware registers appear at specific physical addresses that are not RAM — they are in a separate physical address range assigned to the device (during PCI enumeration or hardcoded in the SoC memory map for embedded platforms).
To access these registers from a kernel driver, you need a kernel virtual address that maps to the physical address of the registers. That is exactly what ioremap() does — it creates a mapping in the vmalloc region from your physical address range to a kernel virtual address range.
#include <linux/io.h>
#define MY_DEVICE_REGS_PHYS 0xFE200000UL /* physical base of hardware registers */
#define MY_DEVICE_REGS_SIZE 0x1000 /* 4 KB of register space */
static void __iomem *regs_base;
static int mydev_probe(struct platform_device *pdev)
{
/* Map the hardware registers into kernel virtual space */
regs_base = ioremap(MY_DEVICE_REGS_PHYS, MY_DEVICE_REGS_SIZE);
if (!regs_base) {
dev_err(&pdev->dev, "Failed to ioremap registers\n");
return -ENOMEM;
}
/* Now read and write registers using the mapped virtual address */
/* Always use readl/writel — never dereference __iomem pointers directly */
u32 ctrl = readl(regs_base + 0x00); /* read control register */
writel(ctrl | 0x01, regs_base + 0x00); /* set enable bit */
return 0;
}
static int mydev_remove(struct platform_device *pdev)
{
/* Always unmap what you mapped */
iounmap(regs_base);
return 0;
}
__iomem. Always use readl()/writel() (or the 8/16-bit variants readb/writeb, readw/writew) to access them. Direct dereference of an __iomem pointer is undefined behavior on some architectures. The sparse static analysis tool will warn you if you dereference without the proper accessors.devm_ioremap() — Automatic Cleanup
Modern Linux drivers (since kernel 3.x) prefer the devm_ variants of resource allocation functions. These are tied to the device’s lifecycle — when the device is removed or the driver fails to probe, the kernel automatically calls the corresponding cleanup function.
/* devm_ioremap — automatically calls iounmap when device is removed */
regs_base = devm_ioremap(&pdev->dev, MY_DEVICE_REGS_PHYS, MY_DEVICE_REGS_SIZE);
if (!regs_base)
return -ENOMEM;
/* No need to call iounmap() in remove() — kernel does it automatically */
For platform devices, the even cleaner approach is to use devm_platform_ioremap_resource(), which reads the MMIO range directly from the device tree or ACPI tables:
/* Get resource from device tree / ACPI, then map it — all in one call */
regs_base = devm_platform_ioremap_resource(pdev, 0);
if (IS_ERR(regs_base))
return PTR_ERR(regs_base);
DMA Memory Allocation
Direct Memory Access (DMA) allows hardware devices to transfer data directly to and from system RAM without involving the CPU. This is how network cards, storage controllers, and many other devices achieve high throughput. DMA imposes specific requirements on the memory buffer:
- The buffer must be physically contiguous (the hardware DMA engine works with physical addresses)
- The buffer must be in a physical address range the device can reach (some older devices can only address 32-bit or even 24-bit physical addresses)
- The buffer must be cache-coherent — if the CPU caches the data but the device writes new data directly to RAM, the CPU could read stale cached data
#include <linux/dma-mapping.h>
struct device *dev = &pdev->dev;
dma_addr_t dma_handle;
size_t size = 4096;
/* Allocate a DMA-coherent buffer */
void *cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, GFP_KERNEL);
if (!cpu_addr)
return -ENOMEM;
/*
* cpu_addr — virtual address for CPU to use
* dma_handle — physical/bus address to give to the hardware DMA engine
*/
/* Give dma_handle to the hardware to set up DMA transfer */
writel((u32)dma_handle, regs_base + DMA_ADDR_REG);
writel(size, regs_base + DMA_LEN_REG);
/* When done */
dma_free_coherent(dev, size, cpu_addr, dma_handle);
physical address
Cache-coherent
kernel virtual addr
Practical Decision Guide: Which Allocator to Use
| Situation | Use This | Reason |
|---|---|---|
| Small alloc (<4 KB), process context | kmalloc(size, GFP_KERNEL) |
Fast, physically contiguous, slab-backed |
| Small alloc, interrupt / atomic context | kmalloc(size, GFP_ATOMIC) |
Non-sleeping path |
| Zeroed small alloc | kzalloc(size, GFP_KERNEL) |
Same as kmalloc + memset(0) |
| Large alloc (>4 MB), no DMA needed | vmalloc(size) |
Can be physically scattered |
| DMA buffer (device transfers) | dma_alloc_coherent() |
Physically contiguous, cache-coherent |
| Map hardware MMIO registers | devm_ioremap() |
Creates kernel VA for device registers |
| Frequent fixed-size objects | kmem_cache_create() + kmem_cache_alloc() |
Dedicated slab cache, best performance |
| Whole pages, custom use | __get_free_pages() / alloc_pages() |
Direct page allocator access |
Understanding Allocation Limits in the Linux Kernel
One common question from new kernel developers is: how much can I allocate in a single call? The answer depends on the allocator:
kmalloc()— practical maximum is around 4 MB (222 bytes, orKMALLOC_MAX_SIZE). For anything larger, the physically-contiguous requirement makes allocation increasingly likely to fail as memory gets fragmented.vmalloc()— limited only by the size of the vmalloc region (tens of GB on 64-bit systems) and available physical pages. Individual calls can be hundreds of megabytes.dma_alloc_coherent()— limited by contiguous physical memory availability and the device’s DMA address mask. For large DMA buffers, consider using the CMA (Contiguous Memory Allocator) which reserves a region of contiguous physical RAM at boot time.
Debugging Kernel Memory Issues
Kernel memory bugs are among the most dangerous — they can cause silent data corruption, intermittent crashes, or security vulnerabilities. Linux provides several built-in tools to catch them.
KASAN — Kernel Address Sanitizer
KASAN is a dynamic memory error detector built into the kernel. Enable it with CONFIG_KASAN=y. It instruments all memory accesses and detects:
- Use-after-free bugs (accessing memory after
kfree()) - Out-of-bounds reads and writes
- Stack buffer overflows
SLUB Debug
Boot with slub_debug=FZ on the kernel command line to enable SLUB’s built-in corruption detection. It fills freed objects with a poison pattern and verifies the pattern before reuse — catching use-after-free bugs reliably.
/proc/meminfo and /proc/slabinfo
Monitor your driver’s memory usage during development using these proc files. Watch VmallocUsed grow when your driver maps resources, and check /proc/slabinfo if you use custom slab caches.
Key Takeaways
- All kernel memory ultimately comes from the buddy (page) allocator
- The SLUB allocator sits on top and provides efficient small-object allocation
kmalloc()is fast and physically contiguous — use it for small allocations, DMA-accessible buffers, and atomic contexts with GFP_ATOMICvmalloc()handles large allocations but is slower due to page table setup and TLB shootdownioremap()maps device hardware registers into the kernel’s vmalloc region — every device driver that accesses MMIO uses thisdma_alloc_coherent()is the correct API for DMA buffers — never usevmallocfor DMA- Prefer
devm_variants in device drivers to eliminate resource leak bugs
Frequently Asked Questions
kmalloc() allocates physically contiguous memory from the lowmem region — it is fast and suitable for DMA. vmalloc() allocates physically scattered pages and maps them as a contiguous virtual range in the vmalloc region — it can satisfy large allocations but is slower and cannot be used for DMA. Choose based on whether you need physical contiguity and how large the allocation is.
DMA engines work with physical addresses. vmalloc() memory is physically scattered — individual pages are at non-contiguous physical addresses. A DMA engine cannot follow these scattered physical addresses. You need a single contiguous physical buffer, which is what dma_alloc_coherent() provides.
ioremap() creates a kernel virtual address mapping for a physical MMIO range (device registers). Use it whenever your driver needs to access hardware registers that appear at a fixed physical address. Always use readl()/writel() accessors to read/write through the mapped address, and always call iounmap() in your cleanup path.
GFP_ATOMIC tells the memory allocator that it must not sleep. Use it whenever you need to allocate memory from a context where sleeping is not allowed — interrupt handlers, softirqs, tasklets, and any code holding a spinlock. Using GFP_KERNEL in these contexts can deadlock the system.
KASAN (Kernel Address Sanitizer) is a kernel build option (CONFIG_KASAN=y) that instruments memory accesses to detect use-after-free, out-of-bounds reads/writes, and stack overflows at runtime. It is invaluable during driver development to catch memory corruption bugs early. Enable it in your development kernel, run your driver, and KASAN reports bugs with detailed stack traces.
The devm_ (device-managed) variants of allocation functions — like devm_kmalloc(), devm_ioremap() — automatically release the resource when the device is removed or when the driver’s probe function fails. This eliminates an entire class of resource leak bugs and makes driver cleanup code much simpler.
The practical maximum for a single kmalloc() call is around 4 MB on most configurations (defined by KMALLOC_MAX_SIZE). Beyond this, the requirement for physical contiguity makes success very unlikely on a system that has been running for a while. For larger needs, use vmalloc() or dma_alloc_coherent() with CMA for large DMA buffers.
Master Linux Kernel Programming — For Free
This lecture is part of EmbeddedPathashala’s completely free Linux kernel development course and free Linux device drivers course. No fees, no registration, no limits.
🏠 Course Index 📖 Next: Process Memory and Page Tables