What Is the Difference Between kmalloc and vmalloc? – Free Linux Device Driver Course

 

Linux Kernel Memory Regions Deep Dive
vmalloc, ioremap, lowmem, and How Kernel Allocators Work
🎓 Free Course
⚙️ Kernel 6.x
🧠 Intermediate
🕐 ~25 min read

Linux Kernel Memory Regions — vmalloc, ioremap and Lowmem Explained

Topics Covered:
vmalloc internals ioremap for device drivers lowmem vs highmem kmalloc vs vmalloc DMA memory Slab allocator Linux 6.x memory APIs

What You Will Learn

  • How the lowmem region works and why physical contiguity matters for DMA
  • What vmalloc does internally — page allocation, page table mapping, and TLB shootdown
  • How ioremap maps device registers into kernel space so drivers can access hardware
  • The slab/slub allocator and why it exists on top of the page allocator
  • Practical decision guide: which allocator to use in which situation
  • How to use devm_ variants for automatic resource cleanup in device drivers
Part of: EmbeddedPathashala’s free Linux kernel development course and free Linux device drivers course. This is the second lecture in the kernel memory management module. If you have not read Part 1 on the Kernel VAS layout, start there first.

In the previous lecture we looked at the big picture — how the Linux kernel divides virtual address space into distinct regions. Now we go deeper. Each region has a specific allocator and a specific set of rules. Knowing which allocator to use, when, and why is one of the most practical skills in Linux kernel programming and Linux device driver development.

The Foundation: The Page Allocator (Buddy System)

Everything in the Linux kernel memory subsystem is ultimately built on top of one thing: the page allocator, also called the buddy allocator. This allocator manages physical memory in units of pages (4 KB on most architectures). When any kernel subsystem needs physical memory, it ultimately calls into the page allocator.

The buddy allocator gets its name from how it works. Physical pages are organized into free lists. Each list holds blocks of pages that are powers of two in size — 1 page, 2 pages, 4 pages, 8 pages, up to 210 (1024) pages. When you request memory, it finds the smallest block that satisfies the request and splits it if needed. When you free memory, it tries to merge the freed block with its “buddy” (the adjacent block of the same size) to form a larger block. Over time this keeps large physically contiguous regions available.

Buddy Allocator — Splitting a Free Block
Request: 1 page. Smallest available free block: 4 pages
Free
4 pages
↓ split into two buddies of 2 pages
Free
2 pages
Free
2 pages
↓ split left buddy into two 1-page buddies
ALLOCATED
1 page
Free
1 page
Free
2 pages
On free: if the 1-page buddy is also free, they merge back to 2 pages, then 4 pages

You call the page allocator directly using functions like alloc_pages(), __get_free_pages(), and free_pages(). But most kernel code does not call these directly — it uses higher-level allocators built on top of the page allocator.

The Slab/SLUB Allocator — Efficient Small Allocations

The page allocator works in page-sized chunks (4 KB minimum). But most kernel data structures are far smaller — a struct task_struct is a few kilobytes, a network socket structure might be a few hundred bytes. Allocating an entire 4 KB page for a 200-byte structure wastes over 95% of the memory.

The slab allocator (modernized as SLUB in Linux 2.6.23 and later) solves this. It pre-allocates pages from the buddy allocator and then carves them up into fixed-size slots. Each type of frequently-used kernel object gets its own slab cache. When you need a new object, SLUB takes the next free slot from the cache. When you free it, SLUB puts it back — ready for the next allocation without going back to the page allocator.

/* Creating a slab cache for your driver's private data structure */
struct my_driver_data {
    int id;
    unsigned long flags;
    void *hw_ptr;
};

static struct kmem_cache *my_cache;

static int __init mydriver_init(void)
{
    /* Create a cache of my_driver_data objects */
    my_cache = kmem_cache_create(
        "my_driver_data",           /* name shown in /proc/slabinfo */
        sizeof(struct my_driver_data), /* object size */
        0,                          /* alignment (0 = natural) */
        SLAB_HWCACHE_ALIGN,         /* flags */
        NULL                        /* constructor (optional) */
    );

    if (!my_cache)
        return -ENOMEM;

    return 0;
}

static void __exit mydriver_exit(void)
{
    kmem_cache_destroy(my_cache);
}

/* Allocating and freeing objects */
struct my_driver_data *obj = kmem_cache_alloc(my_cache, GFP_KERNEL);
if (!obj)
    return -ENOMEM;

/* ... use the object ... */

kmem_cache_free(my_cache, obj);

You can see all active slab caches and their current usage at /proc/slabinfo.

kmalloc() — The General Purpose Kernel Allocator

kmalloc() is the kernel equivalent of user-space malloc(). Under the hood it uses a set of pre-defined SLUB caches for common sizes (8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096, 8192 bytes, and so on). When you call kmalloc(size, flags), the kernel picks the smallest cache slot that fits your size and returns a pointer.

The crucial property of kmalloc() is that the returned memory is physically contiguous. This is required for DMA operations — if your driver tells a hardware DMA engine to copy data to or from a kernel buffer, the DMA engine uses physical addresses and expects the buffer to be a single contiguous block of physical memory.

#include <linux/slab.h>

/* Basic allocation */
void *buf = kmalloc(1024, GFP_KERNEL);
if (!buf)
    return -ENOMEM;

/* Always free what you allocate */
kfree(buf);

/* kzalloc — same as kmalloc but zeroes the memory */
void *zbuf = kzalloc(1024, GFP_KERNEL);

/* krealloc — resize an existing kmalloc buffer */
buf = krealloc(buf, 2048, GFP_KERNEL);

GFP Flags — Telling the Allocator What You Need

The second argument to kmalloc() is a set of flags that tell the allocator about your context and requirements. Getting these wrong causes bugs that are hard to diagnose.

Flag When to Use Can Sleep?
GFP_KERNEL Normal kernel context (process context, can sleep) Yes
GFP_ATOMIC Interrupt handlers, spinlock held, any atomic context No
GFP_DMA DMA operations requiring memory in the low 16 MB (legacy x86 ISA DMA) Yes
GFP_DMA32 DMA to devices that can only address 32-bit physical addresses Yes
GFP_NOWAIT Similar to GFP_ATOMIC — do not sleep, do not reclaim No
GFP_ZERO Add to any flag to zero-fill the allocated memory —
Rule of thumb: If you are in process context and can sleep, use GFP_KERNEL. If you are in interrupt context, a softirq, or holding a spinlock, use GFP_ATOMIC. Using GFP_KERNEL in atomic context causes the kernel to warn or deadlock.

vmalloc() — Large, Virtually-Contiguous Allocations

When you need more memory than kmalloc() can give you in a single call (the per-allocation limit for kmalloc is around 4 MB, though this is configurable) — or when physical contiguity is not required — use vmalloc().

vmalloc() allocates individual physical pages (via the page allocator) and then maps them as a single contiguous virtual range in the vmalloc region of the kernel VAS. To the caller, the memory appears contiguous. In physical memory, the pages may be scattered everywhere.

How vmalloc Maps Scattered Physical Pages into a Contiguous Virtual Range
Kernel Virtual (vmalloc region)
VA: 0xFFFFA000
Page A
VA: 0xFFFFA001
Page B
VA: 0xFFFFA002
Page C
VA: 0xFFFFA003
Page D
← contiguous to caller
⟶
page
tables
Physical Memory (scattered)
PA: 0x12345000
Page A
PA: 0x80011000
Page B
PA: 0x40200000
Page C
PA: 0x1F000000
Page D
← non-contiguous in RAM
#include <linux/vmalloc.h>

/* Allocate 2 MB — fine for vmalloc, too large for kmalloc */
void *buf = vmalloc(2 * 1024 * 1024);
if (!buf)
    return -ENOMEM;

/* Use the buffer — it looks contiguous in virtual space */
memset(buf, 0, 2 * 1024 * 1024);

/* Free when done */
vfree(buf);

The Cost of vmalloc

vmalloc is not free. For each call, the kernel must:

  1. Allocate individual pages from the buddy allocator
  2. Create new page table entries mapping those pages into the vmalloc region
  3. Perform a TLB shootdown — sending IPIs to all CPU cores to flush their TLBs so they pick up the new mappings

This makes vmalloc() considerably slower than kmalloc(). Use it only when you actually need large allocations. Do not use it in hot paths or in frequently-called code.

ioremap() — Mapping Device Registers Into Kernel Space

Every device driver writer needs to understand ioremap(). Hardware devices expose their control registers and status registers through Memory-Mapped I/O (MMIO). On most modern systems, these hardware registers appear at specific physical addresses that are not RAM — they are in a separate physical address range assigned to the device (during PCI enumeration or hardcoded in the SoC memory map for embedded platforms).

To access these registers from a kernel driver, you need a kernel virtual address that maps to the physical address of the registers. That is exactly what ioremap() does — it creates a mapping in the vmalloc region from your physical address range to a kernel virtual address range.

#include <linux/io.h>

#define MY_DEVICE_REGS_PHYS  0xFE200000UL  /* physical base of hardware registers */
#define MY_DEVICE_REGS_SIZE  0x1000         /* 4 KB of register space */

static void __iomem *regs_base;

static int mydev_probe(struct platform_device *pdev)
{
    /* Map the hardware registers into kernel virtual space */
    regs_base = ioremap(MY_DEVICE_REGS_PHYS, MY_DEVICE_REGS_SIZE);
    if (!regs_base) {
        dev_err(&pdev->dev, "Failed to ioremap registers\n");
        return -ENOMEM;
    }

    /* Now read and write registers using the mapped virtual address */
    /* Always use readl/writel — never dereference __iomem pointers directly */
    u32 ctrl = readl(regs_base + 0x00);    /* read control register */
    writel(ctrl | 0x01, regs_base + 0x00); /* set enable bit */

    return 0;
}

static int mydev_remove(struct platform_device *pdev)
{
    /* Always unmap what you mapped */
    iounmap(regs_base);
    return 0;
}
Important: Always mark ioremap pointers with __iomem. Always use readl()/writel() (or the 8/16-bit variants readb/writeb, readw/writew) to access them. Direct dereference of an __iomem pointer is undefined behavior on some architectures. The sparse static analysis tool will warn you if you dereference without the proper accessors.

devm_ioremap() — Automatic Cleanup

Modern Linux drivers (since kernel 3.x) prefer the devm_ variants of resource allocation functions. These are tied to the device’s lifecycle — when the device is removed or the driver fails to probe, the kernel automatically calls the corresponding cleanup function.

/* devm_ioremap — automatically calls iounmap when device is removed */
regs_base = devm_ioremap(&pdev->dev, MY_DEVICE_REGS_PHYS, MY_DEVICE_REGS_SIZE);
if (!regs_base)
    return -ENOMEM;
/* No need to call iounmap() in remove() — kernel does it automatically */

For platform devices, the even cleaner approach is to use devm_platform_ioremap_resource(), which reads the MMIO range directly from the device tree or ACPI tables:

/* Get resource from device tree / ACPI, then map it — all in one call */
regs_base = devm_platform_ioremap_resource(pdev, 0);
if (IS_ERR(regs_base))
    return PTR_ERR(regs_base);

DMA Memory Allocation

Direct Memory Access (DMA) allows hardware devices to transfer data directly to and from system RAM without involving the CPU. This is how network cards, storage controllers, and many other devices achieve high throughput. DMA imposes specific requirements on the memory buffer:

  • The buffer must be physically contiguous (the hardware DMA engine works with physical addresses)
  • The buffer must be in a physical address range the device can reach (some older devices can only address 32-bit or even 24-bit physical addresses)
  • The buffer must be cache-coherent — if the CPU caches the data but the device writes new data directly to RAM, the CPU could read stale cached data
#include <linux/dma-mapping.h>

struct device *dev = &pdev->dev;
dma_addr_t dma_handle;
size_t size = 4096;

/* Allocate a DMA-coherent buffer */
void *cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, GFP_KERNEL);
if (!cpu_addr)
    return -ENOMEM;

/*
 * cpu_addr  — virtual address for CPU to use
 * dma_handle — physical/bus address to give to the hardware DMA engine
 */

/* Give dma_handle to the hardware to set up DMA transfer */
writel((u32)dma_handle, regs_base + DMA_ADDR_REG);
writel(size,             regs_base + DMA_LEN_REG);

/* When done */
dma_free_coherent(dev, size, cpu_addr, dma_handle);
DMA Transfer Flow — Device to Memory
🖥️
Hardware Device
e.g., NIC, Storage
Bus Address (dma_handle)
⟶
DMA engine uses
physical address
💾
DMA Buffer
Physically contiguous
Cache-coherent
Virtual Address (cpu_addr)
⟶
CPU reads data via
kernel virtual addr
⚙️
Driver Code
Processes received data

Practical Decision Guide: Which Allocator to Use

Situation Use This Reason
Small alloc (<4 KB), process context kmalloc(size, GFP_KERNEL) Fast, physically contiguous, slab-backed
Small alloc, interrupt / atomic context kmalloc(size, GFP_ATOMIC) Non-sleeping path
Zeroed small alloc kzalloc(size, GFP_KERNEL) Same as kmalloc + memset(0)
Large alloc (>4 MB), no DMA needed vmalloc(size) Can be physically scattered
DMA buffer (device transfers) dma_alloc_coherent() Physically contiguous, cache-coherent
Map hardware MMIO registers devm_ioremap() Creates kernel VA for device registers
Frequent fixed-size objects kmem_cache_create() + kmem_cache_alloc() Dedicated slab cache, best performance
Whole pages, custom use __get_free_pages() / alloc_pages() Direct page allocator access

Understanding Allocation Limits in the Linux Kernel

One common question from new kernel developers is: how much can I allocate in a single call? The answer depends on the allocator:

  • kmalloc() — practical maximum is around 4 MB (222 bytes, or KMALLOC_MAX_SIZE). For anything larger, the physically-contiguous requirement makes allocation increasingly likely to fail as memory gets fragmented.
  • vmalloc() — limited only by the size of the vmalloc region (tens of GB on 64-bit systems) and available physical pages. Individual calls can be hundreds of megabytes.
  • dma_alloc_coherent() — limited by contiguous physical memory availability and the device’s DMA address mask. For large DMA buffers, consider using the CMA (Contiguous Memory Allocator) which reserves a region of contiguous physical RAM at boot time.

Debugging Kernel Memory Issues

Kernel memory bugs are among the most dangerous — they can cause silent data corruption, intermittent crashes, or security vulnerabilities. Linux provides several built-in tools to catch them.

KASAN — Kernel Address Sanitizer

KASAN is a dynamic memory error detector built into the kernel. Enable it with CONFIG_KASAN=y. It instruments all memory accesses and detects:

  • Use-after-free bugs (accessing memory after kfree())
  • Out-of-bounds reads and writes
  • Stack buffer overflows

SLUB Debug

Boot with slub_debug=FZ on the kernel command line to enable SLUB’s built-in corruption detection. It fills freed objects with a poison pattern and verifies the pattern before reuse — catching use-after-free bugs reliably.

/proc/meminfo and /proc/slabinfo

Monitor your driver’s memory usage during development using these proc files. Watch VmallocUsed grow when your driver maps resources, and check /proc/slabinfo if you use custom slab caches.

Key Takeaways

  • All kernel memory ultimately comes from the buddy (page) allocator
  • The SLUB allocator sits on top and provides efficient small-object allocation
  • kmalloc() is fast and physically contiguous — use it for small allocations, DMA-accessible buffers, and atomic contexts with GFP_ATOMIC
  • vmalloc() handles large allocations but is slower due to page table setup and TLB shootdown
  • ioremap() maps device hardware registers into the kernel’s vmalloc region — every device driver that accesses MMIO uses this
  • dma_alloc_coherent() is the correct API for DMA buffers — never use vmalloc for DMA
  • Prefer devm_ variants in device drivers to eliminate resource leak bugs

Frequently Asked Questions

Q1. What is the difference between kmalloc and vmalloc in the Linux kernel?

kmalloc() allocates physically contiguous memory from the lowmem region — it is fast and suitable for DMA. vmalloc() allocates physically scattered pages and maps them as a contiguous virtual range in the vmalloc region — it can satisfy large allocations but is slower and cannot be used for DMA. Choose based on whether you need physical contiguity and how large the allocation is.

Q2. Why can’t I use vmalloc memory for DMA?

DMA engines work with physical addresses. vmalloc() memory is physically scattered — individual pages are at non-contiguous physical addresses. A DMA engine cannot follow these scattered physical addresses. You need a single contiguous physical buffer, which is what dma_alloc_coherent() provides.

Q3. What does ioremap do and when should I use it in a driver?

ioremap() creates a kernel virtual address mapping for a physical MMIO range (device registers). Use it whenever your driver needs to access hardware registers that appear at a fixed physical address. Always use readl()/writel() accessors to read/write through the mapped address, and always call iounmap() in your cleanup path.

Q4. What is GFP_ATOMIC and when is it required?

GFP_ATOMIC tells the memory allocator that it must not sleep. Use it whenever you need to allocate memory from a context where sleeping is not allowed — interrupt handlers, softirqs, tasklets, and any code holding a spinlock. Using GFP_KERNEL in these contexts can deadlock the system.

Q5. What is KASAN and how does it help in Linux kernel driver development?

KASAN (Kernel Address Sanitizer) is a kernel build option (CONFIG_KASAN=y) that instruments memory accesses to detect use-after-free, out-of-bounds reads/writes, and stack overflows at runtime. It is invaluable during driver development to catch memory corruption bugs early. Enable it in your development kernel, run your driver, and KASAN reports bugs with detailed stack traces.

Q6. What is the benefit of devm_ allocation functions?

The devm_ (device-managed) variants of allocation functions — like devm_kmalloc(), devm_ioremap() — automatically release the resource when the device is removed or when the driver’s probe function fails. This eliminates an entire class of resource leak bugs and makes driver cleanup code much simpler.

Q7. What is the maximum size I can allocate with kmalloc?

The practical maximum for a single kmalloc() call is around 4 MB on most configurations (defined by KMALLOC_MAX_SIZE). Beyond this, the requirement for physical contiguity makes success very unlikely on a system that has been running for a while. For larger needs, use vmalloc() or dma_alloc_coherent() with CMA for large DMA buffers.

Master Linux Kernel Programming — For Free

This lecture is part of EmbeddedPathashala’s completely free Linux kernel development course and free Linux device drivers course. No fees, no registration, no limits.

🏠 Course Index 📖 Next: Process Memory and Page Tables

Leave a Reply

Your email address will not be published. Required fields are marked *