⬅ Previous Lecture
Next Lecture ➡
If you are learning Linux device driver programming, sooner or later you will run into the Linux kernel page allocator. It is the lowest-level memory allocator inside the kernel, and almost every other memory API — including kmalloc() and vmalloc() — is built on top of it. In this lecture of our free Linux kernel development course, we break down how the Linux kernel page allocator actually works, how it maps physical RAM to kernel virtual addresses, and why careless use of the Linux kernel page allocator can silently waste large chunks of system memory through internal fragmentation.
This tutorial is part of EmbeddedPathashala’s free Linux kernel development course and free Linux device drivers course. No prior experience with kernel memory internals is assumed, but a basic understanding of kernel modules will help.
Buddy system algorithm
__get_free_page()
alloc_pages()
Direct-mapped memory
Internal fragmentation
PAGE_OFFSET
Free Linux kernel course
By the end of this lecture you will understand:
- What the Linux kernel page allocator is and why drivers need it
- How the buddy system algorithm organizes free memory
- The core page allocator APIs and when to use each one
- How physical RAM maps to kernel virtual addresses (direct-mapped / lowmem region)
- What internal fragmentation is and why it matters for embedded systems
- Practical best practices to avoid wasting kernel memory
- Basic C programming knowledge
- A working Linux Kernel Module (LKM) build setup
- Familiarity with writing a basic “Hello World” kernel module
- Conceptual idea of virtual memory and physical memory
What Is the Linux Kernel Page Allocator?
The Linux kernel manages physical RAM in fixed-size chunks called pages — typically 4 KB on most architectures (you can confirm this with PAGE_SIZE on your target). The Linux kernel page allocator is the subsystem responsible for handing out and reclaiming these pages whenever the kernel, a driver, or a kernel module needs raw memory.
Unlike user-space malloc(), which can return memory of almost any size, the Linux kernel page allocator always works in units of whole pages, or powers-of-two groups of pages. This design choice keeps the allocator extremely fast, but it also has consequences — particularly around wasted memory — that every driver author should understand.
The Buddy System: How the Page Allocator Organizes Free Memory
Internally, the Linux kernel page allocator uses an algorithm known as the buddy system. Free pages are grouped into “order” levels, where order 0 means a single page, order 1 means 2 pages, order 2 means 4 pages, and so on, up to MAX_ORDER. Each order maintains its own free list per memory zone.
| Order | Block Size | Free List (example) |
|---|---|---|
| 0 | 4 KB (1 page) | page |
| 1 | 8 KB (2 pages) | 2 pages 2 pages |
| 2 | 16 KB (4 pages) | 4 pages |
| 3 | 32 KB (8 pages) | 8 pages 8 pages |
When a request comes in, the allocator finds the smallest order that can satisfy it. If only a larger block is free, that block is split in half repeatedly — each half is called a “buddy” of the other — until a block of the right order is produced. When memory is freed, the allocator checks if the buddy block is also free; if so, the two are merged back into a larger block. This merging is what keeps free memory from becoming permanently scattered.
Core Linux Kernel Page Allocator APIs
Here are the page allocator APIs you will use most often inside a kernel module. (Always remember to free what you allocate, and pair each call with its matching free function.)
| API | Returns | Typical Use |
|---|---|---|
__get_free_page(gfp_mask) |
Virtual address of 1 page | Single page, content uninitialized |
get_zeroed_page(gfp_mask) |
Virtual address of 1 page | Single page, zero-filled |
__get_free_pages(gfp_mask, order) |
Virtual address of 2^order pages | Multiple physically contiguous pages |
alloc_page(gfp_mask) |
struct page * |
When you need the page descriptor itself |
alloc_pages(gfp_mask, order) |
struct page * |
Multiple pages, returned as page descriptor |
A minimal example of allocating and releasing a single page inside a kernel module:
#include <linux/module.h>
#include <linux/gfp.h>
static unsigned long my_page;
static int __init my_pgalloc_init(void)
{
my_page = __get_free_page(GFP_KERNEL);
if (!my_page)
return -ENOMEM;
pr_info("my_pgalloc: allocated page at virtual addr %lx\n", my_page);
return 0;
}
static void __exit my_pgalloc_exit(void)
{
free_page(my_page);
pr_info("my_pgalloc: page freed\n");
}
module_init(my_pgalloc_init);
module_exit(my_pgalloc_exit);
MODULE_LICENSE("GPL");
And requesting several physically contiguous pages using an order value:
unsigned long block;
/* order = 3 means 2^3 = 8 pages = 32 KB on a 4 KB page system */
block = __get_free_pages(GFP_KERNEL, 3);
if (!block)
return -ENOMEM;
/* ... use the memory ... */
free_pages(block, 3);
free_pages() is a common bug that corrupts the buddy system’s free lists.How the Linux Kernel Page Allocator Maps Physical RAM to Virtual Addresses
One of the more interesting things you can verify yourself is that pages returned by the Linux kernel page allocator are physically contiguous, and that there is a simple, predictable relationship between the kernel virtual address (KVA) and the physical address (PA) for memory in the kernel’s direct-mapped (also called “lowmem”) region. On most modern 64-bit systems, this relationship is:
kernel_virtual_address = physical_address + PAGE_OFFSET
PAGE_OFFSET is an architecture-specific constant marking the start of the direct-mapped region in kernel virtual address space. Because of this fixed offset, translating between PA and KVA in this region is just simple arithmetic — no page table walk needed for the kernel itself, although the MMU still performs the actual translation for hardware accesses.
| Page # | Kernel Virtual Address (KVA) | Physical Address (PA) |
|---|---|---|
| 0 | PAGE_OFFSET + 0x0000 | 0x0000 |
| 1 | PAGE_OFFSET + 0x1000 | 0x1000 |
| 2 | PAGE_OFFSET + 0x2000 | 0x2000 |
| … | … | … |
On current mainline kernels (6.x series, including the latest LTS releases), the broad layout concept hasn’t changed, though exact address ranges can shift slightly between kernel versions and depend on whether 4-level or 5-level page tables are in use on x86_64, or the VA size configured on arm64. Rather than memorizing fixed hex ranges, the more future-proof approach is to read the layout directly from your running kernel’s documentation tree (Documentation/arch/x86/x86_64/mm.rst on modern kernels, or the arm64 equivalent) or inspect /proc/iomem and kernel boot logs on your target board.
| Region (low → high) | Purpose |
|---|---|
| User space | Per-process address space, separate for each running task |
| Guard hole | Unmapped gap, traps stray pointer errors and reserved space |
| Direct-mapped / lowmem region | Where the Linux kernel page allocator’s memory lives; PA + PAGE_OFFSET = KVA |
| vmalloc / ioremap space | Virtually contiguous but not necessarily physically contiguous mappings |
| Kernel modules / fixmaps / other reserved regions | Module text/data, per-CPU fixed mappings, and similar special regions |
Internal Fragmentation: The Hidden Cost of the Page Allocator
Here’s where many beginners get caught out. Because the Linux kernel page allocator only hands out memory in power-of-two page counts, any request that isn’t a “nice” power-of-two size gets rounded up — and the difference is wasted. This wastage is called internal fragmentation.
| Requested Size | Order Used | Actually Allocated | Wasted |
|---|---|---|---|
| 132 KB | order 6 | 256 KB | 124 KB |
| 20 KB | order 3 | 32 KB | 12 KB |
| 16 KB | order 2 | 16 KB | 0 KB |
In a one-off allocation, a few wasted kilobytes hardly matter. But on memory-constrained embedded boards, or in drivers that perform many such allocations repeatedly (for example, per-packet buffers in a networking driver), this wastage adds up quickly and can be the difference between a board that runs stably and one that runs into out-of-memory conditions under load.
How to Avoid Internal Fragmentation in Driver Code
- Use
kmalloc()/kzalloc()for small, oddly-sized allocations. The slab/slub allocator sits on top of the page allocator and is specifically designed to efficiently serve small, arbitrary-sized requests without large rounding waste. - Reserve the raw page allocator for allocations that are naturally page-aligned, such as DMA buffers or memory you intend to map into user space.
- Batch small allocations into a single larger structure when you control the data layout, instead of allocating many tiny chunks separately.
- Profile actual usage with tools like
/proc/buddyinfoand/proc/slabinforather than guessing.
Real-World Use Cases
| Use Case | Why the Page Allocator Is Used |
|---|---|
| DMA buffer allocation | DMA-capable hardware often needs physically contiguous memory |
| Network driver ring buffers | Page-aligned, predictable-size buffers for packet descriptors |
| Memory-mapped device buffers | Pages can be mapped directly into user space via mmap() |
| Building custom slab caches | The slab allocator itself sources its raw memory from the page allocator |
Common Mistakes & Troubleshooting
- Forgetting to free pages — leads to kernel memory leaks that persist until reboot.
- Mismatched order on free — corrupts buddy free lists and can crash the system.
- Using
GFP_KERNELin atomic/interrupt context — this flag can sleep; useGFP_ATOMICinstead in such contexts. - Assuming vmalloc’d memory is physically contiguous — it is not; only direct-mapped/page-allocator memory is.
- Hardcoding address ranges from an older kernel — always verify against your actual running kernel version and architecture.
Best Practices
- Always check the return value for
NULL/zero before using allocated memory. - Free memory in the exact reverse order of allocation, in your module’s cleanup path.
- Prefer the smallest API that meets your needs — don’t reach for
alloc_pages()whenkmalloc()would do. - Use
%p(hashed) printk specifiers in production code; avoid exposing raw kernel addresses, which is a kernel security hardening concern.
Performance & Security Considerations
From a performance standpoint, the buddy system’s split/merge approach gives O(log n) allocation and free operations, which is why the Linux kernel page allocator stays fast even under heavy churn. From a security standpoint, modern kernels deliberately randomize and hash displayed kernel addresses (KASLR and printk address hashing) to make it harder for an attacker to use leaked addresses for exploitation — another reason to avoid printing raw addresses outside of controlled debugging sessions.
Summary / Key Takeaways
- The Linux kernel page allocator hands out memory in page-sized, power-of-two units using the buddy system algorithm.
- Direct-mapped (lowmem) kernel memory has a simple, fixed-offset relationship to physical addresses via
PAGE_OFFSET. - Requests that don’t align to a power-of-two page count cause internal fragmentation and wasted memory.
- Use
kmalloc()for small/odd-sized allocations and reserve the raw page allocator for page-aligned, DMA, or mmap use cases.
Conclusion
Understanding the Linux kernel page allocator is foundational to writing efficient, memory-safe kernel modules and device drivers — especially on resource-constrained embedded systems. Once you’re comfortable with how the page allocator and the buddy system work, the rest of the kernel’s memory management subsystems (slab, vmalloc, kmalloc) will make a lot more sense, since they all build on these same fundamentals.
FAQ
Q1. What is the Linux kernel page allocator used for?
It is the core kernel subsystem that allocates and frees physical memory in page-sized units, used directly by drivers and indirectly by higher-level allocators like kmalloc.
Q2. What is the difference between the page allocator and kmalloc?
The page allocator works in whole pages or power-of-two page groups. kmalloc() is built on the slab allocator, which sits on top of the page allocator and is optimized for small, arbitrary-sized allocations.
Q3. Why does the page allocator use the buddy system?
The buddy system allows fast splitting and merging of free memory blocks, which keeps allocation and deallocation efficient while minimizing external fragmentation.
Q4. What is internal fragmentation in the context of the page allocator?
It’s the wasted memory that results when a request is rounded up to the nearest power-of-two page count, since the allocator cannot hand out a partial block.
Q5. Is memory from the page allocator always physically contiguous?
Yes — memory returned by APIs like __get_free_pages() and alloc_pages() is guaranteed to be physically contiguous, unlike vmalloc() memory.
Q6. What does PAGE_OFFSET mean?
It’s the architecture-specific constant marking where the kernel’s direct-mapped (lowmem) virtual address region begins, used to translate between physical and kernel virtual addresses in that region.
Q7. When should I use alloc_pages() instead of kmalloc()?
Use alloc_pages() when you specifically need page-aligned, physically contiguous memory — for example, for DMA buffers or memory you plan to map into user space.
Q8. What happens if I free pages with the wrong order?
It corrupts the buddy system’s internal free lists and can lead to memory corruption or a kernel crash. Always match the order used at allocation and free time.
Q9. Is this tutorial part of a free Linux kernel development course?
Yes, this lecture is part of EmbeddedPathashala’s free Linux kernel development course, which also covers free Linux device drivers and embedded systems content.
Q10. Does the page allocator behave differently on ARM vs x86_64?
The core buddy system algorithm is architecture-independent, but exact constants like PAGE_OFFSET, virtual address ranges, and page table levels differ between architectures such as x86_64 and arm64.
This lecture is part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.
