How Does the Buddy System Work in Linux? – Linux Device Drivers Course Online

« Previous Lecture  |  Next Lecture »

Linux Kernel Page Allocator (Buddy System) Explained
Free Linux Kernel Development Course – Memory Management Internals for Module Authors
Modern Kernel
6.x Series
Hands-On
Kernel Module Demo
100% Free
EmbeddedPathashala

If you are learning Linux kernel programming or writing your first Linux device driver, sooner or later you will need to allocate memory inside kernel space. The kernel cannot call malloc() like a normal user-space application. Instead, every byte of kernel memory ultimately comes from a low-level engine called the page allocator, also known as the Buddy System Allocator (BSA). In this free lesson from our free Linux kernel development course, we break down exactly how the page allocator works, why it exists, and how you can use it safely inside your own kernel modules.

page allocator
buddy system allocator
free linux kernel development course
free linux device drivers course
kernel memory management
__get_free_pages

What You Will Learn

  • What the page allocator (Buddy System Allocator) is and why the kernel needs it
  • How the buddy algorithm splits and merges free memory blocks
  • The difference between page-frame allocation and byte-level allocation
  • How to allocate and free pages safely from a kernel module
  • Real allocation flags (GFP flags) and when to use each one
  • Common mistakes, debugging tips, and best practices

Prerequisites

Before this lesson, you should be comfortable with basic kernel module structure (module_init, module_exit), have a working Linux VM with kernel headers installed, and have a general idea of what physical memory and virtual memory mean. If any of this is new to you, go back to the earlier lectures in this free embedded systems course before continuing.

Why the Kernel Needs a Dedicated Page Allocator

Physical RAM is a finite, shared resource. Every running process, every driver, the filesystem cache, and the kernel itself compete for the same pool of memory. The kernel needs one central authority that hands out physical page frames in a fast, fragmentation-resistant way, and reclaims them just as quickly when they are no longer needed. That central authority is the page allocator.

On modern Linux, almost every kernel memory request eventually funnels down to this same engine, whether it comes from the scheduler, the networking stack, the virtual filesystem layer, or a small device driver you write yourself. Understanding it is therefore one of the highest-leverage topics in kernel programming.

Who Talks To The Page Allocator
[ Scheduler ] [ Filesystem ] [ Networking ] [ Your Driver ]
| all request memory |
v v
============================================
PAGE ALLOCATOR (Buddy System)
============================================
|
v
Physical RAM (page frames)

The Buddy System Algorithm, Explained Simply

The buddy system organizes free physical memory into power-of-two sized blocks: 1 page, 2 pages, 4 pages, 8 pages, and so on, up to a maximum order (typically order 10, meaning 1024 pages, though this is configurable). Each “order” has its own free list. When a request comes in for, say, 4 pages, the allocator looks at the free list for that exact order first.

If no block of the right size is free, the allocator looks one order higher, takes a larger block, and splits it in half. One half satisfies the request; the other half — its “buddy” — goes back onto the free list at the lower order. This splitting can cascade down several levels if needed.

The reverse happens during deallocation. When a block is freed, the kernel checks whether its buddy block is also free. If it is, the two are merged back into a single larger block, and the kernel checks again one level up. This constant splitting and merging is what keeps physical memory from fragmenting into countless unusable tiny pieces over time.

Buddy Splitting Example (Allocating 1 Page From an 8-Page Block)
Order 3 (8 pages): [ free block ]
Order 2 (4 pages): [ used ][ buddy: free ]
Order 1 (2 pages): [used][buddy free]
Order 0 (1 page): [used][buddy free]

Page Allocator API: Using It Inside a Kernel Module

You rarely call the buddy allocator directly by name; instead you use a small family of functions that wrap it. The two you will use most often are __get_free_pages() and its sibling alloc_pages(). Below is a minimal, original example module demonstrating page allocation and the matching free call.

#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/gfp.h>

static unsigned long my_page = 0;

static int __init pagealloc_demo_init(void)
{
    /* request a single physical page, order 0 */
    my_page = __get_free_pages(GFP_KERNEL, 0);
    if (!my_page) {
        pr_err("pagealloc_demo: allocation failed\n");
        return -ENOMEM;
    }
    pr_info("pagealloc_demo: got page at virtual addr 0x%lx\n", my_page);
    return 0;
}

static void __exit pagealloc_demo_exit(void)
{
    if (my_page)
        free_pages(my_page, 0);
    pr_info("pagealloc_demo: freed page\n");
}

module_init(pagealloc_demo_init);
module_exit(pagealloc_demo_exit);
MODULE_LICENSE("GPL");

The second argument to __get_free_pages() and free_pages() is the order — order 0 means one page, order 1 means two pages, order 2 means four pages, and so on. Always pass the exact same order to the free call that you used for the allocation, or you will corrupt the buddy free lists.

Understanding GFP Flags

The first argument, commonly GFP_KERNEL, tells the allocator the context and urgency of your request. Choosing the wrong flag is one of the most common bugs new kernel developers introduce.

GFP Flag When To Use It Can It Sleep?
GFP_KERNEL Normal process-context allocation, most common case Yes
GFP_ATOMIC Inside interrupt handlers or while holding a spinlock No
GFP_NOWAIT Similar to atomic, fails fast without reclaiming memory No
GFP_DMA Memory must be reachable by older DMA-capable hardware Depends

Real-World Use Cases

  • Network drivers allocating DMA buffers for incoming packets
  • Block device drivers reserving large contiguous buffers for I/O
  • The slab allocator itself, which is layered on top of the page allocator (covered in the next lecture)
  • Graphics and video drivers requesting large contiguous frame buffers

Common Mistakes and Troubleshooting

  • Mismatched order on free: always free with the same order you allocated with.
  • Using GFP_KERNEL inside an interrupt handler: this can sleep and will crash or hang your system; use GFP_ATOMIC instead.
  • Requesting very high orders: large contiguous allocations (order 6+) often fail under memory pressure because contiguous free blocks become scarce. Prefer the slab allocator or vmalloc for large, non-contiguous needs.
  • Forgetting to free on the error path: always match every allocation with a corresponding free, including in your module’s error-handling branches.

Best Practices

  • Use the smallest order that satisfies your need; smaller orders are far easier for the allocator to satisfy.
  • Prefer slab-based APIs (kmalloc/kzalloc) for small, frequent, byte-sized allocations instead of raw pages.
  • Always check the return value for NULL/zero before using the memory.
  • Document the GFP flag choice in your code comments so future maintainers understand the calling context.

Performance and Security Considerations

From a performance angle, the buddy system gives O(log n) allocation and free time and naturally reduces external fragmentation through merging. From a security angle, freshly allocated pages are not guaranteed to be zeroed unless you explicitly request a zeroing variant; leaking stale kernel memory contents to user space through an uninitialized page is a classic vulnerability class, so always zero memory that will ever be exposed outside the kernel.

Summary / Key Takeaways

  • The page allocator (Buddy System Allocator) is the single engine that ultimately hands out all physical memory in the Linux kernel.
  • It organizes free memory into power-of-two block sizes called orders, splitting and merging blocks as needed.
  • __get_free_pages() / free_pages() are the core APIs you will use directly in modules.
  • Choosing the correct GFP flag for your execution context is critical for stability.
  • The slab allocator, covered next, builds on top of this same engine to handle smaller allocations efficiently.

Conclusion

The page allocator is the foundation everything else in kernel memory management is built on. Once you understand how the buddy algorithm splits and merges blocks, and how to call the page allocator APIs correctly with the right GFP flags, you have unlocked one of the most important building blocks of Linux kernel and device driver development. In the next free lecture of this course, we move one layer up the stack to the slab allocator, which is what most kernel modules actually use day to day for everyday memory needs.

Frequently Asked Questions

1. What is the page allocator in the Linux kernel?

The page allocator, also called the Buddy System Allocator, is the core kernel subsystem responsible for allocating and freeing physical memory page frames.

2. Why is it called the “buddy” system?

Because every memory block has a paired “buddy” block of the same size; when both buddies become free, they are merged back into one larger block.

3. What is the difference between the page allocator and the slab allocator?

The page allocator hands out memory in whole page-sized chunks, while the slab allocator sits on top of it and slices pages into smaller, fixed-size objects for efficient small allocations.

4. Can I use GFP_KERNEL inside an interrupt handler?

No. GFP_KERNEL may sleep, which is not allowed in interrupt context. Use GFP_ATOMIC instead.

5. What is an allocation “order” in the buddy system?

Order is a power-of-two multiplier of page size; order 0 is one page, order 1 is two pages, order 2 is four pages, and so on.

6. Is page allocator memory zeroed automatically?

Not unless you specifically request a zeroing allocation; otherwise the memory may contain stale data from a previous use.

7. What happens if I free memory with the wrong order?

It corrupts the buddy system’s internal free lists and can crash the kernel or cause silent memory corruption later.

8. Is this lesson updated for modern kernel versions?

Yes, this lesson reflects current mainline kernel behavior rather than older textbook-era kernels, and the concepts apply across recent stable kernel series.

This lesson is part of EmbeddedPathashala’s 100% free Linux Kernel Programming and Linux Device Drivers course.

Explore Free Kernel Course
Join Free Embedded Systems Course

« Previous Lecture  |  Next Lecture »

Leave a Reply

Your email address will not be published. Required fields are marked *