Kernel Course Lesson
Hardware Explained
Modern Kernel Focus
Friendly
Demand paging is one of the most important yet least understood concepts in Linux kernel memory management, and this lesson from our free Linux kernel programming course explains it from first principles. When a kernel API like vmalloc() or a user space malloc() call succeeds, it does not mean physical RAM has actually been handed to you. It only means a virtual address range has been reserved. Physical memory shows up later, one page at a time, only when that memory is actually touched.
This free Linux device drivers course lesson walks through exactly how that “later” moment works, using the CPU’s memory management hardware and the kernel’s page fault handler.
Role of the MMU and TLB
What happens on a TLB miss
The page fault handler
Lazy vs immediate allocation
Why vmalloc uses demand paging
Why slab/buddy allocations don’t
This lesson builds directly on the previous lecture in our free Linux kernel development course. Before continuing, you should understand:
vmalloc() and kmalloc() basics
What a page frame is
Part of EmbeddedPathashala’s free courses:
Free Linux Device Drivers Course
Free Embedded Systems Course
Kernel Memory Management
What Is Demand Paging?
Demand paging, sometimes called lazy allocation, is a memory management strategy where the operating system delays assigning a physical page frame to a virtual page until the program actually accesses that memory. This applies to both kernel-side vmalloc() allocations and user-space malloc() allocations. The allocator’s job at call time is only to reserve a range of virtual addresses; the real work of finding a physical page frame is postponed until it is genuinely needed.
Why Demand Paging Matters for Driver Developers
If you are working through our free Linux device drivers course, understanding demand paging changes how you think about memory usage. A large vmalloc() call that succeeds does not immediately consume physical RAM — RAM usage grows gradually as your driver actually reads or writes into the buffer. This has real consequences for how you size buffers and how you interpret memory usage tools while debugging.
The Hardware Side: MMU and TLB
Every virtual address your kernel code touches has to be translated into a physical address before the CPU can actually read or write memory. This translation is handled by the Memory Management Unit (MMU), a dedicated piece of hardware built into the CPU core.
To avoid walking page tables on every single memory access, the MMU keeps a small hardware cache of recent virtual-to-physical translations called the Translation Lookaside Buffer (TLB). If the translation is already cached, this is called a TLB hit, and the access completes immediately. If not, it is a TLB miss, and the MMU must walk the page tables to find the mapping.
What Happens When No Physical Page Exists Yet?
Here is the crucial part for demand paging: sometimes the page table walk itself fails to find a matching physical address, because — as with a fresh vmalloc() allocation — no physical page frame has been assigned yet. The MMU cannot resolve this on its own, so it raises a page fault exception, handing control to the kernel’s page fault handler, which runs in the context of the process that triggered the access.
The page fault handler requests a single physical page frame, at allocation order zero, from the kernel’s page (buddy system) allocator, and maps it into the faulting virtual address. From this point on, that specific page has real physical memory behind it — but only that one page, not the entire original allocation.
Demand Paging vs Immediate Allocation: A Comparison
| Allocator Type | Allocation Timing | Example |
|---|---|---|
| Page / Buddy Allocator | Immediate, at call time | alloc_pages() |
| Slab Allocator | Immediate, at call time | kmalloc() |
| vmalloc-based / user malloc() | Deferred, on first access | vmalloc(), glibc malloc() |
Notice that kmalloc() and the low-level page allocator do not use demand paging — physical page frames are assigned immediately, because the buddy system’s freelists already map all usable physical RAM into the kernel’s low-memory region, making the assignment essentially free in terms of extra work.
A Small Demonstration: Observing Lazy Allocation
You can observe demand paging indirectly by comparing memory reported immediately after a large vmalloc() style allocation versus after the buffer is actually written to. Here is an original demo module for this lesson that touches memory in two separate stages so you can watch resident memory grow between them using /proc/meminfo.
#include <linux/init.h>
#include <linux/module.h>
#include <linux/vmalloc.h>
#include <linux/delay.h>
#define EP_BIG_SIZE (16 * 1024 * 1024)
static char *ep_bigbuf;
static int __init ep_demand_paging_demo_init(void)
{
ep_bigbuf = vmalloc(EP_BIG_SIZE);
if (!ep_bigbuf)
return -ENOMEM;
pr_info("ep_demo: 16MB reserved virtually, check /proc/meminfo now\n");
ssleep(5);
memset(ep_bigbuf, 0xAA, EP_BIG_SIZE);
pr_info("ep_demo: buffer fully touched, physical pages now assigned\n");
return 0;
}
static void __exit ep_demand_paging_demo_exit(void)
{
vfree(ep_bigbuf);
pr_info("ep_demo: buffer freed\n");
}
module_init(ep_demand_paging_demo_init);
module_exit(ep_demand_paging_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala demand paging demo");
Load the module and, during the five-second sleep window, watch memory usage in another terminal:
sudo insmod ep_demand_paging_demo.ko
watch -n 1 grep -E "MemFree|VmallocUsed" /proc/meminfo
Common Mistakes and Troubleshooting
- Assuming a successful allocation means memory is “used”: free memory only drops as pages are actually touched.
- Translating vmalloc virtual addresses to physical addresses manually: this is unsafe. Only direct-mapped (lowmem) addresses from the page or slab allocators can be safely translated this way.
- Confusing demand paging with swapping: they are related but distinct mechanisms — demand paging is about first-time allocation, while swapping reclaims already-used pages under memory pressure.
- Ignoring page fault overhead in hot paths: code that touches a huge vmalloc buffer for the first time on a performance-critical path will pay page fault costs at the worst possible moment.
Best Practices
- Pre-touch large vmalloc buffers during initialisation if you need predictable, fault-free performance later.
- Use immediate allocators like kmalloc() or the page allocator when you need guaranteed physical memory up front.
- Monitor /proc/meminfo and /proc/vmallocinfo when debugging unexpected memory growth in drivers.
Performance and Security Considerations
Performance: the first access to any demand-paged region costs extra time due to the page fault trap and handler execution. Subsequent accesses to the same page are fast, since the mapping is now cached in the TLB.
Security: because physical frames are assigned lazily and can come from anywhere in RAM, kernel code must never assume a fixed relationship between a vmalloc virtual address and any physical address, and must never leak raw virtual addresses that could aid attackers in probing kernel memory layout.
Summary and Key Takeaways
- Demand paging defers physical memory assignment until first access.
- The MMU and TLB handle translation in hardware; a TLB miss triggers a page table walk.
- A page fault occurs when no physical mapping exists yet, invoking the kernel’s page fault handler.
- vmalloc() and user-space malloc() both rely on demand paging; kmalloc() and the buddy allocator do not.
Conclusion
This lesson of our free Linux kernel development course explained how demand paging bridges the gap between a virtual memory reservation and real physical RAM, using the MMU, TLB, and the kernel’s page fault handler. Combined with the previous lesson on vmalloc(), vzalloc(), and kvmalloc(), you now have a solid, modern-kernel understanding of how kernel memory allocation truly works under the hood — a foundational skill for our free Linux device drivers course and free embedded systems course.
Frequently Asked Questions
Q1. What is demand paging in simple terms?
It is the practice of assigning physical memory to a virtual page only when that page is actually accessed for the first time.
Q2. Does kmalloc() use demand paging?
No. kmalloc() and the buddy/page allocator assign physical memory immediately at call time.
Q3. What triggers a page fault?
A page fault occurs when the MMU cannot find a valid virtual-to-physical mapping for an accessed address.
Q4. What is the difference between a TLB hit and a TLB miss?
A TLB hit means the translation is already cached and access is immediate; a TLB miss means the MMU must walk the page tables.
Q5. Is demand paging the same as swapping?
No. Demand paging concerns first-time allocation of physical pages, while swapping reclaims memory from already-allocated pages under pressure.
Q6. Can I convert a vmalloc virtual address to a physical address manually?
No, this is unsafe and unsupported; only direct-mapped lowmem addresses can be translated this way.
Q7. Why does this matter for embedded and driver developers?
Because it explains why large buffer allocations don’t immediately consume physical RAM, which affects how you size buffers and debug memory usage.
