What Is a VMA in the Linux Kernel? – Free Linux Device Driver Course

 

← Previous Lecture Linux Kernel Programming – EmbeddedPathashala Next Lecture →
Virtual Memory Areas (VMAs) in the Linux Kernel
Free Linux Kernel Development Course | EmbeddedPathashala
Level
Beginner–Intermediate
Part
2 of 2
Topic
VMAs & Memory Regions

In Part 1 of this free Linux kernel development course tutorial, we learned how the Process Virtual Address Space (VAS) is organized into user space and kernel space. We saw the major segments — text, data, BSS, heap, stack — and learned how to view them through /proc/PID/maps. Now, in Part 2, we go one layer deeper: we examine the kernel data structure that actually implements each of those segments. That structure is called the Virtual Memory Area, or VMA.

Understanding VMAs is critical for anyone pursuing free Linux kernel programming or writing Linux device drivers, because VMAs are the structures you encounter directly when implementing mmap() in a driver, handling page faults, or doing memory-mapped I/O from the kernel side.

What You Will Learn

  • What a Virtual Memory Area (VMA) is and why the kernel needs it
  • How the kernel stores and manages VMAs for a process
  • The important fields inside the vm_area_struct kernel structure
  • How VMAs relate to what you see in /proc/PID/maps
  • How VMAs are used in Linux device drivers (the mmap file operation)
  • VMA flags and protection bits explained clearly
  • How page faults and demand paging connect to VMAs
  • Best practices for working with VMAs in kernel development
Prerequisites: Complete Part 1 of this tutorial (Process Virtual Address Space) before reading this one. Basic knowledge of C structures and pointers is required.
Topics Covered
VMA vm_area_struct Linux Kernel Memory mmap driver Page Fault Handling VMA flags mm_struct Demand Paging

What Is a Virtual Memory Area (VMA)?

Every line you see in /proc/PID/maps represents one Virtual Memory Area (VMA). A VMA is the kernel’s internal representation of a contiguous range of virtual addresses that share the same properties — the same permissions, the same backing store, and the same behavior when accessed.

Think of the user space VAS as a large, mostly empty number line. A VMA is a labeled, colored segment painted onto that number line, saying “this range of addresses belongs to the program’s text segment, is read-only and executable, and is backed by the program’s ELF binary on disk.” The kernel creates a VMA for each distinct segment in the process.

In this free Linux kernel development course, we will look at how these VMAs are actually stored in the kernel and what information each one carries.

Process User Space VAS — VMAs in the /proc/PID/maps View
Address RangeVMA DescriptionPermissions
High UVAstack [stack]rw-p
…sparse (no VMA)—
…[vvar] — kernel-shared varsr–p
…[vdso] — virtual DSOr-xp
…libc.so (mmap region)r-xp
…libc.so (data)rw-p
…[heap]rw-p
…/usr/bin/myapp (data+BSS)rw-p
Low UVA/usr/bin/myapp (text)r-xp

The vm_area_struct: The Kernel’s VMA Data Structure

In the Linux kernel source, each VMA is represented by a C structure called vm_area_struct, defined in linux/mm_types.h. Every VMA in a process has one instance of this structure allocated in kernel memory.

The key fields you need to understand for this free Linux kernel programming course are:

struct vm_area_struct { /* The VMA covers virtual addresses from vm_start to vm_end */ unsigned long vm_start; /* start address (inclusive) */ unsigned long vm_end; /* end address (exclusive) */ /* Link to other VMAs of this process */ struct vm_area_struct *vm_next; /* linked list – older kernels */ struct vm_area_struct *vm_prev; /* Permission and type flags */ pgprot_t vm_page_prot; /* page-level HW protection */ unsigned long vm_flags; /* VM_READ, VM_WRITE, VM_EXEC, … */ /* Pointer back to this process’s memory descriptor */ struct mm_struct *vm_mm; /* Operations — function pointers for fault handling, etc. */ const struct vm_operations_struct *vm_ops; /* If file-backed: the mapped file and offset */ struct file *vm_file; /* NULL if anonymous */ unsigned long vm_pgoff; /* offset in pages */ };
Kernel Version Note: In Linux kernel 6.1 and later, the VMA list management was changed from a simple doubly-linked list to a maple tree data structure for better scalability with processes that have very large numbers of VMAs. The core fields like vm_start, vm_end, vm_flags, and vm_ops remain conceptually the same. The maple tree makes lookup, insertion, and deletion of VMAs significantly faster for processes with thousands of mappings.

vm_start and vm_end

These two fields define the exact virtual address range this VMA covers. The range is half-open: vm_start is the first byte address that belongs to this VMA, and vm_end is the first byte address that does not belong to it. Both values are always page-aligned (multiples of the system page size, typically 4096 bytes).

You can compute the size of a VMA as:

size_in_bytes = vma->vm_end - vma->vm_start;

vm_flags: VMA Permission and Behavior Flags

The vm_flags field is a bitmask that describes both the permissions and the behavior of the VMA. The most important flags for this free Linux kernel development course are:

Flag Meaning
VM_READ Pages in this VMA may be read
VM_WRITE Pages in this VMA may be written
VM_EXEC Pages in this VMA may be executed (code)
VM_SHARED Pages are shared between processes (not private COW)
VM_MAYREAD / VM_MAYWRITE / VM_MAYEXEC Maximum allowed permissions (used during mprotect calls)
VM_GROWSDOWN VMA can grow downward (stack)
VM_GROWSUP VMA can grow upward (some architectures)
VM_IO This VMA maps device I/O memory — important for driver mmap
VM_PFNMAP VMA maps raw PFNs (page frame numbers), not normal pages — used in device drivers
VM_DONTEXPAND Cannot be expanded with mremap
VM_LOCKED Pages are locked in RAM (not swappable)

vm_ops: VMA Operations

The vm_ops field points to a vm_operations_struct, which is a table of function pointers. These functions are callbacks that the kernel calls when specific events happen on this VMA. The most important callback for Linux device drivers is:

struct vm_operations_struct { /* Called when the VMA is opened (e.g. forked or duped) */ void (*open)(struct vm_area_struct *vma); /* Called when the VMA is closed (process exits or munmap) */ void (*close)(struct vm_area_struct *vma); /* Called when a page fault occurs in this VMA */ vm_fault_t (*fault)(struct vm_fault *vmf); };

The fault() callback is at the heart of demand paging: the kernel does not load every page of a program into RAM when it starts. Instead, it creates the VMA but leaves the page table entries empty. The first time a program touches an address in that VMA, the CPU raises a page fault. The kernel’s page fault handler looks up which VMA the faulting address belongs to, then calls that VMA’s fault() function to load the required page into RAM.

mm_struct: The Process Memory Descriptor

All the VMAs for a single process are managed through a structure called mm_struct. Every process has exactly one mm_struct (or shares one, in the case of kernel threads). This is the top-level memory descriptor for the process.

Relationship: task_struct → mm_struct → vm_area_struct chain
task_struct
(Process descriptor)
→ mm
→
mm_struct
(Memory descriptor)
mmap (VMA list root)
mm_count, map_count
pgd (page-global-dir)
→
vm_area_struct #1
text segment
r-xp
↓ vm_next
vm_area_struct #2
data segment
rw-p
↓ vm_next
vm_area_struct #N
heap / stack / libs
…

The kernel finds the mm_struct for the current process via current->mm (where current is a pointer to the current process’s task_struct). Inside mm_struct, the field mmap points to the first VMA in a sorted linked list. In kernels from 6.1 onwards, a maple tree structure is used for faster lookup, but the same concept applies.

VMAs and Demand Paging: How Pages Are Loaded On Demand

One of the most important jobs of the VMA system is enabling demand paging. When the kernel loads a program, it does not immediately copy the entire executable binary into RAM. Instead, it:

  1. Creates VMAs describing each segment (text, data, etc.)
  2. Leaves the actual page table entries empty (not present)
  3. Waits until the CPU tries to access a specific address
  4. Receives the resulting page fault interrupt
  5. Looks up which VMA contains the faulting address
  6. Calls that VMA’s fault() handler to load the needed page from disk into RAM
  7. Updates the page table entry and resumes the CPU from the faulting instruction
Demand Paging via Page Fault — Step by Step
① CPU: Access virtual address 0x5555557a0000
↓
② MMU: Page table entry NOT PRESENT → raise Page Fault
↓
③ Kernel page fault handler: find the VMA for this address
↓
④ Call vma->vm_ops->fault() → load page from disk into RAM
↓
⑤ Update page table entry → mark page PRESENT
↓
⑥ CPU resumes from faulting instruction — transparent to user

From the user program’s perspective, none of this happens. The program simply accesses memory and everything works. The entire page fault, disk read, and page table update happen invisibly inside the kernel. This is one of the most elegant aspects of the virtual memory system.

VMAs in Linux Device Drivers: The mmap File Operation

For anyone learning this free Linux kernel programming course with an interest in writing device drivers, VMAs are directly encountered when implementing the mmap() file operation in a character device driver.

When a user program calls mmap() on your device file, the kernel:

  1. Creates a new VMA for the requested address range
  2. Calls your driver’s .mmap file operation callback
  3. Passes a pointer to the newly created VMA to your callback
  4. Expects your driver to set up the page table entries (or set vm_ops) so the user can access the device memory

A Basic Driver mmap Implementation

Here is a minimal example of how a character device driver implements mmap() to map a region of physical device memory (for example, an MMIO buffer) into user space:

static int mydriver_mmap(struct file *filp, struct vm_area_struct *vma)
{
    unsigned long pfn;
    unsigned long size;

    size = vma->vm_end - vma->vm_start;

    /* Physical address of device memory, converted to page frame number */
    pfn = DEVICE_PHYS_ADDR >> PAGE_SHIFT;

    /*
     * VM_IO: marks this VMA as device I/O memory.
     * VM_DONTEXPAND: prevents mremap from expanding this VMA.
     * These flags are mandatory for device memory mappings.
     */
    vma->vm_flags |= VM_IO | VM_DONTEXPAND | VM_DONTDUMP;

    /* Set non-cacheable page protection for device registers */
    vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);

    /*
     * remap_pfn_range: maps physical pages (by PFN) into the VMA.
     * Returns 0 on success, negative on failure.
     */
    if (remap_pfn_range(vma,
                        vma->vm_start,
                        pfn,
                        size,
                        vma->vm_page_prot)) {
        return -EAGAIN;
    }

    return 0;
}
Driver Developer Warning: Always set VM_IO and VM_DONTEXPAND when mapping device memory. Failing to set VM_IO can cause the kernel to attempt to swap out device memory, which will crash the system. Also set VM_DONTDUMP to prevent device memory from appearing in core dumps, which would be meaningless and could expose sensitive data.

How Many VMAs Does a Typical Process Have?

The number of VMAs in a process depends on how many distinct memory regions have been mapped. A minimal statically-linked program might have only 3–5 VMAs. A typical dynamically-linked application that uses several shared libraries might have 50–100 VMAs. A complex application like a web browser or JVM can easily have several hundred to a few thousand VMAs.

You can check the VMA count for any process by looking at the VmPeak / VmSize and VmMapped fields in /proc/PID/status, or by counting lines in /proc/PID/maps:

wc -l /proc/$(pgrep firefox)/maps
Kernel Limit: The Linux kernel imposes a default limit on the number of VMAs a process can have. This limit is controlled by /proc/sys/vm/max_map_count and defaults to 65530 on most distributions. Applications that need more (like certain JVMs or security-hardened builds) may need this limit increased.

Best Practices When Working With VMAs in Kernel Development

✅ Always check VMA boundaries before walking the VMA list

When writing kernel code that iterates over VMAs, always acquire the appropriate lock (mmap_read_lock(mm) for read access) before walking the VMA list. Failing to do so is a race condition that can cause kernel crashes or data corruption.

✅ Use find_vma() to locate a VMA by address

The kernel provides the helper function find_vma(mm, addr) to efficiently find the VMA that contains or starts after a given virtual address. Use this instead of walking the list manually.

✅ Set appropriate vm_flags in your driver’s mmap callback

Always set VM_IO for device I/O memory mappings. Set VM_PFNMAP when using raw PFN mappings. These flags change how the kernel’s memory management treats the VMA and are critical for correct driver behavior.

✅ Set vm_ops if you need page fault handling

If your driver needs to map memory lazily (one page at a time, on demand), implement the fault() callback in a vm_operations_struct and assign it to vma->vm_ops in your mmap() implementation.

Key Takeaways

VMA = one map lineEach line in /proc/PID/maps corresponds to one vm_area_struct in the kernel.
vm_flags matterVM_READ, VM_WRITE, VM_EXEC, VM_IO, VM_SHARED control how the kernel treats each VMA.
vm_ops callbacksThe fault() callback enables demand paging — loading pages only when first accessed.
mm_struct is the rootAll VMAs belong to a process’s mm_struct, accessed via current->mm in kernel code.
Driver mmapWhen writing a device driver, your mmap callback receives the VMA and must map physical device pages into it.
Maple tree (6.1+)Modern kernels use a maple tree for VMA storage, replacing the old sorted linked list for better performance.

Frequently Asked Questions (FAQ)

Q1. What is a VMA (Virtual Memory Area) in the Linux kernel?
A VMA is the kernel’s internal data structure (vm_area_struct) that represents a contiguous range of virtual addresses in a process’s user space with uniform properties — same permissions, same backing store, same vm_ops callbacks. Every segment you see in /proc/PID/maps is one VMA.
Q2. What is the difference between vm_flags and vm_page_prot?
vm_flags is a kernel-level bitmask describing the logical properties of the VMA (VM_READ, VM_WRITE, VM_EXEC, VM_IO, etc.). vm_page_prot is the hardware-level page protection value that gets written into the actual page table entries for pages in this VMA. vm_page_prot is derived from vm_flags during page mapping.
Q3. What happens when a page fault occurs inside a VMA?
The CPU raises a page fault exception. The kernel’s page fault handler identifies the VMA containing the faulting address. If a vm_ops->fault() callback is set, it calls that function to load the needed page into RAM and update the page table entry. The faulting CPU instruction is then transparently re-executed.
Q4. What does it mean if a page fault occurs outside any VMA?
If the faulting address is not covered by any VMA (it falls in a sparse region), the kernel determines this is an invalid memory access. For a user process, it sends a SIGSEGV (segmentation fault) signal, which normally terminates the program. In kernel mode, this would be a kernel oops.
Q5. Why does a device driver need to set VM_IO on a VMA?
VM_IO tells the kernel’s memory management system that this VMA maps device I/O memory, not normal RAM. This prevents the kernel from trying to include these pages in core dumps, from attempting to swap them out, and from doing other page-management operations that would be incorrect or harmful for device register memory.
Q6. What is the maple tree used for in Linux kernel 6.1+?
The maple tree replaces the red-black tree and doubly-linked list previously used to store a process’s VMAs. It provides better performance for processes with large numbers of VMAs, with faster insertion, deletion, and range-lookup operations. The behavior from a driver or module author’s perspective is unchanged — you still use helper functions like find_vma() and the VMA fields remain the same.
Q7. How does fork() handle VMAs?
When a process calls fork(), the child process gets a complete copy of the parent’s VMA list. However, the underlying pages are not immediately duplicated. Instead, both parent and child are given read-only mappings to the same physical pages (Copy-On-Write, or COW). When either process tries to write to one of these shared pages, the kernel creates a private copy for the writing process at that point.
Q8. Can two VMAs in the same process overlap?
No. The kernel enforces that all VMAs within a single process are non-overlapping and sorted by starting address. When you call mmap() and request a specific address that overlaps an existing VMA, the kernel either unmaps the overlapping part of the old VMA or returns an error, depending on the flags you passed.

Conclusion

Virtual Memory Areas are the kernel’s way of giving precise meaning to every byte in a process’s user space. They are not just abstract data structures — they directly control what happens when your program accesses memory, what the CPU is allowed to do with each page, and how device drivers expose hardware registers to user programs.

In this free Linux kernel development course tutorial, we walked through the vm_area_struct in detail, explained the mm_struct that ties all VMAs together, showed how demand paging works through the fault callback, and demonstrated how VMAs appear in the context of a character device driver’s mmap() implementation.

Together, Parts 1 and 2 of this tutorial give you a solid foundation in Linux process memory management. These concepts will reappear constantly as you go deeper into the kernel — in memory allocation internals, page table management, slab allocators, and more advanced Linux device driver topics. All of that is covered in the rest of this free Linux kernel programming course on EmbeddedPathashala.

Continue Your Free Linux Kernel Development Journey

EmbeddedPathashala — Free Linux Kernel, Linux Device Drivers, Embedded Systems, and BLE courses

← Previous Lecture EmbeddedPathashala | Free Linux Kernel Programming Next Lecture →

Leave a Reply

Your email address will not be published. Required fields are marked *