How Does Linux Manage Virtual Address Space? – Free Linux Device Driver Coure

Linux Virtual Address Space & VM Split Explained
Free Linux Kernel Development Course — Chapter 7 | EmbeddedPathashala
📚 Beginner to Intermediate
⏰ ~20 min read
💻 Kernel 6.x Covered

What You Will Learn

This tutorial is part of our free Linux kernel development course at EmbeddedPathashala. By the end of this article you will understand:

  • What the Virtual Address Space (VAS) is and why every process gets its own
  • How the Linux kernel splits the VAS between user space and kernel space
  • Why 64-bit systems use only 48 or 57 bits of the full 64-bit address range
  • What canonical and non-canonical addresses mean on x86-64
  • How the kernel VAS is shared by all processes while user VAS is private
  • How to read /proc/<pid>/maps to inspect the VAS of any process
  • Differences in VM split across ARM, x86-64, and AArch64 architectures

Prerequisites

Before reading this article you should be comfortable with:

  • Basic C programming (pointers, memory allocation)
  • What a process is and how the OS schedules tasks
  • Binary and hexadecimal number representation
  • General idea of RAM vs virtual memory (paging is explained here from scratch)

No prior kernel programming experience is required.

What Is a Virtual Address Space?

When your C program calls malloc() or accesses an array, it uses virtual addresses, not physical RAM addresses. The CPU’s Memory Management Unit (MMU) transparently translates each virtual address to a physical address at run time using a data structure called a page table.

This translation layer gives every process the illusion that it owns the full address range of the CPU. On a 32-bit system that range is 4 GB; on a 64-bit system it is theoretically 16 EB (exabytes). This full range is the Virtual Address Space (VAS) of a process.

The VAS is almost entirely sparse — most addresses are unmapped. Only the regions that are explicitly mapped (text segment, data segment, heap, stack, library mappings, kernel segment) consume actual page-table entries and, when accessed, consume physical RAM.

Figure 1 — Process Virtual Address Space Overview (32-bit IA-32, 3 GB:1 GB split)
0xFFFFFFFF
KERNEL SEGMENT
Shared by ALL processes
Code, data, stacks, page tables
1 GB (3:1 split)
0xC0000000  (PAGE_OFFSET)
Stack & mmap region
[rw-] grows downward
↑ grows up     ↓ grows down
Heap
[rw-] grows upward
BSS (Uninitialised data)
[rw-]
Data segment (Initialised)
[rw-]
Text segment (Code)
[r-x] read + execute only
0x00000000

The Linux Kernel VM Split: User Space vs Kernel Space

Every process VAS on Linux is divided into two halves: user space and kernel space. This is called the VM split (or user:kernel split). The boundary is set at compile time by the kernel configuration parameter PAGE_OFFSET.

The kernel space half is identical in every process VAS. When a process makes a system call or an interrupt fires, the CPU switches from user mode to kernel mode and starts executing kernel code that lives in the upper portion of the same virtual address space. This design avoids a full context switch just to access kernel memory — the kernel segment is always mapped in.

Figure 2 — VM Split Ratios for Common Architectures (4 KB Page Size)
Architecture Addr Bits VM Split (User:Kernel) User End Address Kernel Start Address
IA-32 32 3 GB : 1 GB 0xBFFFFFFF 0xC0000000
ARM (32-bit) 32 2 GB : 2 GB 0x7FFFFFFF 0x80000000
x86-64 (4-level paging) 48 128 TB : 128 TB 0x0000 7FFF FFFF FFFF 0xFFFF 8000 0000 0000
x86-64 (5-level paging) 57 64 PB : 64 PB 0x00FF FFFF FFFF FFFF 0xFF00 0000 0000 0000
AArch64 (39-bit) 39 512 GB : 512 GB 0x0000 007F FFFF FFFF 0xFFFF FF80 0000 0000
AArch64 (48-bit) 48 256 TB : 256 TB 0x0000 FFFF FFFF FFFF 0xFFFF 0000 0000 0000

The highlighted row (128 TB:128 TB) is the most common on modern Linux x86-64 desktops, laptops, and servers.

⚠ Key Insight: No 64-bit processor in production today actually uses all 64 bits for addressing. The hardware truncates the address bus — x86-64 currently uses 48 bits (4-level paging), giving 256 TB of total VAS. AMD and Intel are extending this to 57 bits with 5-level paging for workloads that need more than 128 TB of user memory.

Canonical vs Non-Canonical Addresses on x86-64

With 48-bit addressing on x86-64, only bits 0–47 carry address information. The hardware requires that bits 48–63 be a sign extension of bit 47. This means they must all be 0 (if bit 47 is 0) or all be 1 (if bit 47 is 1). An address satisfying this rule is called a canonical address.

Addresses where bits 48–63 are neither all-0 nor all-1 are non-canonical. The CPU raises a general-protection fault if you try to access a non-canonical address. This is exactly the large unused “hole” in the middle of the 64-bit VAS — it is not wasted; it is a hardware enforcement zone.

Figure 3 — x86-64 Canonical Address Layout (48-bit, 4-level paging)
0xFFFF FFFF FFFF FFFF   0xFFFF 8000 0000 0000     0x0000 7FFF FFFF FFFF   0x0000 0000 0000 0000
Canonical Higher Half
Kernel Segment
128 TB
Non-Canonical
(Unused hole)
~16 EB minus 256 TB
Canonical Lower Half
User VAS
128 TB

The two canonical halves together give each process a total of 256 TB of usable virtual address space — 128 TB for user programs and 128 TB for the kernel. The enormous non-canonical region in between is why we describe the VAS as sparse: almost no addresses in the full 64-bit range are actually valid.

What Lives in the Kernel Segment?

The kernel segment (canonical higher half) holds everything the Linux kernel needs:

  • Kernel text & data: The compiled kernel binary — vmlinux — its read-only code and read-write global data.
  • Per-process kernel stacks: Every process has a small kernel-mode stack (typically 16 KB on 64-bit systems) used during system calls and interrupts.
  • Page tables: The multi-level page tables that translate virtual to physical addresses for all processes.
  • Task structures (struct task_struct): The kernel’s process descriptor that tracks state, scheduling, file descriptors, memory maps, and more.
  • Kernel modules: Loadable modules (.ko files) are mapped here when inserted with insmod or modprobe.
  • Device driver memory: DMA buffers, ioremap regions, and driver-specific allocations.
  • vmalloc / vmap region: Virtually-contiguous but physically-discontiguous memory used by the kernel for large allocations.
  • Direct physical memory map: On 64-bit kernels, all physical RAM is directly mapped into the kernel segment. This is how the kernel accesses any physical page without needing to map it first.
💡 Performance Note: Kernel memory is never swapped out. The kernel cannot page-fault its own code or critical data structures. User-space pages, on the other hand, can be paged to swap space. Kernel code must handle this asymmetry carefully when copying data to/from user space.

The Full Process VAS: Every Process Is Unique in User Space, Identical in Kernel Space

This is one of the most important concepts in Linux kernel internals: all processes running on the system share the exact same kernel segment, but each process has its own private user VAS.

Figure 4 — Multiple Processes Share the Kernel Segment
Process P1
Kernel Segment
(shared)
Stack [rw-]
lib mappings
Heap
BSS / Data
Text [r-x]
Process P2
Kernel Segment
(shared)
Stack [rw-]
lib mappings
Heap
BSS / Data
Text [r-x]
Process Pn
Kernel Segment
(shared)
Stack [rw-]
lib mappings
Heap
BSS / Data
Text [r-x]

Each process has a unique user VAS but all map to the same physical kernel segment pages.

Inspecting the VAS of a Running Process

On any Linux system you can view the actual VAS segments of a process through the procfs filesystem. The file /proc/<pid>/maps lists every mapped memory region with its address range, permissions, and backing file.

# Show the VAS of the current shell
cat /proc/self/maps

# Show the VAS of a specific process (replace 1234 with actual PID)
cat /proc/1234/maps

# More readable output using pmap
pmap -x 1234

A typical output line looks like:

55a3b2c00000-55a3b2c02000 r--p 00000000 08:01 1245678  /usr/bin/bash
55a3b2c02000-55a3b2caf000 r-xp 00002000 08:01 1245678  /usr/bin/bash
55a3b2caf000-55a3b2ce4000 r--p 000af000 08:01 1245678  /usr/bin/bash
55a3b2ce4000-55a3b2ce8000 r--p 000e3000 08:01 1245678  /usr/bin/bash
55a3b2ce8000-55a3b2ceb000 rw-p 000e7000 08:01 1245678  /usr/bin/bash
55a3b3a5a000-55a3b3b0d000 rw-p 00000000 00:00 0        [heap]
7f8b1f000000-7f8b1f200000 rw-p 00000000 00:00 0
7ffceb1a4000-7ffceb1c5000 rw-p 00000000 00:00 0        [stack]
7ffceb1e1000-7ffceb1e5000 r--p 00000000 00:00 0        [vvar]
7ffceb1e5000-7ffceb1e7000 r-xp 00000000 00:00 0        [vdso]
ffffffffff600000-ffffffffff601000 --xp 00000000 00:00 0 [vsyscall]

The permission string has the format rwxp: r=read, w=write, x=execute, and p=private (copy-on-write) or s=shared. Notice that the text segment is r-xp (read + execute, no write), the heap is rw-p, and the stack is also rw-p.

Addresses starting with 7f... are in the upper 128 TB canonical lower half (user space) on x86-64. Addresses starting with ffff... belong to the kernel segment and appear only in kernel-mode context — they are not visible in /proc/maps for security reasons.

Changes in Recent Kernels (Linux 5.x / 6.x)

The fundamental VM split concept has been stable for many years, but modern kernels have introduced important refinements:

Feature Kernel Version Impact
5-level paging (LA57) 4.14+ Extends user VAS to 128 PB on supported Intel/AMD CPUs
KASLR (Kernel ASLR) 3.14+ Randomizes the kernel load address at boot within the kernel segment
KPTI (Page Table Isolation) 4.15+ (Meltdown fix) Separates kernel and user page tables to mitigate Meltdown; kernel mappings in user-mode PGD are minimal
Memory Tagging (MTE) 5.10+ (AArch64) Top-byte tags in pointers; affects the effective canonical address range on ARM
GUP (Get User Pages) hardening 5.6+ Stricter rules on kernel pinning of user pages; affects DMA drivers

KPTI: When the Kernel Segment Is No Longer Fully Mapped in User Mode

Before the Meltdown vulnerability (CVE-2017-5754), the full kernel segment was always mapped in every process’s page table even during user-mode execution — it was just protected by the supervisor/user permission bit. Meltdown showed that speculative execution could leak kernel data across this boundary.

KPTI (Kernel Page-Table Isolation) introduced a two-page-table model: a minimal user-mode page table (containing only the kernel trampoline needed to handle system calls) and the full kernel page table. On every context switch between user and kernel mode, the CPU switches between these two page tables. This adds a small performance overhead but closes the Meltdown side-channel.

From a conceptual standpoint, the VM split still exists; the implementation is just more security-conscious now.

Common Mistakes and Misconceptions

  • Confusing virtual with physical addresses: The addresses in /proc/maps are all virtual. Two processes can have the same virtual address but they map to completely different physical pages.
  • Thinking 64-bit means 16 EB of usable RAM: The VAS is 16 EB on a 64-bit system, but the CPU only implements 48 or 57 address bits. Physical RAM support is much smaller (typically up to a few TB on server platforms).
  • Assuming the stack grows up: On x86 and ARM, the user-mode stack grows downward (toward lower addresses). The heap grows upward. This is why they can coexist in the same VAS without a fixed boundary.
  • Thinking kernel memory can be swapped: It cannot. Kernel memory is always resident. Only user-space pages participate in page-out to swap.
  • Confusing PAGE_OFFSET with a physical address: PAGE_OFFSET is the virtual address where the kernel segment starts. On 32-bit IA-32 with 3:1 split it is 0xC0000000. It is not a physical address.

Best Practices for Kernel and Driver Developers

  • Always use copy_to_user() / copy_from_user() when exchanging data between kernel and user space — never dereference a raw user pointer in kernel mode.
  • Use access_ok() before accessing user-space pointers to verify the address falls within the user VAS.
  • Prefer kmalloc() for small kernel allocations and vmalloc() for large ones that do not need physical contiguity.
  • When writing device drivers that use DMA, allocate coherent memory with dma_alloc_coherent() rather than kmalloc() to avoid cache-coherency issues.
  • Use IS_ERR() and PTR_ERR() to handle error pointers returned by kernel APIs — a value like -ENOMEM encoded in a pointer looks like a valid kernel-space address but is not.

📋 Key Takeaways

VAS = virtual illusion of full address range
VM split divides user & kernel space
x86-64 uses only 48 bits (128 TB each)
Kernel segment is shared by all processes
Non-canonical addresses = hardware hole
KPTI splits page tables for security
/proc/<pid>/maps shows live VAS
Kernel memory is never swappable

Frequently Asked Questions

Q1: Why does a 64-bit system not use all 64 bits for addressing?

Implementing all 64 address lines in silicon would require routing hundreds of additional traces on the CPU die and the motherboard. Since no workload today needs more than 128 TB of virtual address space per process, hardware vendors limit the physical implementation to 48 bits (and recently 57 bits with 5-level paging). Unused bits are checked for canonical form to catch programming errors early.

Q2: What happens if a program accesses a non-canonical address?

The CPU raises a #GP (General Protection) fault immediately. The kernel converts this into a SIGSEGV signal for the offending process, which normally terminates it. Non-canonical accesses are a common symptom of corrupted function pointers or wild pointer arithmetic.

Q3: Does each thread in a multi-threaded process get its own VAS?

No. All threads within the same process share the same VAS (and thus the same page table). Each thread has its own stack region within that shared VAS, but they can freely read and write each other’s memory — which is why synchronization is needed. This is one of the key differences between threads and processes.

Q4: What is PAGE_OFFSET and where is it defined in the kernel source?

PAGE_OFFSET is a compile-time macro that defines the virtual address where the kernel segment begins. On x86-64 it is defined in arch/x86/include/asm/page_64_types.h. For 64-bit kernels it is typically 0xffff888000000000 (the start of the direct physical memory map), though the exact value depends on KASLR and the kernel configuration.

Q5: How does KPTI affect performance, and how can I check if it is enabled?

KPTI adds a TLB flush on every user-to-kernel and kernel-to-user transition because the two page tables use different CR3 values. On workloads that make frequent system calls (e.g., a web server doing many small I/O operations), KPTI overhead can be 5–30%. Newer CPUs use PCID (Process-Context Identifiers) to tag TLB entries and avoid full flushes, reducing the overhead significantly. Check your kernel boot parameters with cat /proc/cmdline and look for pti=on or nopti. The file /sys/devices/system/cpu/vulnerabilities/meltdown shows the current mitigation status.

Q6: Can user-space code ever see kernel virtual addresses?

No. The kernel segment pages are marked with the supervisor-only bit in the page table. Even if a user-space program speculates about a kernel address, the CPU will not allow it to be read in user mode (and KPTI removes most of the kernel mappings from the user-mode page table entirely). The /proc/kallsyms file shows kernel symbol addresses but only to root, and the values are randomised when KASLR is active.

Q7: What is the vmalloc region and how is it different from the direct map?

The direct map (also called the linear map) in the kernel segment maps every byte of physical RAM at a fixed, predictable virtual address. This lets the kernel access any physical page instantly without needing to set up a special mapping. The vmalloc region is a separate area used for large allocations where physical contiguity is not required. Pages in the vmalloc region may be physically scattered; the kernel sets up page-table entries to make them appear virtually contiguous. kmalloc() returns physically contiguous memory (from the direct map); vmalloc() returns only virtually contiguous memory.

Q8: How does the 5-level paging differ from 4-level paging?

Standard 4-level paging uses a 4-level page table hierarchy: PGD → PUD → PMD → PTE, supporting 48-bit virtual addresses (256 TB total). 5-level paging (also called LA57) adds one more level — P4D — making the hierarchy PGD → P4D → PUD → PMD → PTE, supporting 57-bit virtual addresses (128 PB per user/kernel half). Linux supports 5-level paging from kernel 4.14 onward when the CPU advertises the LA57 CPUID bit. Most current consumer hardware still uses 4-level paging.

Conclusion

Understanding the Linux Virtual Address Space and VM split is foundational to everything else in kernel development and low-level systems programming. Whether you are writing a device driver, debugging a kernel oops, or working on memory-sensitive embedded firmware, you constantly reason about which half of the address space you are in, what permissions apply, and how the hardware enforces those boundaries.

In the next lecture in this free Linux kernel development course, we will dig into page tables in detail — how the kernel walks the four (or five) levels of the page table to translate a virtual address, how huge pages work, and how the kernel’s own memory allocators (slab, slub, buddy system) carve up the kernel segment.

Keep experimenting with /proc/<pid>/maps and pmap on your own Linux machine — the VAS layout will quickly become second nature.

Continue Your Free Linux Kernel Development Journey

EmbeddedPathashala offers completely free, in-depth courses on Linux kernel programming, Linux device drivers, and embedded systems development.

Visit EmbeddedPathashala

Leave a Reply

Your email address will not be published. Required fields are marked *