What Is the Linux Kernel Segment Layout? – Free Linux Device Driver Training

Linux Kernel Segment Layout
Understanding the Kernel Virtual Address Space — Free Linux Kernel Development Course
Free
Linux Device Drivers Course
~30 min
Read Time
x86_64 + ARM
Covered

What You Will Learn

  • How the total virtual address space is split between user space and kernel space
  • What the kernel segment (kernel VAS) contains and why it is shared by all processes
  • The lowmem region, kernel logical addresses, and the PAGE_OFFSET macro
  • The vmalloc/ioremap region and when drivers use it
  • Where kernel modules land in memory when you insmod them
  • How the layout differs between 32-bit and 64-bit kernels
  • A practical kernel module that queries and prints live kernel segment boundaries
Prerequisites
Virtual vs Physical address concepts Process VAS and VMAs (previous lecture) Basic kernel module building C programming — pointers and addresses

The Big Picture: Splitting the Virtual Address Space

Every process on Linux has its own virtual address space. But here is the key insight that surprises many beginners: the kernel shares a portion of that virtual address space with every process simultaneously. The total virtual address space is divided into two halves — one for the user-mode code of that particular process, and one for the kernel that is identical across all processes.

This shared upper portion is what we call the kernel segment or kernel VAS. When a user-space program makes a system call and the CPU switches to kernel mode, the kernel code runs using addresses from this shared kernel segment. No TLB flush is needed when crossing the user-kernel boundary because the kernel mappings were already present in the upper half of the page table.

Full Virtual Address Space Split (x86_64, 48-bit VA)
VA
0xffff ffff ffff ffff
Kernel Segment (~128 TB)
non-canonical hole
User VAS (~128 TB)
0x0000 0000 0000 0000
Kernel VAS (upper half)
Shared across ALL processes.
Contains: lowmem, vmalloc, modules, fixmap, vsyscall, etc.
Non-canonical hole
Hardware limitation — x86_64 only uses 48-bit VA currently. These addresses are invalid.
User VAS (lower half)
Unique per process.
Contains: text, data, heap, stack, shared libs (VMAs).

On 64-bit x86 (x86_64) with a 48-bit virtual address space, each half is approximately 128 TB. With the newer 5-level paging (57-bit VA space, enabled in kernel 5.5+ on supported hardware), each half grows to about 64 PB. On 32-bit x86, the traditional split is 3 GB for user space and 1 GB for the kernel (3:1 split), though other splits like 2:2 are also possible via kernel config.

Kernel Segment Regions Explained

The kernel segment is not a flat, undifferentiated block of memory. It is organized into several distinct regions, each serving a specific purpose. As a Linux device driver developer, knowing these regions is essential because the kernel memory allocation APIs you use — kmalloc(), vmalloc(), ioremap() — return addresses from different regions.

Kernel Segment Regions (x86_64)
Kernel Virtual Address Space (shared by all processes)
Fixmap &
vsyscall
Fixed virtual addresses used for kernel-internal mappings. vsyscall page maps here for backward compatibility.
Kernel
Modules
Where insmod places your LKM’s text and data. On x86_64, this region is typically near the core kernel text to allow short (32-bit relative) jumps between module and kernel code.
vmalloc /
ioremap
Virtually contiguous, physically discontiguous allocations via vmalloc(). Also where ioremap() maps device MMIO registers. Very large region on 64-bit (~32 TB).
Direct-mapped
RAM (lowmem)
All physical RAM mapped here at PAGE_OFFSET. kmalloc() returns addresses from here. Physically contiguous. Contains kernel code (.text, .data, .bss) and all struct page descriptors.
PAGE_OFFSET (0xffff888000000000 on x86_64 with KASLR disabled) ← base of kernel VAS

The Lowmem Region and Kernel Logical Addresses

The lowmem region is the most fundamental part of the kernel segment. Starting at the virtual address given by PAGE_OFFSET, the kernel maps all physical RAM in a linear, direct fashion. This means that for every physical address pa, there is a corresponding kernel virtual address kva = pa + PAGE_OFFSET.

These kernel virtual addresses that have a fixed, predictable relationship to their physical counterparts are called kernel logical addresses. The Linux kernel provides two macros to convert between them:

/* Converting between physical and kernel logical addresses */
#include <linux/mm.h>

/* Physical address → Kernel logical address (kva) */
void *kva = phys_to_virt(physical_addr);
/* Equivalently: kva = (void *)(physical_addr + PAGE_OFFSET) */

/* Kernel logical address → Physical address */
phys_addr_t pa = virt_to_phys(kernel_logical_addr);
/* Equivalently: pa = (phys_addr_t)kernel_logical_addr - PAGE_OFFSET */

/* These ONLY work for lowmem addresses — not vmalloc or ioremap! */
⚠️ Important: virt_to_phys() only works for kernel logical addresses (lowmem). Calling it on a vmalloc address or an ioremap address will give you wrong results. For vmalloc, use vmalloc_to_pfn() or vmalloc_to_page().

What Lives in Lowmem?

  • The kernel’s own compiled binary: .text (code), .rodata (read-only data), .data (writable data), .bss (zero-initialized data)
  • All struct page descriptors — one per physical page frame, used by the page allocator
  • Memory allocated by kmalloc(), kzalloc(), and the slab/SLUB allocator
  • Memory allocated by __get_free_pages() and alloc_pages()
  • DMA-capable memory on systems where all RAM is accessible for DMA

PAGE_OFFSET Value by Architecture

Architecture Typical PAGE_OFFSET User VAS Size Notes
x86_64 (48-bit VA) 0xffff888000000000 ~128 TB Varies with KASLR
x86_64 (57-bit 5-level paging) 0xff11000000000000 ~64 PB Kernel 5.5+ with hardware support
x86 32-bit (3:1 split) 0xc0000000 (3 GB) 3 GB Classic split, still used in embedded
ARM64 (48-bit VA) 0xffff000000000000 ~128 TB TTBR0=user, TTBR1=kernel
RISC-V 64-bit (Sv48) 0xffff800000000000 ~128 TB Growing in embedded Linux systems

The vmalloc and ioremap Region

Above the lowmem region (in virtual address terms) lies the vmalloc region. Memory allocated here is virtually contiguous — meaning it appears as one large block from the kernel’s point of view — but the underlying physical pages may be scattered all over RAM. This makes vmalloc suitable for large kernel allocations where you cannot guarantee physical contiguity.

The vmalloc region doubles as the ioremap space. When a device driver calls ioremap() to map a hardware register window into kernel address space, the kernel virtual address returned comes from this same region.

kmalloc vs vmalloc: Physical Layout Difference
kmalloc / kzalloc
KVA: A
→ PA: page 5
KVA: A+1
→ PA: page 6
KVA: A+2
→ PA: page 7
✓ Physically contiguous
✓ DMA-safe
✓ Low overhead
✗ Limited to ~4MB (typical)
vmalloc
KVA: B
→ PA: page 2
KVA: B+1
→ PA: page 47
KVA: B+2
→ PA: page 103
✓ Virtually contiguous
✓ Can be very large
✗ Not DMA-safe
✗ Higher overhead (TLB)
/* vmalloc usage example in a kernel module */
#include <linux/vmalloc.h>

void *buf;

/* Allocate 10 MB — no way kmalloc handles this */
buf = vmalloc(10 * 1024 * 1024);
if (!buf) {
    pr_err("vmalloc failed\n");
    return -ENOMEM;
}

/* Use the buffer ... */

/* Free when done — always pair with vfree */
vfree(buf);
/* ioremap usage — map device registers into kernel VAS */
#include <linux/io.h>

#define MY_DEVICE_BASE  0xFEA00000UL
#define MY_DEVICE_SIZE  0x1000        /* 4 KB register window */

void __iomem *reg_base;

reg_base = ioremap(MY_DEVICE_BASE, MY_DEVICE_SIZE);
if (!reg_base) {
    pr_err("ioremap failed\n");
    return -ENOMEM;
}

/* Read a 32-bit register at offset 0x10 */
u32 val = readl(reg_base + 0x10);
pr_info("Device register: 0x%08x\n", val);

/* Always unmap when done */
iounmap(reg_base);

The Kernel Modules Region

When you load a kernel module using insmod or modprobe, the kernel allocates memory for that module’s compiled code and data. This memory comes from the kernel modules region, a dedicated area within the kernel segment.

On x86_64, this region is deliberately placed close to the core kernel’s .text section. The reason is a CPU architecture constraint: CALL and JMP instructions on x86_64 can only encode relative offsets up to ±2 GB. If a module called a kernel function and the two were more than 2 GB apart in virtual address space, the relative jump would not reach. Placing modules close to the kernel avoids this problem.

Why Kernel Modules Live Close to Kernel Text (x86_64)
Module text
← ≤ 2 GB → must reach kernel text
relative CALL/JMP range: ±2 GB
Kernel .text
e.g., 0xffffffff81000000
Module region start: typically 0xffffffffa0000000 on x86_64
Distance to kernel text: well within ±2 GB ✓

You can see exactly where a loaded module lives in virtual memory:

# List loaded modules and their kernel virtual addresses
cat /proc/modules

# Or with more detail
sudo cat /sys/module/your_module_name/sections/.text

Practical Kernel Module: Print Kernel Segment Boundaries

Here is a kernel module that queries and prints the actual virtual address boundaries of key kernel segment regions on your running system. This is a great diagnostic tool when learning the free Linux kernel development course material.

/* kernel_seg_info.c — Print kernel segment region addresses
 * Works on Linux kernel 5.x and 6.x
 * Build: make -C /lib/modules/$(uname -r)/build M=$(pwd) modules
 */

#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/mm.h>
#include <linux/vmalloc.h>
#include <linux/init.h>
#include <asm/pgtable.h>

MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("Print kernel segment virtual address boundaries");

static int __init kseg_info_init(void)
{
    pr_info("=== Kernel Segment Layout on This System ===\n");

    /* Lowmem / direct-mapped region */
    pr_info("PAGE_OFFSET (start of kernel VAS / lowmem): 0x%lx\n",
            PAGE_OFFSET);

    /* Kernel text and data boundaries */
    pr_info("_text  (kernel .text start):  0x%lx\n",
            (unsigned long)_text);
    pr_info("_etext (kernel .text end):    0x%lx\n",
            (unsigned long)_etext);
    pr_info("_sdata (kernel .data start):  0x%lx\n",
            (unsigned long)_sdata);
    pr_info("_edata (kernel .data end):    0x%lx\n",
            (unsigned long)_edata);
    pr_info("_end   (kernel image end):    0x%lx\n",
            (unsigned long)_end);

    /* vmalloc region */
    pr_info("VMALLOC_START: 0x%lx\n", VMALLOC_START);
    pr_info("VMALLOC_END:   0x%lx\n", VMALLOC_END);

    /* Module region */
#ifdef MODULES_VADDR
    pr_info("MODULES_VADDR: 0x%lx\n", MODULES_VADDR);
    pr_info("MODULES_END:   0x%lx\n", MODULES_END);
#endif

    /* Address of this module itself */
    pr_info("This module .text: 0x%lx\n",
            (unsigned long)kseg_info_init);

    pr_info("============================================\n");
    return 0;
}

static void __exit kseg_info_exit(void)
{
    pr_info("kseg_info: unloaded\n");
}

module_init(kseg_info_init);
module_exit(kseg_info_exit);
# Build and run
make
sudo insmod kernel_seg_info.ko
dmesg | grep -A 20 "Kernel Segment Layout"
sudo rmmod kernel_seg_info

KASLR: Kernel Address Space Layout Randomization

Modern kernels ship with KASLR (Kernel Address Space Layout Randomization) enabled by default. This security feature randomizes the base virtual address at which the kernel is loaded on every boot. This means the values you see for PAGE_OFFSET, kernel text addresses, and module addresses will be different each time the system starts.

KASLR makes it significantly harder for attackers to exploit vulnerabilities that require knowing the kernel’s virtual address layout (such as return-oriented programming attacks). The randomization range is determined by the kernel config and hardware capabilities.

# Check if KASLR is active
grep -i kaslr /proc/cmdline

# To disable KASLR for debugging (boot parameter)
# Add "nokaslr" to kernel boot parameters in GRUB
# GRUB_CMDLINE_LINUX="nokaslr"

# On a running system — see the actual kernel text load address
sudo cat /proc/kallsyms | grep " _text"
Note for module developers: Your module’s virtual address will also be randomized. Never hardcode kernel virtual addresses in driver code. Always obtain addresses at runtime using kernel APIs like kallsyms_lookup_name() (only available if the kernel is compiled with CONFIG_KALLSYMS) or through proper exported symbols.

Common Mistakes When Working with the Kernel Segment

❌ Mistake 1: Using virt_to_phys() on vmalloc addresses

This is undefined behavior and will give you a garbage physical address. Use vmalloc_to_page() and page_to_phys() for vmalloc regions. Only use virt_to_phys() on addresses from the lowmem region.

❌ Mistake 2: Assuming PAGE_OFFSET is 0xc0000000 on 64-bit

The value 0xc0000000 applies only to 32-bit x86 with a 3:1 split. On 64-bit x86, PAGE_OFFSET is a much larger value and varies with KASLR. Always use the PAGE_OFFSET macro rather than any hardcoded constant.

❌ Mistake 3: kmalloc for large allocations

On x86_64 the practical upper limit for a single kmalloc is 4 MB (determined by the maximum order in the page allocator and slab constraints). Requesting more will fail. Use vmalloc() for large buffers, but remember they are not DMA-safe without extra work.

❌ Mistake 4: Accessing ioremap’d memory without I/O accessors

Never use plain pointer dereferences to access ioremap’d device memory. Always use readl()/writel() (32-bit), readb()/writeb() (8-bit), etc. These ensure proper memory barriers and compiler ordering on all architectures.

Key Takeaways

Kernel segment is shared by ALL processes
Lowmem = all RAM direct-mapped at PAGE_OFFSET
Kernel logical addresses ↔ physical: fixed offset
vmalloc = virtually contiguous, physically scattered
ioremap lives in the vmalloc region
Modules placed near kernel .text for jump range
KASLR randomizes all kernel addresses at boot
Never hardcode kernel virtual addresses

Frequently Asked Questions

Q: Why is the kernel segment shared across all processes?

For performance. When a user process makes a system call, the CPU switches to kernel mode and needs to access kernel code and data immediately. If the kernel had separate virtual mappings per process, the CPU would need to flush the TLB (Translation Lookaside Buffer) on every system call. By keeping the kernel mapped in the upper half of every process’s page table, the kernel mappings are always available without any TLB flush overhead.

Q: Is it a security risk that all processes share the kernel VAS?

This was indeed exploited by the Meltdown hardware vulnerability (2018). The fix was KPTI (Kernel Page Table Isolation), enabled in Linux 4.15+. With KPTI, the kernel is mostly unmapped while user-space code is running, and the mapping is restored only when entering kernel mode. This does cost a small performance overhead (a full TLB flush on kernel entry/exit) but closes the Meltdown attack vector.

Q: What happens if the system has more RAM than the lowmem region can map?

On 32-bit systems with only 1 GB of kernel VAS, this was a real problem — a system with 4 GB of RAM could not map all of it in lowmem. The solution on 32-bit was highmem: a mechanism to temporarily map the extra physical RAM into kernel VAS on demand. On 64-bit systems with terabytes of kernel VAS available, this problem is practically non-existent.

Q: Can I use vmalloc memory for DMA transfers?

Not directly. DMA controllers typically need physically contiguous memory (or scatter-gather lists). You have two options: use kmalloc()/get_free_pages() for small DMA buffers where physical contiguity is guaranteed, or use the DMA API (dma_alloc_coherent()) which handles physical contiguity and cache coherency correctly for DMA regardless of the underlying physical layout.

Q: How do I find out the exact kernel segment layout on my machine?

Several ways: run the kernel module shown in this tutorial, inspect /proc/iomem and /proc/vmallocinfo, or look at /proc/kallsyms for symbol addresses. The crash utility with a vmcore or a live system gives you a complete view. For ARM and ARM64, Documentation/arm64/memory.rst in the kernel source tree has detailed region maps.

Q: What is the fixmap region?

The fixmap is a small area of kernel VAS where virtual addresses are fixed at compile time (hence the name) but the physical pages they point to can be changed at runtime. The kernel uses it for tasks that need a known, compile-time-predictable virtual address — for example, mapping the local APIC registers, early boot memory, and the vsyscall page on x86.

Q: Does ARM64 handle the kernel-user split differently?

Yes. ARM64 hardware has two translation table base registers: TTBR0 for user-space translations and TTBR1 for kernel-space translations. These are switched on every user-kernel transition. This means the ARM64 split is done at the hardware level, not purely by convention. The user VAS starts at address 0 (TTBR0) and the kernel VAS uses the top half (TTBR1). KPTI on ARM64 works similarly but with some architecture-specific optimizations.

Leave a Reply

Your email address will not be published. Required fields are marked *