Linux Process Virtual Address Space (VAS): Segments & Memory Layout – Best Linux Device Drivers Course

Free Linux Kernel Development Course — EmbeddedPathashala

Linux Process Virtual Address Space (VAS): Segments & Memory Layout Explained

Understand how the Linux kernel organises every process in memory — from text and data segments to heap, stack and shared library mappings — with up-to-date Linux 6.x internals.

~15 min
Read Time
Free
No Paywall
Linux 6.x
Up to Date

What You Will Learn

  • What the Process Virtual Address Space (VAS) is and why Linux sandboxes every process inside it
  • The exact memory segments / mappings every Linux process contains (text, data, heap, stack, libraries)
  • How the heap grows dynamically and what MMAP_THRESHOLD means in Linux 6.x
  • How the stack grows downward and why the “top” of the stack is actually the lowest address
  • How shared libraries are mapped into a process’s VAS at runtime
  • Why every thread needs its own private stack inside the shared VAS
  • How to inspect a live process VAS using /proc/<pid>/maps on your Linux system

Prerequisites

Before reading this lecture you should be comfortable with:

  • Basic C programming (pointers, variables, functions)
  • What a process is and how the OS loads a program
  • Elementary understanding of memory addresses (virtual vs physical is a plus)

No kernel-level experience is required. This lecture is beginner-friendly and part of the free Linux kernel development course at EmbeddedPathashala.

What is the Process VAS in Linux?

When a program runs on Linux, the operating system does not give it access to all of physical RAM. Instead, the kernel creates a virtual address space (VAS) — an isolated, private view of memory that belongs exclusively to that process. Think of it as a walled garden: the process can only access addresses that exist inside its own VAS. Any attempt to read or write outside this boundary immediately triggers a segmentation fault and the kernel kills the offending process.

This isolation is a cornerstone of modern OS security and stability. One process cannot corrupt another process’s memory, even if they are running the same program simultaneously. The Linux kernel enforces this sandbox using the CPU’s Memory Management Unit (MMU), which translates virtual addresses to physical RAM locations transparently.

In the Linux kernel development world, you will frequently hear “process VAS”, “process image”, and “address space” used interchangeably. They all mean the same thing: the range of virtual memory addresses that a process is allowed to use.

Figure 1 — Linux Process Virtual Address Space Layout (x86-64)
HIGH VIRTUAL ADDRESS
▼ Stack [rw-]  grows downward
   thread stacks (thrd2, thrd3 …)
📚 Library mappings [r-x / rw-]
📚 Library mappings [r-x / rw-]
▲ Heap [rw-]  grows upward
malloc() < 128KB → from heap
BSS — Uninitialized data [rw-]
auto-zeroed at runtime
Initialized data [rw-]
Text / Code [r-x]  read-execute only
LOW VIRTUAL ADDRESS (0x0000…)

The Memory Segments Inside a Linux Process VAS

The VAS is not one big block of memory. The Linux kernel divides it into distinct regions called segments (the older UNIX term) or, more precisely, mappings (the modern kernel term, since they are managed via the mmap subsystem). Each mapping has its own set of permissions and purpose. Let’s go through each one.

1. Text Segment (Code Segment) — r-x

The text segment holds the compiled machine instructions of your program — what the CPU actually executes. It is marked read and execute only (r-x). Writing to this region is forbidden, which prevents accidental or malicious modification of running code.

This segment is static — its size is fixed at program load time and does not change during execution. On Linux 6.x, the kernel uses W^X (write XOR execute) policy: a memory page can be either writable or executable, never both simultaneously. This is enforced at the hardware level via the CPU’s NX (No-Execute) bit and is a crucial Linux kernel security feature.

When multiple processes run the same binary (for example, 10 shell instances), the kernel maps the same physical pages of the text segment into each process’s VAS. The code is shared in RAM but appears private to each process through the MMU’s page tables.

Quick check on your Linux system:
cat /proc/$$/maps | grep -E " r-xp "

This shows the executable mappings of your current shell process. The r-xp flag means read+execute, private mapping — that is your text segment.

2. Data Segments — rw-

The data segment stores global and static variables declared in your C program. It is read-write (rw-) because these variables need to be modified at runtime. Internally, it splits into two sub-regions:

Initialized Data (.data)

Variables given an explicit initial value in your source code. E.g., int counter = 100; at file scope. These values are stored in the binary on disk and loaded into RAM at startup.

Uninitialized Data (.bss)

Variables declared without an initial value. E.g., int buffer[1024]; at file scope. The kernel automatically zeroes this region at process startup. The BSS section barely occupies space in the binary on disk — only its size is stored.

/* Examples of data segment variables in C */

int global_init   = 42;          /* .data section  -- initialized  */
int global_uninit;                /* .bss  section  -- zeroed by OS */
static int module_counter = 0;   /* .data section  -- static       */

int main(void) {
    static int call_count;        /* .bss  -- static local, zeroed  */
    /* ... */
    return 0;
}

3. Heap Segment — Dynamic Memory (rw-)

The heap is where dynamically allocated memory lives — memory you request at runtime using malloc(), calloc(), realloc() or the C++ new operator. Unlike the text and data segments, the heap is dynamic: it can grow or shrink as the program allocates and frees memory.

The boundary between the heap and the unmapped region above it is called the program break. The classic brk() and sbrk() system calls move this break upward (to allocate more heap) or downward (to release it). The glibc malloc() implementation calls these internally on your behalf.

However, there is an important detail in modern Linux (including Linux 6.x):

Figure 2 — How malloc() chooses allocation strategy in Linux 6.x
📦
Small allocation
malloc(size) < 128 KB
Uses heap (brk/sbrk)
Fast, contiguous
⟶
🗺️
Large allocation
malloc(size) ≥ 128 KB
Uses anonymous mmap()
Separate VAS mapping

The threshold (128 KB default) is called MMAP_THRESHOLD. It is tunable via mallopt(M_MMAP_THRESHOLD, value).

#include <stdlib.h>
#include <stdio.h>

int main(void)
{
    /* Small: comes from heap segment (brk) */
    char *small = malloc(4096);   /* 4 KB -- from heap */

    /* Large: kernel creates a NEW anonymous mmap region */
    char *large = malloc(1024 * 1024);  /* 1 MB -- anonymous mmap */

    printf("small ptr : %p\n", small);
    printf("large ptr : %p\n", large);

    /* Inspect: cat /proc/<pid>/maps while sleeping */

    free(large);   /* mmap region is released back to OS immediately */
    free(small);   /* heap shrinks only when brk() is called */
    return 0;
}
⚡ Linux Kernel Dev Insight: When free() is called for a large mmap-based allocation, the kernel immediately unmaps those pages and the physical RAM is returned to the system. For small heap allocations, glibc typically keeps the freed memory in its own internal free-list for reuse — actual brk() reduction happens lazily. This is why kernel module code that manages memory directly uses the kernel allocators (kmalloc, vmalloc) instead of userspace malloc.

4. Library Mappings — Shared Libraries (r-x / rw-)

Every modern Linux program uses shared libraries — the most fundamental being the C library (glibc, libc.so.6). These libraries are not copied into your binary at compile time (that would be static linking). Instead, the dynamic linker/loader (ld.so) maps them into the process VAS at runtime, right after execve() starts the process.

Each shared library gets two separate mappings in the process VAS:

  • Text mapping (r-x): The library’s compiled code. Shared in physical RAM across all processes using that library — huge memory saving.
  • Data mapping (rw-): The library’s global/static variables. Each process gets its own copy via Copy-On-Write (COW), so one process cannot affect another’s library state.
# See all library mappings of the bash shell:
cat /proc/$(pgrep -n bash)/maps

# Typical output (abbreviated):
# 7f2a1c000000-7f2a1c200000 r-xp  libc.so.6   <-- libc text
# 7f2a1c200000-7f2a1c400000 ---p  libc.so.6   <-- guard (no access)
# 7f2a1c400000-7f2a1c404000 r--p  libc.so.6   <-- libc read-only data
# 7f2a1c404000-7f2a1c406000 rw-p  libc.so.6   <-- libc writable data

5. Stack Segment — Function Call Machinery (rw-)

The stack is the memory region that implements your programming language’s function-call mechanism. Every time a function is called, the CPU allocates a stack frame for it. This frame holds:

  • The local variables declared inside the function
  • The function arguments (parameters)
  • The return address — where to go after the function finishes
  • Saved registers (CPU state preservation)

When the function returns, its stack frame is destroyed instantly — the Stack Pointer (SP) register is just moved back. This is why returning a pointer to a local variable is a bug in C: the memory is “gone” once the function returns.

Figure 3 — Stack Grows Downward (Fully Descending Stack)
HIGH address
main() frame
func_a() frame
func_b() frame ← SP
unmapped (grows here next)
LOW address ↓
SP (Stack Pointer) always points to the lowest address in use — the “top” of the stack.
Stack grows downward toward address 0. Each new function call pushes SP lower.
On x86-64 and ARM64, the ABI defines exactly how each frame is laid out: arguments, saved registers, locals.
⚠️ Common Mistake — Stack Overflow: If functions call each other recursively too deeply (or a function declares a huge local array), the stack keeps growing downward. Eventually it collides with the heap or library mappings below it. The kernel detects this and sends a SIGSEGV signal to the process. This is a stack overflow. Default stack size on Linux is 8 MB (check with ulimit -s).

Multiple Threads — Each Gets Its Own Stack

A process can contain multiple threads of execution. All threads inside one process share the same VAS — they see the same text, data, heap, and library mappings. This is what makes inter-thread communication fast (no kernel-boundary crossing needed).

However, there is one critical exception: each thread has its own private stack. If threads shared a stack, they could not run in parallel — each function call would corrupt the other thread’s local variables and return addresses. The Linux kernel allocates a separate stack region in the VAS for every thread created with pthread_create().

Figure 4 — Thread Stacks Inside the Shared Process VAS
HIGH ADDRESS
Stack of main() thread [rw-]
Stack of thread_2 [rw-]
Stack of thread_3 [rw-]
📚 Library mappings (SHARED by all threads)
⬆ Heap (SHARED by all threads)
BSS + Data (SHARED by all threads)
Text / Code (SHARED by all threads)
LOW ADDRESS

Inspect a Real Process VAS on Linux 6.x

You do not need a special tool to look at a process’s VAS. Linux exposes it through the procfs virtual filesystem. Every process has a file at /proc/<pid>/maps that lists all its current mappings in real time.

# 1. Run a simple C program in the background
sleep 60 &
PID=$!

# 2. Inspect its VAS
cat /proc/$PID/maps

# Example output (x86-64, Linux 6.x):
# Address range           Perms  Offset  Dev  Inode  Pathname
# 5555555551000-555555556000 r-xp 0000 08:01 12345  /usr/bin/sleep   <-- text
# 555555756000-555555757000 r--p 0000 08:01 12345  /usr/bin/sleep   <-- rodata
# 555555757000-555555758000 rw-p 0000 08:01 12345  /usr/bin/sleep   <-- data
# 555555758000-555555779000 rw-p 0000 00:00 0                        <-- heap
# 7f2a...         r-xp ...         libc.so.6                         <-- libc text
# 7ffd...         rw-p ...         [stack]                           <-- stack

# 3. A more readable summary:
pmap -x $PID

Quick Reference: VAS Segments at a Glance

Segment Permissions Static / Dynamic What lives here Grows toward
Text r-x Static Machine code (instructions) —
Data (.data) rw- Static Initialized global/static vars —
BSS (.bss) rw- Static Uninitialized global/static vars (zeroed) —
Heap rw- Dynamic malloc/calloc allocations (<128KB) Higher addresses ↑
Libraries r-x / rw- Dynamic (loaded at start) Shared library code and data —
Stack rw- Dynamic Local vars, return addresses, function frames Lower addresses ↓

Common Mistakes and Debugging Tips

❌ Returning a pointer to a local variable

Local variables live on the stack. Their stack frame is destroyed when the function returns. Returning a pointer to them causes undefined behaviour. Use malloc() or pass a caller-allocated buffer instead.

⚠️ Heap memory leak (forgetting free())

Every malloc() must have a matching free(). Memory not freed stays allocated until the process exits. Use valgrind --leak-check=full ./your_program to detect leaks.

🔵 Writing to the text segment (SIGSEGV)

String literals like char *s = "hello"; live in the read-only text/rodata segment. Writing to s[0] = 'H'; crashes with SIGSEGV. Use char s[] = "hello"; instead — this puts the string on the stack.

✅ Use /proc/<pid>/maps for debugging

When you get a mysterious SIGSEGV, open a second terminal, find the process PID, and run cat /proc/<pid>/maps. Correlate the faulting address from dmesg/GDB with the segment boundaries to immediately know which region was accessed illegally.

Key Takeaways

  • The Linux process VAS is an isolated, sandboxed virtual address space unique to each process.
  • It is divided into segments / mappings: text (r-x), data (.data + .bss, rw-), heap (rw-), library mappings, and stack (rw-).
  • The heap grows upward; the stack grows downward. On modern Linux, large malloc() allocations (≥128KB) are served via anonymous mmap(), not the heap.
  • Shared libraries are memory-mapped at process startup; their text pages are physically shared across all processes using that library.
  • All threads share the VAS except for their individual stacks — each thread has its own private stack.
  • Use /proc/<pid>/maps or pmap to inspect any live process’s memory layout on Linux 6.x.

Frequently Asked Questions

Q1: Is the process VAS the same as physical RAM?

No. The VAS contains virtual addresses. The CPU’s MMU translates these to physical RAM locations using page tables managed by the Linux kernel. A process can have a VAS much larger than available physical RAM — unused pages are swapped to disk.

Q2: What happens if two processes use the same virtual address?

That is perfectly fine and normal. Virtual addresses are per-process. Two processes can both have a local variable at virtual address 0x7fff...1234 and these map to completely different physical RAM locations.

Q3: Why does the stack grow downward instead of upward?

This is a hardware convention inherited from early CPU architectures (including Intel 8086 and early ARM). The stack is placed at the top of the VAS and grows toward the heap, maximising the distance between them to allow both to grow as large as possible.

Q4: What is ASLR and how does it affect the VAS layout?

ASLR (Address Space Layout Randomisation) is a security feature in all modern Linux kernels that randomises the base addresses of the stack, heap, and library mappings each time a program runs. The segment contents remain the same, but their locations in the VAS change, making exploits harder. Check cat /proc/sys/kernel/randomize_va_space — value 2 means full ASLR is active.

Q5: What is the BSS section? Why is it not stored in the binary?

BSS stands for “Block Started by Symbol” — a historical assembler term. Since all BSS variables are zero-initialised, storing their actual zero bytes in the binary would waste disk space. Instead, the binary only records how large the BSS section needs to be. The kernel allocates the memory and zeroes it when loading the process.

Q6: Can a Linux kernel module see the same VAS as a user process?

No. Kernel modules run in kernel space, which is a completely separate address space region above user space. A kernel module cannot directly dereference a user-space pointer without using the dedicated kernel helper functions copy_from_user() and copy_to_user(), which safely cross the boundary.

Q7: How does the Linux kernel track all the VAS segments?

The kernel stores them as a linked list and red-black tree of struct vm_area_struct (VMA) objects inside the process’s struct mm_struct. Each VMA describes one contiguous range of virtual addresses, its permissions, and the backing file (if any). You can study these in the kernel source at include/linux/mm_types.h.

Q8: Is this VAS layout different between 32-bit and 64-bit Linux?

The structure (text, data, heap, libraries, stack) is the same, but the address ranges differ significantly. A 32-bit process has a VAS of 4 GB (typically 3 GB user + 1 GB kernel on x86). A 64-bit process has a VAS of 128 TB of user space on current x86-64 Linux 6.x — vastly larger, with kernel space occupying the top portion.

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *