Free Linux Kernel Development Course — EmbeddedPathashala
Linux Process Virtual Address Space (VAS): Segments & Memory Layout Explained
Understand how the Linux kernel organises every process in memory — from text and data segments to heap, stack and shared library mappings — with up-to-date Linux 6.x internals.
What You Will Learn
- What the Process Virtual Address Space (VAS) is and why Linux sandboxes every process inside it
- The exact memory segments / mappings every Linux process contains (text, data, heap, stack, libraries)
- How the heap grows dynamically and what
MMAP_THRESHOLDmeans in Linux 6.x - How the stack grows downward and why the “top” of the stack is actually the lowest address
- How shared libraries are mapped into a process’s VAS at runtime
- Why every thread needs its own private stack inside the shared VAS
- How to inspect a live process VAS using
/proc/<pid>/mapson your Linux system
Prerequisites
Before reading this lecture you should be comfortable with:
- Basic C programming (pointers, variables, functions)
- What a process is and how the OS loads a program
- Elementary understanding of memory addresses (virtual vs physical is a plus)
No kernel-level experience is required. This lecture is beginner-friendly and part of the free Linux kernel development course at EmbeddedPathashala.
What is the Process VAS in Linux?
When a program runs on Linux, the operating system does not give it access to all of physical RAM. Instead, the kernel creates a virtual address space (VAS) — an isolated, private view of memory that belongs exclusively to that process. Think of it as a walled garden: the process can only access addresses that exist inside its own VAS. Any attempt to read or write outside this boundary immediately triggers a segmentation fault and the kernel kills the offending process.
This isolation is a cornerstone of modern OS security and stability. One process cannot corrupt another process’s memory, even if they are running the same program simultaneously. The Linux kernel enforces this sandbox using the CPU’s Memory Management Unit (MMU), which translates virtual addresses to physical RAM locations transparently.
In the Linux kernel development world, you will frequently hear “process VAS”, “process image”, and “address space” used interchangeably. They all mean the same thing: the range of virtual memory addresses that a process is allowed to use.
The Memory Segments Inside a Linux Process VAS
The VAS is not one big block of memory. The Linux kernel divides it into distinct regions called segments (the older UNIX term) or, more precisely, mappings (the modern kernel term, since they are managed via the mmap subsystem). Each mapping has its own set of permissions and purpose. Let’s go through each one.
1. Text Segment (Code Segment) — r-x
The text segment holds the compiled machine instructions of your program — what the CPU actually executes. It is marked read and execute only (r-x). Writing to this region is forbidden, which prevents accidental or malicious modification of running code.
This segment is static — its size is fixed at program load time and does not change during execution. On Linux 6.x, the kernel uses W^X (write XOR execute) policy: a memory page can be either writable or executable, never both simultaneously. This is enforced at the hardware level via the CPU’s NX (No-Execute) bit and is a crucial Linux kernel security feature.
When multiple processes run the same binary (for example, 10 shell instances), the kernel maps the same physical pages of the text segment into each process’s VAS. The code is shared in RAM but appears private to each process through the MMU’s page tables.
cat /proc/$$/maps | grep -E " r-xp "
This shows the executable mappings of your current shell process. The r-xp flag means read+execute, private mapping — that is your text segment.
2. Data Segments — rw-
The data segment stores global and static variables declared in your C program. It is read-write (rw-) because these variables need to be modified at runtime. Internally, it splits into two sub-regions:
Variables given an explicit initial value in your source code. E.g., int counter = 100; at file scope. These values are stored in the binary on disk and loaded into RAM at startup.
Variables declared without an initial value. E.g., int buffer[1024]; at file scope. The kernel automatically zeroes this region at process startup. The BSS section barely occupies space in the binary on disk — only its size is stored.
/* Examples of data segment variables in C */
int global_init = 42; /* .data section -- initialized */
int global_uninit; /* .bss section -- zeroed by OS */
static int module_counter = 0; /* .data section -- static */
int main(void) {
static int call_count; /* .bss -- static local, zeroed */
/* ... */
return 0;
}
3. Heap Segment — Dynamic Memory (rw-)
The heap is where dynamically allocated memory lives — memory you request at runtime using malloc(), calloc(), realloc() or the C++ new operator. Unlike the text and data segments, the heap is dynamic: it can grow or shrink as the program allocates and frees memory.
The boundary between the heap and the unmapped region above it is called the program break. The classic brk() and sbrk() system calls move this break upward (to allocate more heap) or downward (to release it). The glibc malloc() implementation calls these internally on your behalf.
However, there is an important detail in modern Linux (including Linux 6.x):
malloc(size) < 128 KB
Fast, contiguous
malloc(size) ≥ 128 KB
Separate VAS mapping
The threshold (128 KB default) is called MMAP_THRESHOLD. It is tunable via mallopt(M_MMAP_THRESHOLD, value).
#include <stdlib.h>
#include <stdio.h>
int main(void)
{
/* Small: comes from heap segment (brk) */
char *small = malloc(4096); /* 4 KB -- from heap */
/* Large: kernel creates a NEW anonymous mmap region */
char *large = malloc(1024 * 1024); /* 1 MB -- anonymous mmap */
printf("small ptr : %p\n", small);
printf("large ptr : %p\n", large);
/* Inspect: cat /proc/<pid>/maps while sleeping */
free(large); /* mmap region is released back to OS immediately */
free(small); /* heap shrinks only when brk() is called */
return 0;
}
free() is called for a large mmap-based allocation, the kernel immediately unmaps those pages and the physical RAM is returned to the system. For small heap allocations, glibc typically keeps the freed memory in its own internal free-list for reuse — actual brk() reduction happens lazily. This is why kernel module code that manages memory directly uses the kernel allocators (kmalloc, vmalloc) instead of userspace malloc.4. Library Mappings — Shared Libraries (r-x / rw-)
Every modern Linux program uses shared libraries — the most fundamental being the C library (glibc, libc.so.6). These libraries are not copied into your binary at compile time (that would be static linking). Instead, the dynamic linker/loader (ld.so) maps them into the process VAS at runtime, right after execve() starts the process.
Each shared library gets two separate mappings in the process VAS:
- Text mapping (
r-x): The library’s compiled code. Shared in physical RAM across all processes using that library — huge memory saving. - Data mapping (
rw-): The library’s global/static variables. Each process gets its own copy via Copy-On-Write (COW), so one process cannot affect another’s library state.
# See all library mappings of the bash shell:
cat /proc/$(pgrep -n bash)/maps
# Typical output (abbreviated):
# 7f2a1c000000-7f2a1c200000 r-xp libc.so.6 <-- libc text
# 7f2a1c200000-7f2a1c400000 ---p libc.so.6 <-- guard (no access)
# 7f2a1c400000-7f2a1c404000 r--p libc.so.6 <-- libc read-only data
# 7f2a1c404000-7f2a1c406000 rw-p libc.so.6 <-- libc writable data
5. Stack Segment — Function Call Machinery (rw-)
The stack is the memory region that implements your programming language’s function-call mechanism. Every time a function is called, the CPU allocates a stack frame for it. This frame holds:
- The local variables declared inside the function
- The function arguments (parameters)
- The return address — where to go after the function finishes
- Saved registers (CPU state preservation)
When the function returns, its stack frame is destroyed instantly — the Stack Pointer (SP) register is just moved back. This is why returning a pointer to a local variable is a bug in C: the memory is “gone” once the function returns.
SIGSEGV signal to the process. This is a stack overflow. Default stack size on Linux is 8 MB (check with ulimit -s).Multiple Threads — Each Gets Its Own Stack
A process can contain multiple threads of execution. All threads inside one process share the same VAS — they see the same text, data, heap, and library mappings. This is what makes inter-thread communication fast (no kernel-boundary crossing needed).
However, there is one critical exception: each thread has its own private stack. If threads shared a stack, they could not run in parallel — each function call would corrupt the other thread’s local variables and return addresses. The Linux kernel allocates a separate stack region in the VAS for every thread created with pthread_create().
Inspect a Real Process VAS on Linux 6.x
You do not need a special tool to look at a process’s VAS. Linux exposes it through the procfs virtual filesystem. Every process has a file at /proc/<pid>/maps that lists all its current mappings in real time.
# 1. Run a simple C program in the background
sleep 60 &
PID=$!
# 2. Inspect its VAS
cat /proc/$PID/maps
# Example output (x86-64, Linux 6.x):
# Address range Perms Offset Dev Inode Pathname
# 5555555551000-555555556000 r-xp 0000 08:01 12345 /usr/bin/sleep <-- text
# 555555756000-555555757000 r--p 0000 08:01 12345 /usr/bin/sleep <-- rodata
# 555555757000-555555758000 rw-p 0000 08:01 12345 /usr/bin/sleep <-- data
# 555555758000-555555779000 rw-p 0000 00:00 0 <-- heap
# 7f2a... r-xp ... libc.so.6 <-- libc text
# 7ffd... rw-p ... [stack] <-- stack
# 3. A more readable summary:
pmap -x $PID
Quick Reference: VAS Segments at a Glance
Common Mistakes and Debugging Tips
Local variables live on the stack. Their stack frame is destroyed when the function returns. Returning a pointer to them causes undefined behaviour. Use malloc() or pass a caller-allocated buffer instead.
Every malloc() must have a matching free(). Memory not freed stays allocated until the process exits. Use valgrind --leak-check=full ./your_program to detect leaks.
String literals like char *s = "hello"; live in the read-only text/rodata segment. Writing to s[0] = 'H'; crashes with SIGSEGV. Use char s[] = "hello"; instead — this puts the string on the stack.
When you get a mysterious SIGSEGV, open a second terminal, find the process PID, and run cat /proc/<pid>/maps. Correlate the faulting address from dmesg/GDB with the segment boundaries to immediately know which region was accessed illegally.
Key Takeaways
- The Linux process VAS is an isolated, sandboxed virtual address space unique to each process.
- It is divided into segments / mappings: text (
r-x), data (.data + .bss,rw-), heap (rw-), library mappings, and stack (rw-). - The heap grows upward; the stack grows downward. On modern Linux, large
malloc()allocations (≥128KB) are served via anonymousmmap(), not the heap. - Shared libraries are memory-mapped at process startup; their text pages are physically shared across all processes using that library.
- All threads share the VAS except for their individual stacks — each thread has its own private stack.
- Use
/proc/<pid>/mapsorpmapto inspect any live process’s memory layout on Linux 6.x.
Frequently Asked Questions
No. The VAS contains virtual addresses. The CPU’s MMU translates these to physical RAM locations using page tables managed by the Linux kernel. A process can have a VAS much larger than available physical RAM — unused pages are swapped to disk.
That is perfectly fine and normal. Virtual addresses are per-process. Two processes can both have a local variable at virtual address 0x7fff...1234 and these map to completely different physical RAM locations.
This is a hardware convention inherited from early CPU architectures (including Intel 8086 and early ARM). The stack is placed at the top of the VAS and grows toward the heap, maximising the distance between them to allow both to grow as large as possible.
ASLR (Address Space Layout Randomisation) is a security feature in all modern Linux kernels that randomises the base addresses of the stack, heap, and library mappings each time a program runs. The segment contents remain the same, but their locations in the VAS change, making exploits harder. Check cat /proc/sys/kernel/randomize_va_space — value 2 means full ASLR is active.
BSS stands for “Block Started by Symbol” — a historical assembler term. Since all BSS variables are zero-initialised, storing their actual zero bytes in the binary would waste disk space. Instead, the binary only records how large the BSS section needs to be. The kernel allocates the memory and zeroes it when loading the process.
No. Kernel modules run in kernel space, which is a completely separate address space region above user space. A kernel module cannot directly dereference a user-space pointer without using the dedicated kernel helper functions copy_from_user() and copy_to_user(), which safely cross the boundary.
The kernel stores them as a linked list and red-black tree of struct vm_area_struct (VMA) objects inside the process’s struct mm_struct. Each VMA describes one contiguous range of virtual addresses, its permissions, and the backing file (if any). You can study these in the kernel source at include/linux/mm_types.h.
The structure (text, data, heap, libraries, stack) is the same, but the address ranges differ significantly. A 32-bit process has a VAS of 4 GB (typically 3 GB user + 1 GB kernel on x86). A 64-bit process has a VAS of 128 TB of user space on current x86-64 Linux 6.x — vastly larger, with kernel space occupying the top portion.

1 Comment