Intermediate
Linux 6.x
Free
Understanding the Linux Kernel Mode Stack: A Complete Guide
Every process running on a Linux system has two separate stacks — one that operates in user space and one that operates in kernel space. In this tutorial, part of our free Linux kernel programming course, we focus on the kernel mode stack — what it is, why it exists, how Linux 6.x manages it, and how you can inspect it using the /proc filesystem. Understanding process stacks is essential for anyone learning Linux kernel development or working on Linux device drivers.
🎓 What You Will Learn
- What a kernel mode stack is and why every process has one
- How Linux 6.x allocates and manages the per-process kernel stack
- The structure of a kernel call frame and what each frame means
- How to read the kernel stack of any running process using
/proc/PID/stack - How to interpret kernel stack backtraces including the
??frames - How system calls map to kernel functions (the
SYSCALL_DEFINEnpattern) - Why virtual addresses in stack traces are zeroed out on modern kernels
- Common mistakes beginners make when reading kernel backtraces
What Is the Linux Kernel Mode Stack?
When a running process makes a system call — such as read(), write(), or fork() — control transfers from user space into the Linux kernel. At that moment, the CPU switches from user mode to kernel mode. For the kernel to execute the system call’s logic, it needs a private stack. This private stack is what we call the kernel mode stack.
Every process (and every kernel thread) in Linux has its own dedicated kernel stack. This is completely separate from the user-space stack that grows and shrinks as user-land functions call each other. The kernel stack only exists and is only used while that process is running inside the kernel — either executing a system call, handling an interrupt, or running kernel thread logic.
Why Does the Kernel Need a Separate Stack?
The user-space stack is untrusted. A malicious or buggy program could corrupt its own stack or set the stack pointer to an arbitrary address. If the kernel used the same stack, a compromised process could manipulate kernel execution. By maintaining its own separate, kernel-only stack, the kernel protects its execution context from user-space interference. This is a fundamental security boundary in the Linux process model.
(grows downward ↓)
(grows downward ↓)
Kernel Stack Size in Linux 6.x
On x86_64 systems running Linux 6.x, the default kernel stack size for each process is 16 KiB (two pages). This is a compile-time constant controlled by the THREAD_SIZE macro. On 32-bit x86 systems and some embedded architectures, the default is 8 KiB. The stack size cannot be changed at runtime — it is fixed when the process is created.
How to Inspect the Kernel Mode Stack Using /proc
The Linux kernel exposes the current kernel stack of any process through the /proc pseudo-filesystem. Each running process has an entry at /proc/PID/stack. Reading this file gives you a snapshot of what kernel functions the process is currently executing, displayed as a call chain from the most recent function back to the entry point.
This is one of the most powerful debugging techniques available to Linux kernel developers and is heavily used during driver development and kernel tracing.
Reading the Kernel Stack via /proc
You need root privileges to read another process’s kernel stack. This is a deliberate security restriction — exposing the call chain of a process could reveal sensitive information about what kernel subsystem is running and potentially the kernel’s internal address layout.
The basic command is:
# Find the PID of a process first
$ ps aux | grep bash
# Then read its kernel stack (replace 1234 with actual PID)
$ sudo cat /proc/1234/stack
Example Output and Meaning
Here is an example of what you might see when reading the kernel stack of a bash process that is waiting for a child to finish:
$ sudo cat /proc/1234/stack
[<0>] do_wait+0x1cb/0x230
[<0>] kernel_wait4+0x89/0x130
[<0>] __do_sys_wait4+0x95/0xa0
[<0>] __x64_sys_wait4+0x1e/0x20
[<0>] do_syscall_64+0x5a/0x120
[<0>] entry_SYSCALL_64_after_hwframe+0x44/0xa9
Decoding Each Line of the Kernel Stack Backtrace
Each line in the /proc/PID/stack output follows a specific format. Let’s break it down:
| Part | Example | Meaning |
|---|---|---|
[<0>] |
[<0>] |
Placeholder for the function’s virtual memory address. Zeroed for security on modern kernels (prevents KASLR bypass). |
| Function name | do_wait |
Name of the kernel function executing at this call frame. |
+0x1cb |
+0x1cb/0x230 |
Offset in bytes from the start of the function where execution currently is. |
/0x230 |
+0x1cb/0x230 |
Total size of the function in bytes. |
?? |
?? |
The kernel cannot reliably decode this frame. Treat it as noise and ignore it — it is a leftover artifact, not a real function call. |
Why Are Addresses Shown as Zero?
On modern Linux kernels (since around 4.14), the virtual address shown in [<0>] is deliberately zeroed. In older kernels, the real function address was printed here. Exposing these addresses would let an attacker defeat KASLR (Kernel Address Space Layout Randomization) — a security feature that randomises where the kernel loads itself in memory each boot. Zeroing out the address in proc output removes this information leak.
Understanding How System Calls Map to Kernel Functions
When you call a system call like wait4() from a C program or shell, the Linux kernel does not directly jump to a function called wait4. Instead, there is a layered naming and dispatch system. Understanding this is crucial for reading kernel stack traces.
The SYSCALL_DEFINEn Pattern
In the Linux kernel source, system calls are defined using a family of macros: SYSCALL_DEFINE0, SYSCALL_DEFINE1, …, SYSCALL_DEFINE6. The number at the end indicates how many parameters the system call takes from user space (the range is 0 to 6).
/* Example: wait4 system call definition in kernel/exit.c */
SYSCALL_DEFINE4(wait4,
pid_t, upid,
int __user *, stat_addr,
int, options,
struct rusage __user *, ru)
{
/* ... actual implementation ... */
return kernel_wait4(upid, stat_addr, options, ru);
}
This macro expands at compile time into a function named __x64_sys_wait4 (on x86_64) or the architecture-appropriate equivalent. This is why you see __x64_sys_wait4 in the stack trace rather than just wait4.
Architecture entry point — assembly stub
Looks up syscall table by syscall number
Generated by SYSCALL_DEFINE4(wait4,…)
Actual implementation — worker functions
The key insight is: a user-space call to foo() typically maps to a kernel chain: entry_SYSCALL_64 → do_syscall_64 → __x64_sys_foo → sometimes do_foo() as the actual worker. When you read a kernel stack trace, you are seeing this exact chain.
Practical Example: Tracing a Shell Process
Let’s walk through a real scenario. The Bash shell forks a child process every time you run a command. After forking, the parent Bash process calls wait4() to wait for the child to exit. If you read the kernel stack of the parent Bash while the child is still running, you will see exactly this wait chain in the backtrace.
# Step 1: Find Bash PID
$ echo $$
5421
# Step 2: In another terminal, read the kernel stack
$ sudo cat /proc/5421/stack
[<0>] do_wait+0x1cb/0x230
[<0>] kernel_wait4+0x89/0x130
[<0>] __do_sys_wait4+0x95/0xa0
[<0>] __x64_sys_wait4+0x1e/0x20
[<0>] do_syscall_64+0x5a/0x120
[<0>] entry_SYSCALL_64_after_hwframe+0x44/0xa9
Reading this bottom-up: Bash entered the kernel at entry_SYSCALL_64_after_hwframe, the kernel dispatched the wait4 system call via do_syscall_64, the arch-specific stub __x64_sys_wait4 called __do_sys_wait4, which called kernel_wait4, which is currently blocked inside do_wait waiting for the child process.
/proc/PID/stack file is a live snapshot. Reading it while the process is blocked in a system call gives you a stable picture. Reading it while the process is actively running in user space may show an empty or minimal stack because the process isn’t in kernel mode at that instant.
Common Mistakes When Reading Kernel Stack Backtraces
- Reading top-down instead of bottom-up: The most common mistake. Always start from the bottom. The bottom is the oldest frame (entry point), the top is the most recent (current position).
- Panicking about
??frames: These are invalid leftover frames from stack memory that wasn’t fully initialised. They are not real functions — ignore them completely. - Expecting real addresses in
[<0>]: Since Linux 4.14+, these are always zero due to KASLR protection. You will not get real kernel addresses here without root-level tooling likebpftrace. - Confusing the syscall wrapper with the real function:
__x64_sys_wait4is just a thin wrapper. The real work happens indo_wait. Look for thedo_xxx()functions for the actual implementation. - Trying to read the stack of a zombie process: Zombie processes have already exited. Their kernel stack is gone.
/proc/PID/stackwill be empty or return an error.
Security Considerations: Why /proc/PID/stack Requires Root
The /proc/PID/stack interface requires CAP_SYS_PTRACE capability (which normally means root). This is not arbitrary — there are real security reasons:
- The stack trace reveals which kernel subsystems a process is using, which could be used for reconnaissance in an attack.
- Even though addresses are zeroed in modern kernels, the function names themselves reveal internal kernel state.
- A process waiting inside a cryptographic operation would reveal what crypto functions are in use, potentially aiding side-channel analysis.
/proc/PID/stack output from a production system in public bug reports without sanitising it first. While addresses are zeroed on modern kernels, function names can still reveal sensitive system information.
Key Differences: Kernel Stack in Linux 6.x vs Older Kernels
| Feature | Linux 4.x / 5.x (older) | Linux 6.x (current) |
|---|---|---|
| Default stack size (x86_64) | 8 KiB (1 page) | 16 KiB (2 pages) — configurable |
| Address shown in /proc/PID/stack | Real address (pre-4.14) or zero | Always zero (KASLR protection) |
| Stack guard pages | Optional, not always enabled | Enabled by default |
| Stack depot for debugging | Limited | Improved, used by KASAN/KMSAN |
| Syscall naming | SyS_foo() wrappers common |
__x64_sys_foo() pattern standardised |
Best Practices for Kernel Stack Analysis
- Always read bottom-up: Internalise this habit. Write it on a sticky note if needed.
- Combine with strace for full picture: Use
strace -p PIDalongside/proc/PID/stackto correlate user-space system call activity with kernel stack state. - Check /proc/PID/syscall too: This file shows the current system call number and its arguments for a blocked process — a natural companion to
/proc/PID/stack. - Use kernel debug builds for development: When writing kernel modules or drivers, compile the kernel with
CONFIG_DEBUG_INFO=yandCONFIG_KASAN=yto get better stack information. - Ignore ?? frames: Do not waste time trying to decode them. They are artifacts.
🏆 Key Takeaways
- Every Linux process has a dedicated kernel mode stack, separate from its user-space stack.
- In Linux 6.x, the default kernel stack size on x86_64 is 16 KiB.
- You can inspect the kernel stack of any process using
sudo cat /proc/PID/stack. - Always read the output bottom-up — bottom is the entry point, top is the current function.
- The
[<0>]addresses are intentionally zeroed for KASLR security. - Syscalls like
wait4()map to kernel functions through theSYSCALL_DEFINEnmacro system, producing__x64_sys_wait4→kernel_wait4→do_waitchains. - Ignore
??frames — they are not real functions.
Frequently Asked Questions (FAQ)
You need root privileges or the CAP_SYS_PTRACE capability. Run the command with sudo. This restriction exists for security reasons — kernel stack information could be used for reconnaissance.
A process only has a kernel stack when it is actively executing inside the kernel — during a system call, interrupt handling, or kernel thread work. If a process is running purely in user space at the moment you read the file, the stack may appear empty because it is not in kernel mode.
The format is offset/total_size. +0x1cb means execution is currently 459 bytes into the function. /0x230 means the function is 560 bytes total. This helps you pinpoint the exact location within a function during debugging.
Yes. Kernel threads also appear in /proc/ with their own PIDs (shown in square brackets in the output of ps aux, e.g., [kworker/0:1]). You can read their stack the same way.
The user-space stack is used for user-mode function calls and grows dynamically (limited by ulimit). The kernel stack is a fixed-size private stack used only when the process is running inside the kernel. They are completely separate memory regions.
No. Reading /proc/PID/stack is a non-invasive, read-only operation. The kernel takes a snapshot and returns it. The monitored process is unaffected and does not pause or slow down.
The Linux kernel limits system calls to a maximum of 6 arguments passed from user space. This is an ABI constraint tied to the number of CPU registers available for passing arguments during a system call transition. Most system calls take far fewer than 6 arguments.
Yes. When the kernel prints an oops or panic, it dumps the kernel stack of the faulting process in the same bottom-up format. Understanding /proc/PID/stack prepares you to read crash dumps and kernel oops messages effectively.
Conclusion
The kernel mode stack is one of the most fundamental data structures in the Linux process model. Every process that makes a system call relies on it, and every kernel developer working on device drivers or kernel modules must understand how it works. In this free Linux kernel programming tutorial, you learned how the kernel stack is allocated, how to inspect it using /proc/PID/stack, and how to decode the call frames you find there. You now understand why addresses are zeroed, how system calls map to kernel functions, and what the SYSCALL_DEFINEn pattern means in practice.
In the next lecture, we explore the user-space stack — how to inspect it using tools like gdb, and the modern approach using eBPF and the BCC toolkit to view both stacks simultaneously.
