Free Linux Kernel Development Course
Iterating Over the Linux Kernel Task List — Exploring Every Process and Thread
Learn how the Linux kernel organizes all processes and threads in a doubly linked task list, and write a kernel module that iterates over every live task on the system.
What You Will Learn
- ✅ How the Linux kernel organizes all process information in a task list (doubly linked circular list)
- ✅ What
init_taskis and why it is the anchor of the entire process tree - ✅ How to use
for_each_process()to iterate over every process in a kernel module - ✅ How to use
do_each_thread()/while_each_thread()to enumerate every thread - ✅ The role of RCU locking when iterating the task list safely
- ✅ The difference between a process and a thread at the kernel level (
task_structfor both) - ✅ How to read task info: PID, TGID, name, state, and task flags from a kernel module
- ✅ Best practices for safe task list traversal in production Linux device driver code
📋 Prerequisites
- Completed Part 1: Process Context in the Linux Kernel (recommended)
- Understanding of
task_structand thecurrentmacro - Familiarity with Linux kernel modules,
insmod,rmmod,dmesg - Basic C knowledge including pointers and linked lists
- Linux system with kernel headers installed (
linux-headers-$(uname -r))
1. What Is the Kernel Task List?
The Linux kernel needs to track every single process and thread running on the system at any given moment. To do this, it uses a data structure called the task list — a circular doubly linked list that chains all task_struct structures together.
Every time a new process is created (via fork() or clone()), the kernel allocates a new task_struct for it and inserts it into this list. When the process exits, the entry is eventually removed. At any point in time, you can walk this list to see every process and thread currently alive on the system — which is exactly how tools like ps, top, and htop gather process information (though they do it via the /proc filesystem, not directly via the task list).
The list_head Structure — Linux’s Universal Linked List
The Linux kernel has its own generic doubly linked list implementation, and it is used everywhere — for the task list, for memory maps, for device queues, and dozens of other kernel subsystems. Understanding it is fundamental to Linux kernel development.
/* Defined in: include/linux/types.h */
struct list_head {
struct list_head *next;
struct list_head *prev;
};
Inside task_struct, there is a field called tasks of type struct list_head. This is the link that chains the task into the global task list. The kernel provides a rich set of macros in include/linux/list.h to work with these lists — list_for_each(), list_entry(), list_add(), etc.
The list is circular — the last entry’s next points back to init_task. Iterating stops when you arrive back at init_task.
What Is init_task?
init_task (also known as the idle task or swapper) is the very first task created by the kernel itself during boot. It has PID 0 and is the anchor of the entire task list. All other tasks are reachable by walking the list from init_task. It is the starting point for for_each_process() traversal.
init_task) and never appears in the output of ps — but it exists in the task list and is used internally by the scheduler when no other task is runnable.
2. Processes vs Threads at the Kernel Level — The Key Distinction
At the Linux kernel level, there is no separate “thread” data structure. Both processes and threads are represented by task_struct. The distinction is made through the PID and TGID fields:
| Concept | PID | TGID | Meaning |
|---|---|---|---|
| Main Process | 1001 |
1001 |
PID == TGID → this is the main/lead thread |
| Thread 1 | 1002 |
1001 |
PID ≠ TGID → this is a thread of process 1001 |
| Thread 2 | 1003 |
1001 |
PID ≠ TGID → also a thread of process 1001 |
When you use for_each_process(), you iterate over processes only — specifically, the main thread of each process group (where PID == TGID). To enumerate all threads including non-main ones, you use do_each_thread() / while_each_thread().
/* Quick check in kernel code:
* Is this task the main thread (process leader)? */
if (task->pid == task->tgid) {
/* This is the main/lead thread = the process */
} else {
/* This is a non-leader thread */
}
In ps aux, what you see as a “process” maps to a group of tasks with the same TGID. Each individual thread in that group has its own unique PID. This is why on a typical Linux desktop, the number of tasks (threads) is much higher than the number of visible processes.
3. Iterating Over All Processes — for_each_process()
The Linux kernel provides the for_each_process() macro in include/linux/sched/signal.h. This macro walks the task list starting from init_task and visits every process (main thread) in the system. Let’s look at how to use it safely in a kernel module.
The Macro Definition
/* include/linux/sched/signal.h */
#define for_each_process(p) \
for (p = &init_task ; (p = next_task(p)) != &init_task ; )
This is a standard C for loop where:
- Init: Start at
&init_task(the swapper/idle task, PID 0) - Condition: Keep going while
next_task(p)doesn’t bring us back toinit_task(circular list — stop when we’ve gone all the way around) - Body: Your code runs for each process pointer
p
Kernel Module: List All Running Processes
Here is a complete, working kernel module that safely enumerates all processes. It uses RCU read lock for safe traversal — this is mandatory on modern kernels:
#include <linux/init.h>
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/sched.h>
#include <linux/sched/signal.h> /* for_each_process() */
#include <linux/rcupdate.h> /* rcu_read_lock/unlock */
#include <linux/sched/task.h> /* task_lock/unlock */
MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("List all processes via task list traversal");
MODULE_VERSION("1.0");
static int __init list_procs_init(void)
{
struct task_struct *p;
int count = 0;
pr_info("=== Process List Traversal (EmbeddedPathashala) ===\n");
pr_info("%-6s %-6s %-5s %s\n",
"PID", "TGID", "STATE", "COMMAND");
pr_info("-------------------------------------------\n");
/*
* rcu_read_lock() is required when traversing the task list.
* It ensures that task_struct entries are not freed under us
* while we are reading them (RCU = Read-Copy-Update mechanism).
* This is the correct and safe way to traverse the task list.
*/
rcu_read_lock();
for_each_process(p) {
/*
* task_pid_nr() and task_tgid_nr() return the PID/TGID
* in the init namespace — which is what you see in 'ps'.
* Using these is safer than accessing p->pid directly.
*/
pr_info("%-6d %-6d %-5u %s\n",
task_pid_nr(p),
task_tgid_nr(p),
p->__state, /* task state flags */
p->comm); /* command name, max 15 chars */
count++;
}
rcu_read_unlock();
pr_info("-------------------------------------------\n");
pr_info("Total processes found: %d\n", count);
pr_info("=== End of Process List ===\n");
return 0; /* 0 = success, module stays loaded */
}
static void __exit list_procs_exit(void)
{
pr_info("list_procs: module removed\n");
}
module_init(list_procs_init);
module_exit(list_procs_exit);
rcu_read_lock() is mandatory: The task list is protected by RCU (Read-Copy-Update) — one of the Linux kernel’s most important synchronization mechanisms. Without holding the RCU read lock, a task could be freed by another CPU while you are still reading its fields, causing a use-after-free bug or kernel crash. Always bracket task list traversal with rcu_read_lock() / rcu_read_unlock().
Makefile for the Module
# Makefile — EmbeddedPathashala kernel module build
obj-m += list_procs.o
KDIR := /lib/modules/$(shell uname -r)/build
all:
make -C $(KDIR) M=$(PWD) modules
clean:
make -C $(KDIR) M=$(PWD) clean
Building and Testing
# Install kernel headers if not already installed
sudo apt install linux-headers-$(uname -r) # Debian/Ubuntu
sudo dnf install kernel-devel # Fedora/RHEL
# Build the module
make
# Insert the module and immediately check output
sudo insmod list_procs.ko
dmesg | tail -60
# Remove when done
sudo rmmod list_procs
After inserting the module, you will see output in dmesg listing every process on your system — PID, TGID, state, and name. Compare this with ps aux output and notice they match up (minus the kernel’s internal PID 0 swapper/idle task which ps may not show).
4. Iterating Over All Threads — do_each_thread() and while_each_thread()
When you need to enumerate not just processes but every thread (including all non-main threads), the Linux kernel provides a pair of macros designed to be used together: do_each_thread() and while_each_thread().
These macros implement a nested loop. The outer loop walks across process group leaders (the TGID holders). The inner loop walks within each process group, visiting every thread that belongs to it.
Thread Iteration Structure
for_each_process() visits only the main thread (PID=TGID row) of each process.do_each_thread() / while_each_thread() visits ALL rows — every thread in every group.
Kernel Module: Enumerate Every Thread
#include <linux/init.h>
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/sched.h>
#include <linux/sched/signal.h>
#include <linux/rcupdate.h>
MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("List all threads in the kernel task list");
MODULE_VERSION("1.0");
static int __init list_threads_init(void)
{
struct task_struct *g; /* g = group leader (process) */
struct task_struct *t; /* t = individual thread */
int total_procs = 0;
int total_threads = 0;
pr_info("=== Thread List Traversal (EmbeddedPathashala) ===\n");
pr_info("%-7s %-7s %s\n", "PID", "TGID", "COMMAND");
pr_info("---------------------------------------------\n");
rcu_read_lock();
/*
* do_each_thread(g, t) ... while_each_thread(g, t)
*
* g = task_struct of the process group leader (TGID holder)
* t = task_struct of each thread within that group
*
* The outer macro (do_each_thread) walks across process leaders.
* For each leader, it walks all threads in the thread group.
* When you enter the loop, t and g are set up correctly.
*/
do_each_thread(g, t) {
pr_info("%-7d %-7d %s%s\n",
task_pid_nr(t),
task_tgid_nr(t),
t->comm,
/* Mark process group leaders for clarity */
(t == g) ? " [PROC]" : " [THR]");
total_threads++;
if (t == g)
total_procs++;
} while_each_thread(g, t);
rcu_read_unlock();
pr_info("---------------------------------------------\n");
pr_info("Processes (group leaders): %d\n", total_procs);
pr_info("Total threads : %d\n", total_threads);
pr_info("=== End of Thread List ===\n");
return 0;
}
static void __exit list_threads_exit(void)
{
pr_info("list_threads: module removed\n");
}
module_init(list_threads_init);
module_exit(list_threads_exit);
5. Reading Task State — Understanding Process States in the Kernel
Each task_struct has a __state field (renamed from state in kernel 5.14) that tells you the current scheduling state of that task. This is important when writing monitoring tools or performance-sensitive Linux device drivers.
| State Constant | Value | Meaning |
|---|---|---|
TASK_RUNNING | 0x00 | On the run queue — running or ready to run |
TASK_INTERRUPTIBLE | 0x01 | Sleeping, can be woken by signal (most common sleep) |
TASK_UNINTERRUPTIBLE | 0x02 | Deep sleep — cannot be interrupted even by signals (D state in ps) |
__TASK_STOPPED | 0x04 | Stopped by SIGSTOP or debugger (T state in ps) |
EXIT_ZOMBIE | 0x20 | Process has exited but parent hasn’t called wait() yet (Z state) |
EXIT_DEAD | 0x10 | Process is being removed from the system |
/* Helper to decode task state into a readable string */
static const char *task_state_str(struct task_struct *t)
{
switch (t->__state) {
case TASK_RUNNING: return "R (running)";
case TASK_INTERRUPTIBLE: return "S (sleeping)";
case TASK_UNINTERRUPTIBLE:return "D (disk wait)";
case __TASK_STOPPED: return "T (stopped)";
default:
if (t->exit_state == EXIT_ZOMBIE) return "Z (zombie)";
if (t->exit_state == EXIT_DEAD) return "X (dead)";
return "? (unknown)";
}
}
state field in task_struct was renamed to __state in Linux 5.14. If you write code that needs to support older kernels, use #if LINUX_VERSION_CODE < KERNEL_VERSION(5,14,0) guards. The linux/version.h header provides LINUX_VERSION_CODE for this purpose.
6. RCU — Why Task List Traversal Needs Read-Copy-Update Locking
RCU (Read-Copy-Update) is one of the most powerful and widely-used synchronization mechanisms in the Linux kernel. Understanding why you need it for task list traversal is essential for anyone doing serious Linux kernel programming.
The Problem Without RCU
Consider this scenario: You start iterating the task list. Midway through, another CPU’s process calls exit() and the kernel begins cleaning up that task’s task_struct. Without any protection, you might be reading a task_struct that is simultaneously being freed — a classic use-after-free bug that causes memory corruption or kernel panic.
CPU 1 simultaneously frees that task_struct after process X exits.
Result: Kernel crash / memory corruption.
CPU 1 marks it for deletion but must wait for the RCU grace period.
Result: Safe traversal, no crash.
rcu_read_lock() / rcu_read_unlock() you must not sleep. Do not call any function that can block (no mutex_lock(), no kmalloc(GFP_KERNEL)). Use GFP_ATOMIC if you need to allocate inside. Keep the RCU read section as short as possible.
/* CORRECT — Safe task list traversal with RCU */
rcu_read_lock();
for_each_process(p) {
/* Safe to read p's fields here */
pr_info("%s [%d]\n", p->comm, task_pid_nr(p));
/* DO NOT call blocking functions here!
* No mutex_lock(), no kmalloc(GFP_KERNEL), no schedule() */
}
rcu_read_unlock();
/* After unlock, p may no longer be valid — do NOT use p here */
7. Best Practices for Task List Traversal in Linux Kernel Modules
Never traverse the task list without RCU read lock protection. This is non-negotiable in modern Linux kernel code.
Prefer task_pid_nr(p) over p->pid, and task_tgid_nr(p) over p->tgid. These handle namespace translation correctly and are safer across kernel versions.
Only read what you need inside the RCU lock. If you need to process data heavily, copy it out first, release the lock, then process.
Task list traversal is read-only by design. Modifying task fields while iterating requires proper write-side locking (tasklist_lock, etc.) — only the kernel scheduler and process management code should do this.
The RCU read-side critical section must be non-blocking. No mutex_lock(), no msleep(), no kmalloc(GFP_KERNEL) inside rcu_read_lock().
for_each_process(), rcu_read_lock(), and many task-related symbols are GPL-only exports. Without MODULE_LICENSE("GPL") your module will fail to load with an “Unknown symbol” error.
🎯 Key Takeaways — Kernel Task List Iteration
- The Linux kernel maintains all process and thread info in a circular doubly linked list called the task list, anchored at
init_task. for_each_process(p)iterates over all processes (group leaders where PID == TGID).do_each_thread(g, t) ... while_each_thread(g, t)iterates over every thread in every process.- All task list traversal must be protected with
rcu_read_lock()/rcu_read_unlock()— no exceptions. - Inside RCU read sections, never sleep, never block, and never modify task fields.
- The
task_struct.__statefield (renamed fromstatein kernel 5.14) tells you whether a task is running, sleeping, stopped, or zombie. - This knowledge directly powers real-world tools: kernel monitors, security scanners, schedulers, and Linux device drivers that manage per-task resources.
Frequently Asked Questions — Linux Kernel Task List
📚 Authoritative References
Conclusion
In this tutorial, part of our free Linux kernel development course on EmbeddedPathashala, we explored one of the most interesting and practically useful capabilities you can implement in a kernel module: iterating over the Linux kernel’s task list to inspect every process and thread on the system.
We covered the internal structure of the task list as a circular doubly linked list, the role of init_task as its anchor, the macros for_each_process() and do_each_thread(), the critical importance of RCU locking for safe traversal, and how to decode task states. We also wrote two complete, working kernel modules with proper Makefiles that you can build and test on your own system today.
This kind of direct access to kernel internals is what separates a developer who merely uses Linux from one who understands Linux at a deep level — which is exactly the foundation you need for serious embedded systems engineering, Linux device driver development, and kernel debugging work.
In the next lecture, we will go further into the kernel’s memory management architecture, exploring how virtual memory is organized per process and how the kernel’s memory subsystem interacts with the scheduler and device drivers.
