How Are Linux Processes and Threads Connected? – Free Linux Device Drivers Course in Hyderabad

Linux Kernel Process & Thread Iteration
How the kernel walks every process and thread — task lists, TGID, PID and task_struct internals explained simply
Free
Linux Kernel Course
Beginner
Friendly
Kernel
Internals

Welcome to this lecture of the free Linux kernel development course on EmbeddedPathashala. In this article you will learn exactly how the Linux kernel keeps track of every running process and every thread on the system, and how you can write a kernel module that walks that list yourself. No prior kernel module experience is assumed beyond knowing how to build and load a simple .ko file.

What You Will Learn

  • What the Linux kernel’s task list is and why it is a circular doubly-linked list
  • How for_each_process() iterates over every process (main thread only)
  • How do_each_thread() / while_each_thread() iterates over every thread
  • The difference between PID and TGID and why both exist in task_struct
  • How user-space tools like ps -LA expose the same kernel data
  • How to write your own kernel module that lists all processes or all threads
  • Common mistakes beginners make when iterating the task list

Prerequisites

  • Basic C programming knowledge
  • Linux installed (Ubuntu 22.04 LTS or later recommended)
  • Kernel headers installed: sudo apt install linux-headers-$(uname -r)
  • Ability to build and load a simple kernel module (insmod / rmmod)

1. The Linux Kernel Task List

Every time a process or thread is created on a Linux system, the kernel allocates a data structure called task_struct. This structure holds everything the kernel needs to know about a running task — its state, its memory map, its open files, its identity, and much more.

All of these task_struct instances are linked together through a circular doubly-linked list. Think of it like a ring — you can start at any point and keep walking forward until you arrive back where you started. The kernel provides a fixed entry point into this ring called init_task, which represents the very first process on the system (historically init, on modern systems systemd with PID 1).

Linux Kernel Task List — Circular Doubly-Linked List
Entry Point
init_task
(systemd / PID 1)
⇄
task_struct
kthreadd
PID 2
⇄
task_struct
bash
PID N
⇄
…
More tasks
wraps back
⇄
Back to
init_task
circular!

Each box is one task_struct. The list is circular — walking it from init_task eventually returns to init_task.

The kernel provides macros that hide the complexity of walking this circular list safely. You do not dereference pointers manually — you use these macros and let the kernel abstractions do the heavy lifting.

2. Iterating Over All Processes with for_each_process()

The first macro to know is for_each_process(). It walks the task list and gives you the task_struct pointer for the main thread of every process. It does not visit the individual threads that may exist within a multithreaded process. That distinction is important and we will come back to it shortly.

How to Use for_each_process() in a Kernel Module

Here is a minimal but complete kernel module that iterates over all processes and prints some basic information about each one:

// prcs_show.c — Linux kernel module to list all processes
// Tested on Linux 6.x kernels

#include <linux/init.h>
#include <linux/module.h>
#include <linux/sched.h>         // task_struct, for_each_process
#include <linux/sched/signal.h>  // signal_struct
#include <linux/cred.h>          // current_uid()

MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("List all processes using for_each_process()");

static int __init prcs_show_init(void)
{
    struct task_struct *p;

    pr_info("prcs_show: loaded\n");
    pr_info("%-20s  %6s  %6s  %6s  %6s\n",
            "Name", "TGID", "PID", "RUID", "EUID");

    /*
     * for_each_process(p) expands to a for-loop over the circular
     * task list. Each iteration gives us the main thread of one process.
     * The RCU read lock protects the list from being modified while
     * we are reading it.
     */
    rcu_read_lock();
    for_each_process(p) {
        const struct cred *cred = __task_cred(p);
        pr_info("%-20s  %6d  %6d  %6u  %6u\n",
                p->comm,
                p->tgid,
                p->pid,
                from_kuid(&init_user_ns, cred->uid),
                from_kuid(&init_user_ns, cred->euid));
    }
    rcu_read_unlock();

    return 0;
}

static void __exit prcs_show_exit(void)
{
    pr_info("prcs_show: removed\n");
}

module_init(prcs_show_init);
module_exit(prcs_show_exit);

When you build and load this module (sudo insmod prcs_show.ko), the kernel log (dmesg) will show one line per process. The key thing to notice is that the TGID and PID columns will always be equal for every row — because for_each_process() only hands you the main thread of each process, and for the main thread those two values are always the same.

⚠ Important Note: Always use rcu_read_lock() / rcu_read_unlock() when iterating the task list. The task list is protected by RCU (Read-Copy-Update). Forgetting this protection can cause crashes on SMP systems because a task could be freed while you are still referencing it.

3. Iterating Over All Threads with do_each_thread()

If you want to visit every single thread on the system — not just the main thread of each process — you need a different pair of macros: do_each_thread(g, t) and while_each_thread(g, t).

These macros use two task_struct pointers: g (the group leader / main thread) and t (the individual thread being visited). The outer loop walks processes; the inner loop walks the threads within each process. Together they visit every thread alive on the system.

Kernel Module Using do_each_thread()

// thrd_show.c — List every thread on the system
// Linux 6.x compatible

#include <linux/init.h>
#include <linux/module.h>
#include <linux/sched.h>
#include <linux/sched/signal.h>

MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("List every thread using do_each_thread / while_each_thread");

static int __init thrd_show_init(void)
{
    struct task_struct *g, *t;  // g = group leader, t = thread
    unsigned long total = 0;

    pr_info("thrd_show: loaded\n");
    pr_info("%-20s  %6s  %6s\n", "Thread Name", "TGID", "PID");

    rcu_read_lock();
    do_each_thread(g, t) {
        /*
         * If PID == TGID, this is the main thread of the process.
         * We mark kernel threads with square brackets (like ps does).
         */
        if (t->mm == NULL) {
            /* kernel thread — no user-space memory map */
            pr_info("[%-18s]  %6d  %6d\n", t->comm, t->tgid, t->pid);
        } else {
            pr_info("%-20s  %6d  %6d\n", t->comm, t->tgid, t->pid);
        }
        total++;
    } while_each_thread(g, t);
    rcu_read_unlock();

    pr_info("thrd_show: total threads on system = %lu\n", total);
    return 0;
}

static void __exit thrd_show_exit(void)
{
    pr_info("thrd_show: removed\n");
}

module_init(thrd_show_init);
module_exit(thrd_show_exit);

When you load this module you will see many more rows than the process-only version. A typical desktop Linux system runs 150–300 threads across perhaps 80–120 processes. Every thread gets its own row, and threads belonging to the same process share the same TGID value but have different PID values.

One Process with Multiple Threads — TGID vs PID
Thread Name TGID PID Role
MyApp (main) 1200 1200 Main Thread — TGID == PID
worker-1 1200 1201 Child Thread — shares TGID
worker-2 1200 1202 Child Thread — shares TGID
io-handler 1200 1203 Child Thread — shares TGID

All threads of the same process share one TGID. Each thread has its own unique PID inside the kernel.

4. Understanding TGID vs PID — The Core Concept of This Free Linux Kernel Development Course

This is the most important concept in this lecture, so let us spend time on it.

Inside the Linux kernel, every thread — whether it is the main thread of a process or a worker thread — has its own task_struct. And every task_struct has a pid field. This means, at the kernel level, every thread has a unique identifier called a PID.

Now here is where POSIX comes in. The POSIX standard (pthreads) says that all threads of the same process must share a common process ID. But the kernel was giving every thread a unique PID internally. This contradiction meant Linux was, for a time, not compliant with the POSIX threading standard and caused porting problems.

The fix was elegant: add a new field to task_struct called tgid — the Thread Group IDentifier. Here is how it works:

  • For a single-threaded process: tgid == pid. Both values are the same.
  • For the main thread of a multithreaded process: tgid == pid. Same as above.
  • For a child thread of a multithreaded process: tgid equals the main thread’s pid; the thread keeps its own unique pid.

So when user space asks “what is the PID of this process?”, the kernel returns tgid. When user space asks “what is this thread’s LWP (lightweight process) number?”, the kernel returns pid. This is exactly what you see in the output of ps -LA:

# Run this in your terminal to see the TGID/PID split for all threads:
$ ps -LA

# Output columns:
#  PID  — the kernel's tgid (shown as "process ID" to user space)
#  LWP  — the kernel's pid  (the unique thread/lightweight-process ID)

# Example: a multithreaded app might show:
#   PID    LWP  TTY   CMD
#   1500  1500  pts/0  myapp       ← main thread: PID == LWP (TGID == PID)
#   1500  1501  pts/0  myapp       ← child thread: PID same, LWP different
#   1500  1502  pts/0  myapp       ← child thread: PID same, LWP different

From inside the kernel module, you access the same data directly:

// Inside a kernel module, for any task_struct pointer 't':

t->pid   // the kernel's internal unique thread ID (== LWP in ps output)
t->tgid  // the process group ID shown as PID to user space

// Check if this is the main thread:
if (t->pid == t->tgid) {
    pr_info("%s is the main thread\n", t->comm);
} else {
    pr_info("%s is a child thread of TGID %d\n", t->comm, t->tgid);
}

5. Identifying Kernel Threads vs User-Space Threads

When you iterate the task list, you will encounter two kinds of tasks: kernel threads and user-space processes/threads. A simple and reliable way to tell them apart is by checking the mm pointer inside task_struct.

Every user-space task has a memory map — a struct mm_struct that describes the virtual address space for that process. Kernel threads have no user-space memory map at all; they run entirely in kernel space. So t->mm == NULL means it is a kernel thread.

// How to detect a kernel thread vs a user-space thread:

rcu_read_lock();
do_each_thread(g, t) {
    if (t->mm == NULL) {
        // This is a kernel thread (e.g., kworker, ksoftirqd, kswapd)
        pr_info("[%s]  tgid=%d  pid=%d  (kernel thread)\n",
                t->comm, t->tgid, t->pid);
    } else {
        // This is a user-space process or thread
        pr_info("%s  tgid=%d  pid=%d  (user space)\n",
                t->comm, t->tgid, t->pid);
    }
} while_each_thread(g, t);
rcu_read_unlock();

By convention (the same one ps uses), kernel threads are displayed with their name in square brackets: [kworker/0:0], [ksoftirqd/0], etc. Your kernel module output will look cleaner and more professional if you follow this same convention.

6. Building Your Kernel Modules — Makefile

Here is a generic Makefile you can use for any of the modules above. Save it in the same directory as your .c file:

# Makefile for simple kernel modules

# Change this to your module's .c filename (without extension)
MODULE_NAME := prcs_show

obj-m += $(MODULE_NAME).o

# Kernel build directory — adjust if your kernel source is elsewhere
KDIR := /lib/modules/$(shell uname -r)/build

all:
	make -C $(KDIR) M=$(PWD) modules

clean:
	make -C $(KDIR) M=$(PWD) clean
# Build, load, check output, then unload:
make
sudo insmod prcs_show.ko
dmesg | tail -60
sudo rmmod prcs_show

7. Common Mistakes and Troubleshooting

Common Mistakes When Iterating the Task List
Mistake Why It Is a Problem Correct Approach
Not using rcu_read_lock() A task can be freed mid-iteration on SMP; causes kernel panic Always wrap iteration in rcu_read_lock() / rcu_read_unlock()
Using for_each_process() to count all threads Only visits main threads; misses child threads completely Use do_each_thread() / while_each_thread()
Sleeping inside the task-list loop RCU read lock must not be held across scheduling points Do all processing outside the locked section, or use a local list
Confusing pid and tgid Printing pid as “process ID” shows wrong value to user space Show tgid as the process ID and pid as the thread/LWP ID
Dereferencing task->mm without a null check Kernel threads have mm == NULL; null dereference crashes the kernel Always check if (t->mm) before accessing user-space memory info

8. Best Practices for Task List Iteration in Linux Kernel Programming

  1. Always use RCU protection — wrap your loop in rcu_read_lock() and rcu_read_unlock().
  2. Never sleep inside the loop — do not call any function that can sleep (like kmalloc(GFP_KERNEL) or copy_to_user()) while holding the RCU read lock.
  3. Use task_lock() if you access credentials — the cred pointer inside task_struct can change; protect it with task_lock(t) / task_unlock(t).
  4. Keep the critical section short — do minimal work inside the loop; copy what you need and process it afterwards.
  5. Test on a VM first — kernel module bugs can crash or corrupt the system; always test on a virtual machine before running on real hardware.
  6. Check signal_pending() in long loops — if you are iterating hundreds of tasks, check periodically whether the module should exit early.

9. Quick Reference — for_each_process vs do_each_thread

Comparison: Process Iteration vs Thread Iteration Macros
Feature for_each_process(p) do_each_thread(g,t) / while_each_thread(g,t)
What it visits Main thread of every process Every thread on the system
Number of iterations Equals number of processes Equals total number of threads
TGID == PID? Always true for visited tasks True only for the main thread
Use case Process monitoring, security auditing Thread profiling, scheduler analysis
Header required <linux/sched.h> <linux/sched.h>, <linux/sched/signal.h>

10. Key Takeaways

Circular task list init_task entry point for_each_process = main threads only do_each_thread = all threads TGID = POSIX process ID PID = unique kernel thread ID mm == NULL → kernel thread Always use RCU read lock ps -LA shows tgid as PID, pid as LWP

Conclusion

In this lecture of our free Linux kernel development course, you learned how the Linux kernel organises all processes and threads into a single circular doubly-linked list rooted at init_task. You saw how to use for_each_process() to visit main threads and do_each_thread() / while_each_thread() to visit every thread. You now understand why the kernel has both a pid and a tgid inside every task_struct — the answer goes back to the need to be POSIX-compliant while still giving every thread a unique internal identifier.

You can verify all of this directly from user space using ps -LA, where the PID column shows tgid and the LWP column shows the kernel’s internal pid for each thread. Write the sample modules, load them, read the dmesg output, and compare it against ps -LA — everything will match perfectly.

In the next lecture we will go deeper into task_struct itself and look at how the kernel manages process state transitions. Keep building, keep exploring!

FAQ

Q1: Why does every thread have its own task_struct in Linux?

The Linux kernel was originally designed around processes. When threading support was added, the simplest approach that reused existing scheduler code was to treat every thread as a schedulable unit with its own task_struct. This is efficient because the scheduler does not need a separate code path for threads versus processes — they are all just tasks.

Q2: What is init_task on a modern system?

init_task is the task_struct for the process with PID 1. On modern systems that use systemd, this is the systemd process. It is the ancestor of all user-space processes and is the convenient starting point the kernel macros use to walk the circular task list.

Q3: Can I use find_task_by_vpid() to get a specific process instead of iterating the whole list?

Yes. If you already know the PID you want, find_task_by_vpid(pid) (with RCU lock held) is much more efficient than iterating the full list. Use full-list iteration only when you genuinely need to process every task.

Q4: Is the task list the same as the run queue?

No. The task list (the tasks list head inside task_struct) links all tasks regardless of their state — running, sleeping, stopped, zombie. The run queue links only tasks that are ready to run. The macros in this lecture walk the global task list, not the run queue.

Q5: Why does ps show PID but actually display TGID?

This is the POSIX compatibility layer at work. User space expects that “PID” identifies a process, not an individual thread. So when user space calls getpid() or reads /proc/self/status, the kernel returns tgid rather than the thread’s own pid. The internal pid is exposed as LWP (lightweight process ID) when you use ps -LA or look in /proc/<pid>/task/<tid>/.

Q6: What happens if I do not use rcu_read_lock() when walking the task list?

On a uniprocessor system nothing may appear to go wrong immediately, but on any SMP system (multicore) you risk a use-after-free bug: another CPU could free a task_struct while your code is still reading it. The result is either a kernel panic (oops) or silent memory corruption. Always use RCU protection.

Q7: How can I find threads of a specific process only?

Use the while_each_thread(g, t) loop starting from the task_struct of the process you care about. Alternatively, look under /proc/<pid>/task/ in user space — each subdirectory there represents one thread of that process.

Q8: What is the difference between a kernel thread and a user-space thread?

A kernel thread runs entirely in kernel space and has no user-space virtual memory mapping (task->mm == NULL). It is created by the kernel using kthread_create() and handles internal kernel work like memory management (kswapd), I/O completion (kworker), and so on. A user-space thread has its own virtual memory mapping and runs user-space code.

Q9: Is this free Linux kernel development course suitable for interview preparation?

Absolutely. Questions about process representation, task_struct layout, PID vs TGID, and task list iteration are very common in embedded Linux and kernel driver interviews at companies like Qualcomm, Texas Instruments, STMicroelectronics, and others. Understanding these fundamentals deeply will give you a real edge.

Q10: Where can I find the authoritative documentation on these kernel APIs?

The best sources are: the Linux kernel source itself (especially include/linux/sched.h and kernel/sched/core.c), the kernel documentation at kernel.org/doc, and the POSIX pthreads specification. Reading the actual kernel headers alongside this tutorial is the fastest way to build real expertise.

Further Reading & References

Continue Your Free Linux Kernel Development Journey

EmbeddedPathashala offers a completely free Linux kernel programming course, free Linux device drivers course, and free embedded systems course — no signup required.

← Previous Lecture Next Lecture →

Leave a Reply

Your email address will not be published. Required fields are marked *