What is Linux Kernel Thread Lifecycle: Reference Counting, Signals & Safe Sleep-Wake Patterns-Linux Device Driver Training

Linux Kernel Thread Lifecycle: Reference Counting, Signals & Safe Sleep-Wake Patterns
Part of the Free Linux Kernel Development Course — Kernel Timers, Threads & Workqueues Series
Kernel Version
6.x (modern)
Level
Beginner → Intermediate
Reading Time
~14 minutes

← Previous Lecture    Next Lecture →

Understanding the Linux kernel thread lifecycle is what separates a driver author who merely gets a kthread running from one who can shut it down cleanly, without races, leaks, or a hung rmmod. In this lecture — part of our free Linux kernel development course — we go beyond simply creating a kernel thread and look at three things every module author eventually needs: how the kernel keeps a thread’s task_struct alive with reference counting, how signals behave differently for kernel threads than for user-space processes, and how to make a kthread sleep and wake up without racing against kthread_stop().

Everything here targets a current 6.x kernel, uses only original example code, and skips the diagrams-as-screenshots approach in favor of simple inline HTML boxes you can read on any device.

Topics covered in this free Linux kernel development course lecture:
kernel thread lifecycle get_task_struct / put_task_struct allow_signal kthread_should_stop wait_event_interruptible kthread_stop internals free linux device drivers course

What You Will Learn

  • Why a kernel thread’s task_struct needs its own reference count, and how get_task_struct() / put_task_struct() protect it.
  • Why kernel threads start with every signal blocked, and how allow_signal() changes that.
  • The difference between a fragile hand-rolled sleep loop and a race-free wait_event_interruptible() pattern.
  • What kthread_stop() actually does under the hood (it is not a signal — a very common misconception).
  • A complete, original example kernel thread module you can build and load on a modern distribution.

Prerequisites

  • Comfortable building and loading a basic out-of-tree kernel module (insmod / rmmod).
  • Familiarity with creating a kernel thread using kthread_run() — covered in the previous lecture of this series.
  • A Linux distribution running kernel 6.x with matching kernel headers installed for module builds.

A Quick Recap: Starting a Kernel Thread

In the previous lecture we created a kernel thread and set it running immediately using the convenience helper that both creates the thread and marks it schedulable in one step. That gets a thread born and running — but a running thread is not the same as a well-managed thread. A well-managed thread in the Linux kernel thread lifecycle needs three more things: a stable reference to its own descriptor, a defined signal policy, and a wake-up path that cannot race with shutdown. We tackle each one below.

Reference Counting a Kernel Thread

Every kernel thread has a task_struct — the same kind of structure the scheduler uses for ordinary processes. The catch: that structure can be freed the moment the thread exits, unless something is holding a reference to it. If your driver keeps a raw pointer to a kthread’s task_struct around (say, to call kthread_stop() on it later from a different code path), you need to guarantee it is not freed out from under you in the meantime.

This is exactly what a small pair of reference-counting helpers are for: one increments the structure’s reference count right after the thread is created, and its counterpart decrements it once your driver no longer needs the pointer. Think of it as a lease — as long as your module holds the lease, the kernel will not tear down the thread’s descriptor, even after the thread function itself has returned.

Why This Matters in Practice

Without holding a reference, a sequence like “thread exits on its own → driver later dereferences the stale pointer to check its state” becomes a use-after-free bug — one of the more painful classes of kernel crashes to debug, because the symptom often shows up far away from the real cause.

Why Kernel Threads Ignore Signals by Default

User-space signal semantics assume a process that can be interrupted, stopped, or killed by another process. Kernel threads are different: they run entirely in kernel address space and have no controlling terminal, so by default every signal delivered to a kthread is either dropped or converted into a forced termination. This is a safety measure — you do not want an ordinary user accidentally killing a thread that a driver depends on for correctness.

When a specific signal should be meaningful to your thread — for example, letting an administrator ask it to stop gracefully with SIGTERM — you opt in for that one signal explicitly. Once opted in, the thread can check whether a signal is pending and react to it inside its own loop, on its own terms, rather than being torn down abruptly.

Signals Are Optional, Not the Recommended Shutdown Path

It is worth stating plainly: for module unload today, the standard shutdown mechanism is kthread_stop(), not signals. Signal handling in a kthread is still useful for cases where you genuinely want an operator or another kernel subsystem to nudge the thread (for example, to cancel a long wait), but it is a secondary tool, not the primary lifecycle control.

The Correct Sleep-Wake Loop for a Kernel Thread

A kernel thread that only has work to do occasionally should not spin in a busy loop burning CPU. It should sleep, and wake only when there is something to do — either real work, or a stop request. The naive approach sets the task’s state to an interruptible sleep state and calls the scheduler directly in a loop. This works most of the time, but it has a subtle race: if the wake-up condition becomes true in the tiny window between checking it and actually going to sleep, the thread can miss the wake-up and sleep far longer than intended, sometimes appearing to hang.

The safer, modern pattern is to sleep on a wait queue with a condition, using a helper that atomically checks the condition and puts the thread to sleep if it is still false, and automatically stops sleeping the moment the condition is signaled. This closes the race window entirely and is the pattern used throughout current kernel code.

Kernel Thread Lifecycle — Visual Flow
1. Create kthread_run() starts the thread and marks it schedulable
2. Reference get_task_struct() takes a lease on the task_struct
3. Sleep wait_event_interruptible() waits for work OR a stop request
4. Wake Work arrives, a signal arrives, or kthread_stop() is called
5. Check kthread_should_stop() decides: loop again, or exit
6. Exit & release Thread returns, put_task_struct() releases the lease

How kthread_stop() Really Works

A persistent myth is that kthread_stop() sends a signal to the target thread. It does not. What it actually does is set an internal flag so that kthread_should_stop() starts returning true for that thread, then it wakes the thread up (so a sleeping thread does not stay asleep forever) and blocks the caller until the thread function returns. This is precisely why any wait condition your thread sleeps on must include kthread_should_stop() — if it does not, kthread_stop() will wake the thread, but the thread has no way to notice it should exit, and the caller of kthread_stop() can block indefinitely.

This single detail explains the majority of “my module hangs on rmmod” reports from developers new to kernel thread programming.

Hands-On Example: A Signal-Aware Kernel Thread

The following is an original demonstration module, written for this lecture, that ties together everything above: reference counting, opt-in signal handling, and a race-free sleep-wake loop. It compiles cleanly against a current 6.x kernel tree.

#include <linux/init.h>
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/sched.h>
#include <linux/signal.h>
#include <linux/sched/signal.h>
#include <linux/wait.h>

static struct task_struct *demo_thread;
static DECLARE_WAIT_QUEUE_HEAD(demo_wq);
static int work_ready;

static int demo_thread_fn(void *arg)
{
    allow_signal(SIGTERM);

    while (!kthread_should_stop()) {

        wait_event_interruptible(demo_wq,
                kthread_should_stop() || work_ready);

        if (kthread_should_stop())
            break;

        if (work_ready) {
            pr_info("lifecycle_demo: doing queued work\n");
            work_ready = 0;
        }

        if (signal_pending(current)) {
            pr_info("lifecycle_demo: SIGTERM received, exiting\n");
            flush_signals(current);
            break;
        }
    }

    pr_info("lifecycle_demo: thread function returning\n");
    return 0;
}

static int __init lifecycle_demo_init(void)
{
    demo_thread = kthread_run(demo_thread_fn, NULL, "lifecycle_demo");
    if (IS_ERR(demo_thread))
        return PTR_ERR(demo_thread);

    get_task_struct(demo_thread);
    pr_info("lifecycle_demo: module loaded, thread started\n");
    return 0;
}

static void __exit lifecycle_demo_exit(void)
{
    kthread_stop(demo_thread);
    put_task_struct(demo_thread);
    pr_info("lifecycle_demo: module removed\n");
}

module_init(lifecycle_demo_init);
module_exit(lifecycle_demo_exit);

MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Original demo: kernel thread lifecycle with refcounting, signals, and safe sleep");

Notice that work_ready is included in the wait condition alongside kthread_should_stop(). Any other subsystem in the driver that sets work_ready = 1 should call wake_up_interruptible(&demo_wq) right after — that is what actually pulls the thread out of sleep to look at the new condition.

Common Mistakes and Troubleshooting

Mistake Consequence Fix
Sleeping without checking kthread_should_stop() in the wait condition rmmod hangs forever Always OR kthread_should_stop() into the wait_event condition
Assuming kthread_stop() sends a signal Thread never notices the stop request Remember it only flips a flag and wakes the thread
Calling allow_signal() for every signal “just in case” Unpredictable behaviour from unrelated user signals Opt in only to the specific signal you intend to handle
Dereferencing a task_struct pointer after the thread exited Use-after-free crash Hold a reference with get_task_struct() until you are done with the pointer

Best Practices for Kernel Thread Lifecycle Management

  • Treat kthread_stop() as the primary shutdown path for any kernel thread lifecycle you design; use signals only for optional, secondary nudges.
  • Always include kthread_should_stop() in whatever condition your thread sleeps on.
  • Pair every get_task_struct() with exactly one matching put_task_struct().
  • Keep the set of signals a thread allows as small as possible — ideally just one.

Performance Considerations

A wait-queue-based sleep costs essentially nothing while idle — the thread consumes no CPU until it is woken. Compare that to a polling loop with a short delay, which wakes on a fixed schedule regardless of whether there is work, wasting cycles and hurting power efficiency on battery-powered and embedded targets alike.

Security Considerations

Because kernel threads run with full kernel privilege, opting a thread in to accept a signal from user space is effectively giving unprivileged callers a way to influence kernel code execution timing. Only allow signals your thread is designed to handle safely, and never allow a signal whose default disposition (like a forced kill) would leave shared kernel resources in an inconsistent state.

Real-World Use Cases

  • A storage driver’s background flush thread that sleeps until dirty data accumulates, then wakes to write it out.
  • A network driver’s monitoring thread that periodically polls link state but can be told to stop instantly at module unload.
  • A watchdog-style thread that an administrator can nudge with a specific signal to force an immediate health check.

Summary / Key Takeaways

  • The Linux kernel thread lifecycle has three pillars beyond simple creation: reference counting, an explicit signal policy, and a race-free sleep-wake loop.
  • get_task_struct() / put_task_struct() keep a thread’s descriptor alive exactly as long as it is needed.
  • Signals are blocked by default for kernel threads and must be explicitly allowed with allow_signal().
  • kthread_stop() does not send a signal — it sets a flag and wakes the thread, so your sleep condition must always include kthread_should_stop().

Conclusion

Getting a kernel thread running is the easy part of kernel thread programming; managing its lifecycle correctly is what makes a driver production-grade. Once you internalize that kthread_stop() is a flag-and-wake mechanism rather than a signal, and that every sleep condition must account for it, the rest of the pattern — reference counting plus opt-in signal handling — falls into place naturally. This lecture continues our free Linux kernel development course, and the next lecture in this series moves on from kernel threads into workqueues, where the kernel manages a pool of worker threads for you.

Frequently Asked Questions

Does kthread_stop() send a signal to the kernel thread?

No. It sets an internal flag so kthread_should_stop() returns true, wakes the thread if it was sleeping, and then waits for the thread function to return.

Why are signals blocked for kernel threads by default?

Kernel threads have no controlling terminal and are relied upon for kernel-internal correctness, so the kernel avoids letting arbitrary user-space signals disrupt them unless a thread explicitly opts in.

What happens if I forget to call get_task_struct() before storing a task_struct pointer?

The thread’s descriptor can be freed once the thread exits, and any later use of that stored pointer becomes a use-after-free bug.

Is wait_event_interruptible() always better than a manual set_current_state() loop?

For almost all cases, yes — it avoids the race window where a wake-up can be missed between checking a condition and going to sleep.

Can I allow more than one signal in a kernel thread?

Technically yes, by calling allow_signal() for each one, but it is best practice to keep the set as small as possible to reduce unpredictable interactions.

Do I need put_task_struct() if I never call get_task_struct()?

No — only call put_task_struct() to release a reference you explicitly took with get_task_struct().

Does this pattern work the same way on older kernels?

The core APIs discussed here have been stable for a long time; the concepts and this example target current 6.x kernels without relying on any deprecated wrapper names.

Continue the Free Linux Kernel Development Course

This lecture is part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course series.

Previous Lecture Next Lecture