← Previous Lecture | Next Lecture →
Intermediate
6.x
~14 min
Getting kernel thread synchronization right is one of the trickiest parts of
writing a Linux device driver that hands work off from an ioctl() call to a
background kernel thread. It sounds simple: a user-space app calls ioctl(),
the driver wakes a kernel thread to do the heavy lifting, and the app comes back later to
collect the result. The trouble starts the moment more than one thread, or more than one
process, calls that same ioctl() at the same time. Without proper
synchronization, two callers can trample the same shared buffer, read half-finished data,
or wake the worker thread twice in a row. In this lecture we design a small, original
driver example that solves this cleanly using struct completion and a mutex,
the way it should be done on a modern kernel 6.x system.
What You Will Learn
- Why handing work from
ioctl()to a kernel thread is race-prone by default - The manager-worker pattern: who does what, and in which process context
- Why busy-polling for “is the work done yet” wastes CPU cycles and is fragile
- How
struct completiongives you race-free, sleep-based signaling - How a mutex protects the shared buffer from concurrent
ioctl()callers - A complete, original example driver you can build and test on kernel 6.x
Prerequisites
- Comfortable writing a basic misc character device driver with an
ioctl()handler - Familiarity with creating a kernel thread using
kthread_create()/kthread_stop() - Basic understanding of process context vs. kernel thread context
- A Linux machine running kernel 6.x with headers installed, for building test modules
The Problem: Why Handing Off Work to a Kernel Thread Can Race
Imagine a driver where ioctl() copies a message from user space, and a kernel
thread transforms it in the background (say, compresses it, checksums it, or encrypts it —
the exact transform doesn’t matter). The natural design looks like this: the ioctl()
handler copies in the data, sets a flag, and calls wake_up_process() on the
kernel thread. The kernel thread wakes up, does the work, and goes back to sleep.
The problem is that ioctl() runs in the process context of whichever user-space
thread called it, while the kernel thread runs in its own, completely independent context.
If a second user-space thread calls ioctl() before the first request has been
picked up by the worker, both requests can stomp on the same shared buffer. Worse, if the
caller tries to read the result back immediately after waking the thread, there’s no
guarantee the thread has even started running yet — this is a classic race condition, and it
gets worse on multi-core systems where the ioctl() caller and the kernel thread
can genuinely execute at the exact same instant on different CPUs.
ioctl(TRANSFORM)
no lock, no signal
ioctl(TRANSFORM)
Both threads touch the same buffer with no ordering guarantee — a textbook race.
The Manager-Worker Pattern, Done Properly
The fix isn’t to abandon the manager-worker split — it’s a genuinely useful pattern, since it
lets time-consuming work run asynchronously in its own kernel thread instead of blocking
whichever process context happens to call ioctl(). The fix is to add two
ingredients that are often skipped in quick-and-dirty examples:
- A mutex around the shared buffer and state, so only one
ioctl()caller can submit or retrieve a job at a time. - A proper completion signal, so the caller can sleep until the kernel thread has actually finished, instead of guessing or spinning.
Older driver examples (including some written against Linux kernel versions from several
years back) reach for a manual polling loop as a stand-in for real signaling: the
ioctl() handler wakes the kernel thread, then spins in a loop checking a plain
variable until it changes, calling schedule() or a short udelay()
on every iteration. It works, technically, but it burns CPU time for no reason and still isn’t
fully race-free unless every access to that variable is atomic. On a modern kernel there is a
purpose-built primitive for exactly this job.
struct completion: The Modern, Race-Free Signal
The kernel’s struct completion API is designed precisely for “one context tells
another context that a specific piece of work is done.” Instead of a caller spinning in a
loop, it calls wait_for_completion_interruptible() and is put to sleep by the
scheduler — using zero CPU while it waits. When the kernel thread finishes its work, it calls
complete(), which wakes the waiter up cleanly. No polling, no wasted cycles, no
manual atomic flag needed.
#include <linux/completion.h>
struct my_drv_ctx {
struct mutex lock; /* serializes ioctl() callers */
struct completion work_done; /* signaled when kthread finishes */
struct task_struct *worker;
char buf[256];
bool job_pending;
};
/* Called once at driver init */
static void my_drv_ctx_init(struct my_drv_ctx *ctx)
{
mutex_init(&ctx->lock);
init_completion(&ctx->work_done);
ctx->job_pending = false;
}
Putting It Together: An Original Example Driver
Below is a simplified, original driver skeleton (not taken from any published source) that
shows the full hand-off: ioctl() takes the mutex, copies in the job, wakes the
kernel thread, waits on the completion, then releases the mutex. The kernel thread does the
actual transform and signals completion when it’s done.
/* ioctl() handler - runs in the caller's process context */
static long my_drv_ioctl(struct file *filp, unsigned int cmd, unsigned long arg)
{
struct my_drv_ctx *ctx = filp->private_data;
long ret = 0;
if (cmd != MY_DRV_IOC_TRANSFORM)
return -ENOTTY;
/* Only one caller can submit/collect a job at a time */
if (mutex_lock_interruptible(&ctx->lock))
return -ERESTARTSYS;
if (copy_from_user(ctx->buf, (char __user *)arg, sizeof(ctx->buf))) {
ret = -EFAULT;
goto out_unlock;
}
reinit_completion(&ctx->work_done);
ctx->job_pending = true;
wake_up_process(ctx->worker);
/* Sleep here - zero CPU burned, no polling loop */
if (wait_for_completion_interruptible(&ctx->work_done)) {
ret = -ERESTARTSYS;
goto out_unlock;
}
if (copy_to_user((char __user *)arg, ctx->buf, sizeof(ctx->buf)))
ret = -EFAULT;
out_unlock:
mutex_unlock(&ctx->lock);
return ret;
}
/* Kernel thread - runs in its own independent context */
static int my_drv_worker(void *arg)
{
struct my_drv_ctx *ctx = arg;
while (!kthread_should_stop()) {
set_current_state(TASK_INTERRUPTIBLE);
if (!ctx->job_pending)
schedule();
set_current_state(TASK_RUNNING);
if (kthread_should_stop())
break;
if (ctx->job_pending) {
my_transform_inplace(ctx->buf, sizeof(ctx->buf));
ctx->job_pending = false;
complete(&ctx->work_done);
}
}
return 0;
}
Notice what each piece is doing: the mutex makes sure a second caller simply waits its turn instead of racing the first one for the buffer. The completion means the caller is asleep, not spinning, while the kernel thread works — and it’s woken up exactly once, exactly when the work is genuinely finished, with no ambiguity about ordering.
Why This Matters for Real Drivers
This pattern — mutex plus completion around a manager-worker hand-off — shows up constantly in real Linux drivers, well beyond toy examples: firmware-loading drivers wait for a kernel thread or workqueue to finish flashing firmware, crypto drivers wait for a hardware accelerator’s completion interrupt, and USB drivers wait for a submitted URB to finish. Once you understand this small example, you’ll recognize the same shape everywhere in the kernel source tree.
Frequently Asked Questions
Q1. Why not just use a global flag and check it in a loop?
A plain flag checked in a busy loop burns CPU the entire time it waits, and reading/writing
it without a memory barrier or atomic type is not guaranteed to be race-free on all
architectures. struct completion handles both problems for you.
Q2. Is a mutex always the right lock here?
For code that may sleep (like waiting on a completion), yes — a mutex is appropriate.
A spinlock must never be held across a sleeping call such as
wait_for_completion_interruptible().
Q3. What happens if the kernel thread is stopped mid-job?
A well-behaved driver should call complete() (or a “job aborted” variant)
before exiting the loop on kthread_should_stop(), so no caller is left
waiting forever.
Q4. Can I use this pattern with workqueues instead of a kernel thread?
Yes — the same mutex-plus-completion idea works whether the background work runs in a
dedicated kernel thread or a workqueue; only the wake-up mechanism changes.
Q5. Does wait_for_completion_interruptible() ever time out?
Not on its own. If you need a deadline, use
wait_for_completion_interruptible_timeout() instead.
Q6. Why use wait_for_completion_interruptible() instead of the
non-interruptible version?
The interruptible variant lets a user-space caller be killed with a signal (like Ctrl-C)
while blocked in the driver, instead of being stuck uninterruptibly.
This lecture is part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.
Browse More Lectures Join the Community
2 Comments