This lecture is part of our free Linux kernel development course and continues our
free Linux device drivers course workqueue mini-series. If you have already gone through the
earlier workqueue lectures, you know how to schedule a work item and how to delay one. In this lecture we do
two things: first, we measure exactly how much delay there is between calling schedule_work()
and the work function actually starting, and second, we build a small original driver called
sed3 — a workqueue-based redesign of the mutex-and-completion driver from our earlier
lecture. All code shown here targets a modern Linux kernel (6.x) and has been written from scratch for this
free embedded systems course.
What You Will Learn
- Why
schedule_work()does not run your function instantly, and how to measure the actual delay - How to confirm, on a modern kernel, whether code is running in process context or interrupt context
- What kernel worker (kworker) threads are and how to identify the one servicing your work
- How to redesign a kthread-based driver as a workqueue-based driver (the sed3 pattern)
- A full original example driver combining a mutex, a workqueue, and a completion
Prerequisites
- Comfortable with
INIT_WORK()andschedule_work()from our earlier workqueue lecture - Basic familiarity with
struct mutexandstruct completion - A Linux system running kernel 6.x with headers installed, for building and testing kernel modules
Why Measure Workqueue Scheduling Latency?
A workqueue hands your function off to a kernel worker thread. That thread has to be scheduled by the CPU
scheduler just like any other thread on the system, so there is always a small gap between the moment you call
schedule_work() and the moment your function actually begins running. On a lightly loaded system
this gap is tiny — typically well under a millisecond. On a busy system, or one running realtime tasks at a
higher priority than kernel worker threads, this gap can grow noticeably. Knowing how to measure it yourself is
far more useful than memorizing a number from any book, because the value depends entirely on your own
hardware, kernel configuration, and system load.
Measuring the Gap Yourself
The idea is simple: take a timestamp right before scheduling the work, take a second timestamp as the very
first line inside the work function, and subtract. We use ktime_get_ns(), the standard modern
kernel API for a monotonic nanosecond timestamp.
#include <linux/workqueue.h>
#include <linux/ktime.h>
static u64 sched_ts;
static void latency_work_fn(struct work_struct *work)
{
u64 run_ts = ktime_get_ns();
pr_info("latency_demo: work started %llu ns after scheduling\n",
run_ts - sched_ts);
}
static DECLARE_WORK(latency_work, latency_work_fn);
static void fire_work(void)
{
sched_ts = ktime_get_ns();
schedule_work(&latency_work);
}
Load a module that calls fire_work() a few times and watch dmesg. You will see a
small nanosecond value each time — that is your real scheduling latency, on your own machine, today.
timestamp taken
usually microseconds
second timestamp taken
Confirming Process Context on a Modern Kernel
A common thing to verify is that your work function really does run in process context, unlike a timer
callback which runs in softirq context. Rather than relying on custom debug macros, the kernel gives you a
ready-made helper for exactly this: in_task(). It returns true when the current code is running
in normal process context and false when running in interrupt or softirq context.
static void latency_work_fn(struct work_struct *work)
{
pr_info("latency_demo: in_task() = %d (should be 1)\n", in_task());
}
Run this alongside a kernel timer callback (from our earlier timer lecture) printing the same check, and you will see the timer print 0 while the workqueue prints 1 — direct, live proof that workqueues execute in process context.
Which Thread Actually Ran Your Work?
By default, schedule_work() queues onto the kernel-global workqueue, which is serviced by a
pool of per-CPU kernel worker threads named kworker/N:x, where N is the CPU number.
You can print current->pid and current->comm from inside your work function to
see exactly which one handled your work, and then look it up with a normal ps command on your
own system.
pr_info("latency_demo: serviced by pid=%d comm=%s\n",
current->pid, current->comm);
The sed3 Mini Project: From Kthread to Workqueue
In an earlier lecture in this series we built a manager-worker driver using a dedicated kthread, a mutex to protect a shared buffer, and a completion to let the ioctl caller wait for the work to finish. That pattern is correct, but it also means your driver owns and manages a long-lived thread for what is really just occasional, short bursts of work. sed3 is the natural next step: keep the same mutex-and-completion synchronization, but replace the dedicated kthread with a workqueue, so the kernel’s own worker pool does the scheduling for you and no thread needs to be created or destroyed by your driver at all.
sed3 Design
- A device-private context structure holds the shared buffer, a
struct mutex, astruct completion, and astruct work_struct - An ioctl call copies data in, locks the mutex, resets the completion, and calls
schedule_work() - The work function does the actual transform, then calls
complete() - The ioctl call waits on
wait_for_completion()before unlocking the mutex and returning to user space
Original sed3 Driver Skeleton
#include <linux/module.h>
#include <linux/miscdevice.h>
#include <linux/uaccess.h>
#include <linux/workqueue.h>
#include <linux/mutex.h>
#include <linux/completion.h>
#define SED3_BUF_LEN 128
#define SED3_ENCRYPT _IOW('s', 1, int)
#define SED3_DECRYPT _IOW('s', 2, int)
struct sed3_ctx {
char data[SED3_BUF_LEN];
struct mutex lock;
struct completion done;
struct work_struct work;
int mode; /* 0 = encrypt, 1 = decrypt */
};
static struct sed3_ctx ctx;
static void sed3_transform(char *buf, size_t len)
{
size_t i;
for (i = 0; i < len; i++)
buf[i] ^= 0x5A; /* toy transform for demo purposes only */
}
static void sed3_work_fn(struct work_struct *work)
{
struct sed3_ctx *c = container_of(work, struct sed3_ctx, work);
sed3_transform(c->data, SED3_BUF_LEN);
pr_info("sed3: transform done by pid=%d (%s)\n",
current->pid, current->comm);
complete(&c->done);
}
static long sed3_ioctl(struct file *filp, unsigned int cmd, unsigned long arg)
{
mutex_lock(&ctx.lock);
ctx.mode = (cmd == SED3_DECRYPT) ? 1 : 0;
reinit_completion(&ctx.done);
schedule_work(&ctx.work);
wait_for_completion(&ctx.done);
mutex_unlock(&ctx.lock);
return 0;
}
static const struct file_operations sed3_fops = {
.owner = THIS_MODULE,
.unlocked_ioctl = sed3_ioctl,
};
static struct miscdevice sed3_miscdev = {
.minor = MISC_DYNAMIC_MINOR,
.name = "sed3_drv",
.fops = &sed3_fops,
};
static int __init sed3_init(void)
{
mutex_init(&ctx.lock);
init_completion(&ctx.done);
INIT_WORK(&ctx.work, sed3_work_fn);
return misc_register(&sed3_miscdev);
}
static void __exit sed3_exit(void)
{
misc_deregister(&sed3_miscdev);
}
module_init(sed3_init);
module_exit(sed3_exit);
MODULE_LICENSE("GPL");
Notice that there is no kthread_create(), no kthread_stop(), and no thread cleanup
logic anywhere in this driver. The kernel-global workqueue’s own worker pool is reused for our short burst of
work, which keeps this driver noticeably simpler than a hand-rolled kthread design while still giving user
space the exact same synchronous, blocking ioctl behavior.
sed2 (Kthread) vs sed3 (Workqueue)
| Aspect | sed2 (kthread) | sed3 (workqueue) |
|---|---|---|
| Thread lifecycle | Driver creates and stops its own thread | Kernel-global worker pool is reused, no thread to manage |
| Best suited for | Long-running, dedicated, continuously active work | Short, occasional bursts of work |
| Cleanup code needed | kthread_stop() plus a stop flag or kthread_should_stop() loop | None — nothing to stop |
| Synchronization used | Mutex + completion | Same mutex + completion pattern, unchanged |
Frequently Asked Questions
1. Is a few hundred microseconds of workqueue latency a problem?
For the vast majority of driver use cases, no. That latency only matters for hard realtime deadlines, where a dedicated high-priority kthread or an hrtimer-based approach, covered earlier in this course, is a better fit.
2. Can I force my work item onto a specific CPU?
Yes, using schedule_work_on() instead of schedule_work(), which takes an explicit
CPU number.
3. Does sed3 lose any correctness compared to sed2?
No. The mutex still serializes access to the shared buffer and the completion still guarantees the ioctl caller only returns after the transform has finished; only the execution vehicle changed.
4. Why use a toy XOR transform instead of a real cipher?
The goal of this driver is to teach the workqueue and synchronization pattern clearly. Plugging in the kernel’s crypto API is a natural follow-on exercise once this pattern is understood.
5. What happens if two ioctl calls arrive at the same time?
The second caller simply blocks on mutex_lock() until the first one finishes and unlocks,
keeping the buffer and completion state consistent.
6. Do I need reinit_completion() every time?
Yes. A struct completion stays “done” once completed, so it must be reinitialized before each
new round of work, or you can call wait_for_completion() immediately without ever blocking.

2 Comments