Workqueue Scheduling Latency and the sed3 Encrypt/Decrypt Driver-Free Linux Device Drivers Course

Previous Lecture Next Lecture
Workqueue Scheduling Latency and the sed3 Encrypt/Decrypt Driver
Free Linux Kernel Programming Course – Workqueues Part 5
Kernel 6.x Ready
Original Driver Code
Beginner Friendly
free linux kernel development course free linux device drivers course free embedded systems course linux kernel workqueue tutorial

This lecture is part of our free Linux kernel development course and continues our free Linux device drivers course workqueue mini-series. If you have already gone through the earlier workqueue lectures, you know how to schedule a work item and how to delay one. In this lecture we do two things: first, we measure exactly how much delay there is between calling schedule_work() and the work function actually starting, and second, we build a small original driver called sed3 — a workqueue-based redesign of the mutex-and-completion driver from our earlier lecture. All code shown here targets a modern Linux kernel (6.x) and has been written from scratch for this free embedded systems course.

What You Will Learn

  • Why schedule_work() does not run your function instantly, and how to measure the actual delay
  • How to confirm, on a modern kernel, whether code is running in process context or interrupt context
  • What kernel worker (kworker) threads are and how to identify the one servicing your work
  • How to redesign a kthread-based driver as a workqueue-based driver (the sed3 pattern)
  • A full original example driver combining a mutex, a workqueue, and a completion

Prerequisites

  • Comfortable with INIT_WORK() and schedule_work() from our earlier workqueue lecture
  • Basic familiarity with struct mutex and struct completion
  • A Linux system running kernel 6.x with headers installed, for building and testing kernel modules

Why Measure Workqueue Scheduling Latency?

A workqueue hands your function off to a kernel worker thread. That thread has to be scheduled by the CPU scheduler just like any other thread on the system, so there is always a small gap between the moment you call schedule_work() and the moment your function actually begins running. On a lightly loaded system this gap is tiny — typically well under a millisecond. On a busy system, or one running realtime tasks at a higher priority than kernel worker threads, this gap can grow noticeably. Knowing how to measure it yourself is far more useful than memorizing a number from any book, because the value depends entirely on your own hardware, kernel configuration, and system load.

Measuring the Gap Yourself

The idea is simple: take a timestamp right before scheduling the work, take a second timestamp as the very first line inside the work function, and subtract. We use ktime_get_ns(), the standard modern kernel API for a monotonic nanosecond timestamp.

#include <linux/workqueue.h>
#include <linux/ktime.h>

static u64 sched_ts;

static void latency_work_fn(struct work_struct *work)
{
    u64 run_ts = ktime_get_ns();

    pr_info("latency_demo: work started %llu ns after scheduling\n",
            run_ts - sched_ts);
}

static DECLARE_WORK(latency_work, latency_work_fn);

static void fire_work(void)
{
    sched_ts = ktime_get_ns();
    schedule_work(&latency_work);
}

Load a module that calls fire_work() a few times and watch dmesg. You will see a small nanosecond value each time — that is your real scheduling latency, on your own machine, today.

Workqueue Scheduling Latency Timeline
schedule_work()
timestamp taken
→
scheduling gap
usually microseconds
→
kworker runs work_fn
second timestamp taken

Confirming Process Context on a Modern Kernel

A common thing to verify is that your work function really does run in process context, unlike a timer callback which runs in softirq context. Rather than relying on custom debug macros, the kernel gives you a ready-made helper for exactly this: in_task(). It returns true when the current code is running in normal process context and false when running in interrupt or softirq context.

static void latency_work_fn(struct work_struct *work)
{
    pr_info("latency_demo: in_task() = %d (should be 1)\n", in_task());
}

Run this alongside a kernel timer callback (from our earlier timer lecture) printing the same check, and you will see the timer print 0 while the workqueue prints 1 — direct, live proof that workqueues execute in process context.

Which Thread Actually Ran Your Work?

By default, schedule_work() queues onto the kernel-global workqueue, which is serviced by a pool of per-CPU kernel worker threads named kworker/N:x, where N is the CPU number. You can print current->pid and current->comm from inside your work function to see exactly which one handled your work, and then look it up with a normal ps command on your own system.

pr_info("latency_demo: serviced by pid=%d comm=%s\n",
        current->pid, current->comm);

The sed3 Mini Project: From Kthread to Workqueue

In an earlier lecture in this series we built a manager-worker driver using a dedicated kthread, a mutex to protect a shared buffer, and a completion to let the ioctl caller wait for the work to finish. That pattern is correct, but it also means your driver owns and manages a long-lived thread for what is really just occasional, short bursts of work. sed3 is the natural next step: keep the same mutex-and-completion synchronization, but replace the dedicated kthread with a workqueue, so the kernel’s own worker pool does the scheduling for you and no thread needs to be created or destroyed by your driver at all.

sed3 Design

  • A device-private context structure holds the shared buffer, a struct mutex, a struct completion, and a struct work_struct
  • An ioctl call copies data in, locks the mutex, resets the completion, and calls schedule_work()
  • The work function does the actual transform, then calls complete()
  • The ioctl call waits on wait_for_completion() before unlocking the mutex and returning to user space
sed3 Architecture Flow
1. User space calls ioctl(ENCRYPT / DECRYPT)
2. Driver locks mutex, reinitializes completion, calls schedule_work()
3. kworker thread runs sed3_work_fn() and transforms the buffer
4. sed3_work_fn() calls complete()
5. ioctl handler wakes from wait_for_completion(), unlocks mutex, returns to user space

Original sed3 Driver Skeleton

#include <linux/module.h>
#include <linux/miscdevice.h>
#include <linux/uaccess.h>
#include <linux/workqueue.h>
#include <linux/mutex.h>
#include <linux/completion.h>

#define SED3_BUF_LEN   128
#define SED3_ENCRYPT   _IOW('s', 1, int)
#define SED3_DECRYPT   _IOW('s', 2, int)

struct sed3_ctx {
    char data[SED3_BUF_LEN];
    struct mutex lock;
    struct completion done;
    struct work_struct work;
    int mode;   /* 0 = encrypt, 1 = decrypt */
};

static struct sed3_ctx ctx;

static void sed3_transform(char *buf, size_t len)
{
    size_t i;

    for (i = 0; i < len; i++)
        buf[i] ^= 0x5A;   /* toy transform for demo purposes only */
}

static void sed3_work_fn(struct work_struct *work)
{
    struct sed3_ctx *c = container_of(work, struct sed3_ctx, work);

    sed3_transform(c->data, SED3_BUF_LEN);
    pr_info("sed3: transform done by pid=%d (%s)\n",
            current->pid, current->comm);

    complete(&c->done);
}

static long sed3_ioctl(struct file *filp, unsigned int cmd, unsigned long arg)
{
    mutex_lock(&ctx.lock);

    ctx.mode = (cmd == SED3_DECRYPT) ? 1 : 0;
    reinit_completion(&ctx.done);
    schedule_work(&ctx.work);
    wait_for_completion(&ctx.done);

    mutex_unlock(&ctx.lock);
    return 0;
}

static const struct file_operations sed3_fops = {
    .owner          = THIS_MODULE,
    .unlocked_ioctl = sed3_ioctl,
};

static struct miscdevice sed3_miscdev = {
    .minor = MISC_DYNAMIC_MINOR,
    .name  = "sed3_drv",
    .fops  = &sed3_fops,
};

static int __init sed3_init(void)
{
    mutex_init(&ctx.lock);
    init_completion(&ctx.done);
    INIT_WORK(&ctx.work, sed3_work_fn);

    return misc_register(&sed3_miscdev);
}

static void __exit sed3_exit(void)
{
    misc_deregister(&sed3_miscdev);
}

module_init(sed3_init);
module_exit(sed3_exit);
MODULE_LICENSE("GPL");

Notice that there is no kthread_create(), no kthread_stop(), and no thread cleanup logic anywhere in this driver. The kernel-global workqueue’s own worker pool is reused for our short burst of work, which keeps this driver noticeably simpler than a hand-rolled kthread design while still giving user space the exact same synchronous, blocking ioctl behavior.

sed2 (Kthread) vs sed3 (Workqueue)

Aspect sed2 (kthread) sed3 (workqueue)
Thread lifecycle Driver creates and stops its own thread Kernel-global worker pool is reused, no thread to manage
Best suited for Long-running, dedicated, continuously active work Short, occasional bursts of work
Cleanup code needed kthread_stop() plus a stop flag or kthread_should_stop() loop None — nothing to stop
Synchronization used Mutex + completion Same mutex + completion pattern, unchanged

Frequently Asked Questions

1. Is a few hundred microseconds of workqueue latency a problem?

For the vast majority of driver use cases, no. That latency only matters for hard realtime deadlines, where a dedicated high-priority kthread or an hrtimer-based approach, covered earlier in this course, is a better fit.

2. Can I force my work item onto a specific CPU?

Yes, using schedule_work_on() instead of schedule_work(), which takes an explicit CPU number.

3. Does sed3 lose any correctness compared to sed2?

No. The mutex still serializes access to the shared buffer and the completion still guarantees the ioctl caller only returns after the transform has finished; only the execution vehicle changed.

4. Why use a toy XOR transform instead of a real cipher?

The goal of this driver is to teach the workqueue and synchronization pattern clearly. Plugging in the kernel’s crypto API is a natural follow-on exercise once this pattern is understood.

5. What happens if two ioctl calls arrive at the same time?

The second caller simply blocks on mutex_lock() until the first one finishes and unlocks, keeping the buffer and completion state consistent.

6. Do I need reinit_completion() every time?

Yes. A struct completion stays “done” once completed, so it must be reinitialized before each new round of work, or you can call wait_for_completion() immediately without ever blocking.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *