« Previous Lecture | Next Lecture »
If you are following this free Linux kernel development course, you already know that a Linux kernel workqueue lets a driver defer work to run later in a safe, sleepable, process context. In this lecture of our free Linux device drivers course we go one level deeper into the Linux kernel workqueue API: why the old create_workqueue() family of calls is deprecated, how to migrate every one of them to the modern concurrency-managed workqueue (cmwq) calls, and exactly how INIT_WORK() and schedule_work() behave under the hood on a current 6.x kernel.
What You Will Learn
- Why the legacy
create_workqueue()style API was deprecated and what replaced it - The exact mapping from every legacy workqueue call to its modern
alloc_workqueue()/alloc_ordered_workqueue()equivalent - How to register a deferred task with
INIT_WORK() - How
schedule_work()actually decides whether to queue your work, and why it is safe to call from interrupt context - A clean pattern for managing several work items in one driver
- A first look at delayed work scheduling with
schedule_delayed_work()
Prerequisites
- Comfortable building and loading kernel modules on a 6.x kernel
- Basic C pointers and structures
- Some familiarity with kernel threads (see the earlier lecture in this series)
- A Linux VM or test machine you don’t mind rebooting
Why the Legacy Workqueue API Was Deprecated
Before 2010, every call to create_workqueue() spun up a brand-new set of dedicated kernel threads — one per CPU. If ten drivers each created their own workqueue, the system ended up with dozens of mostly-idle kernel threads just sitting around consuming memory and scheduler bookkeeping. The concurrency-managed workqueue (cmwq) rewrite fixed this by sharing a common pool of worker threads across the whole system, and handing out threads on demand instead of pinning them permanently to one workqueue.
To avoid breaking every driver in the tree overnight, the kernel developers kept the old function names alive as thin wrapper macros around the new alloc_workqueue() call. That is the situation you still find in the kernel source today: create_workqueue() compiles and works, but it is marked deprecated and new code should never use it.
Legacy API to Modern cmwq API Mapping
Here is the direct replacement for each legacy call, updated for how the mapping actually looks on a modern kernel:
| Legacy Call (deprecated) | Modern cmwq Replacement |
|---|---|
create_workqueue(name) |
alloc_workqueue(name, WQ_MEM_RECLAIM, max_active) |
create_singlethread_workqueue(name) |
alloc_ordered_workqueue(name, WQ_MEM_RECLAIM) |
create_freezable_workqueue(name) |
alloc_workqueue(name, WQ_FREEZABLE | WQ_UNBOUND | WQ_MEM_RECLAIM, max_active) |
Why alloc_ordered_workqueue() Replaced the Old “max_active = 1” Trick
Older documentation suggested that combining WQ_UNBOUND with a max_active of 1 gave you strict, one-at-a-time, in-order execution. On a current kernel that combination is no longer guaranteed to behave that way. If your driver genuinely needs work items to run one after another in the order they were queued, use the dedicated alloc_ordered_workqueue(name, flags) call instead — it exists specifically for this case and makes the intent explicit in the code.
INIT_WORK() — Registering Your Deferred Task
Before anything can run on a workqueue, the kernel needs to know two things: which struct work_struct represents your job, and which function should be called when that job is picked up by a worker thread. That registration is done with the INIT_WORK() macro.
#include <linux/workqueue.h>
INIT_WORK(struct work_struct *work, work_func_t func);
/* the callback signature every work function must follow */
typedef void (*work_func_t)(struct work_struct *work);
The callback only ever receives the work_struct pointer, not your driver’s private data directly. The standard pattern is to embed the work_struct inside your own driver structure and use container_of() inside the callback to get back to your data — the same pattern we already used for kernel timers and kthreads earlier in this series.
schedule_work() — Queuing Work on the Kernel-Global Workqueue
INIT_WORK() only registers the work item; it does not run anything yet. To actually hand the job to a worker thread, call schedule_work():
bool schedule_work(struct work_struct *work);
The return value tells you whether the item was newly queued (true) or was already sitting on the workqueue waiting to run (false) — calling it twice in a row on the same item never creates two pending executions. Because it only links the item into a queue and wakes a worker, schedule_work() does no sleeping of its own, which makes it safe to call from interrupt handlers, spinlock-protected sections, and other atomic contexts. The callback itself, however, always runs later in ordinary process context, where it is free to sleep, allocate memory, or take mutexes.
Original Example: Deferring Work From a Simulated Event
The following original module shows the complete lifecycle: registering a work item at module load, triggering it from a simulated event, and cleaning it up on exit.
#include <linux/module.h>
#include <linux/workqueue.h>
#include <linux/timer.h>
struct epwq_demo {
struct work_struct work;
struct timer_list poke_timer;
int event_count;
};
static struct epwq_demo demo;
static void epwq_work_handler(struct work_struct *work)
{
struct epwq_demo *d = container_of(work, struct epwq_demo, work);
/* Safe to sleep here: we are in process context, not the timer's
* interrupt context. */
pr_info("epwq: handling deferred event #%d\n", d->event_count);
}
static void epwq_timer_cb(struct timer_list *t)
{
struct epwq_demo *d = from_timer(d, t, poke_timer);
d->event_count++;
/* Safe to call from timer (softirq) context. */
schedule_work(&d->work);
mod_timer(&d->poke_timer, jiffies + HZ * 2);
}
static int __init epwq_init(void)
{
INIT_WORK(&demo.work, epwq_work_handler);
timer_setup(&demo.poke_timer, epwq_timer_cb, 0);
mod_timer(&demo.poke_timer, jiffies + HZ * 2);
pr_info("epwq: demo module loaded\n");
return 0;
}
static void __exit epwq_exit(void)
{
timer_delete_sync(&demo.poke_timer);
cancel_work_sync(&demo.work);
pr_info("epwq: demo module unloaded\n");
}
module_init(epwq_init);
module_exit(epwq_exit);
MODULE_LICENSE("GPL");
Note the use of timer_delete_sync() during cleanup rather than the older del_timer_sync() name — consistent with the kernel 6.2+ timer API we covered in the kernel timers lecture. We also call cancel_work_sync() before returning from the exit path, which both cancels a not-yet-started item and waits for an already-running one to finish, so the module can never be unloaded while its callback is mid-execution.
Handling Multiple Work Items in One Driver
Real drivers often need more than one deferred job — for example, one work item for handling incoming data and a separate one for periodic housekeeping. The clean way to do this is simply to call INIT_WORK() once per job, each with its own work_struct and its own callback function:
struct epwq_multi {
struct work_struct data_work;
struct work_struct housekeeping_work;
};
static void epwq_data_handler(struct work_struct *work) { /* ... */ }
static void epwq_housekeeping_handler(struct work_struct *work) { /* ... */ }
static void epwq_multi_init(struct epwq_multi *m)
{
INIT_WORK(&m->data_work, epwq_data_handler);
INIT_WORK(&m->housekeeping_work, epwq_housekeeping_handler);
}
When a driver needs many similar jobs (say, one per hardware queue), the same pattern is usually repeated inside an array of structures rather than written out by hand for each one.
A First Look at schedule_delayed_work()
Sometimes you don’t want your work to run immediately — you want it to run after a short delay. That’s what schedule_delayed_work() is for:
bool schedule_delayed_work(struct delayed_work *dwork, unsigned long delay);
It takes a struct delayed_work (which internally wraps a normal work_struct together with a timer) and a delay expressed in jiffies. We will build a complete original example around this call, along with its _on() CPU-targeted variants, in the next lecture of this series.
Common Mistakes and Troubleshooting
| Mistake | What Happens / Fix |
|---|---|
| Freeing the structure containing a work_struct before the callback runs | Use-after-free crash; always call cancel_work_sync() before freeing |
Calling schedule_work() before INIT_WORK() | Undefined behaviour — the work item’s function pointer is garbage; always initialize first |
Assuming WQ_UNBOUND + max_active=1 guarantees strict ordering | No longer guaranteed on modern kernels; use alloc_ordered_workqueue() instead |
| Blocking for a long time inside a work callback on the shared kernel-global workqueue | Starves other users of that shared pool; allocate a dedicated workqueue for long-running work |
Best Practices
- Prefer the kernel-global workqueue for short, quick jobs; allocate a dedicated one only when you truly need isolation or ordering guarantees
- Always pair
INIT_WORK()with a matchingcancel_work_sync()in your cleanup path - Use
alloc_ordered_workqueue()when you need in-order execution — don’t rely on flag combinations that no longer guarantee it - Keep work callbacks short where possible; split unrelated jobs into separate work items instead of one large callback
Performance Considerations
Sharing worker threads across the kernel-global workqueue is efficient precisely because idle threads aren’t wasted per driver. But that efficiency is a shared resource: a callback that blocks for a long time on the shared pool can delay unrelated drivers’ work items behind it. If your job routinely takes a long time, allocate your own workqueue rather than leaning on the global one.
Security Considerations
Work callbacks run with kernel privileges in process context, so treat any data captured by a work item the same way you would treat data crossing any other trust boundary in the kernel — validate it before use, and make sure the structure holding your work_struct cannot be freed or reused while a callback might still be pending.
Summary / Key Takeaways
create_workqueue()and friends are deprecated wrappers aroundalloc_workqueue()- Use the mapping table above to migrate legacy calls; prefer
alloc_ordered_workqueue()for strict ordering INIT_WORK()registers a callback;schedule_work()queues it and is safe from atomic/interrupt context- Callbacks always run later in sleepable process context
schedule_delayed_work()adds a timer-based delay on top of the same mechanism
Conclusion
The Linux kernel workqueue API looks small on the surface — a handful of function calls — but understanding exactly when work runs, in what context, and how the legacy and modern calls relate to each other is what separates a driver that merely compiles from one that behaves correctly under load. With INIT_WORK(), schedule_work(), and the legacy-to-modern mapping covered here, you’re ready to move on to delayed work scheduling in the next lecture of this free Linux kernel development course.
FAQ
Q1. Is create_workqueue() completely removed from the kernel?
No. As of current mainline kernels it still exists as a deprecated wrapper, but new code should not use it.
Q2. Can I call schedule_work() from an interrupt handler?
Yes. It only queues the item and wakes a worker; it does not sleep, so it is safe from interrupt and other atomic contexts.
Q3. What replaces create_singlethread_workqueue()?
Use alloc_ordered_workqueue(name, WQ_MEM_RECLAIM), which gives you a dedicated, strictly-ordered workqueue.
Q4. Does calling schedule_work() twice run my callback twice?
No. If the item is already queued, the second call returns false and the item is left in its existing position.
Q5. What is the difference between cancel_work_sync() and simply not calling it?
cancel_work_sync() cancels a pending item or waits for a running one to finish before returning, which prevents use-after-free when you free the containing structure right afterward.
Q6. Why not just always use WQ_UNBOUND with max_active=1 for ordering?
That combination no longer guarantees strict ordering on modern kernels; use alloc_ordered_workqueue() instead.
Q7. Where does schedule_delayed_work() fit in?
It layers a timer-based delay on top of the same work-queuing mechanism, letting you defer execution by a chosen number of jiffies. We cover it fully in the next lecture.
This lecture is part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.
Explore More Lectures
2 Comments