« Previous Lecture | Next Lecture »
Linux CPU Scheduler Internals – Part 2: Who Runs the Scheduler Code?
Busting the “scheduler thread” myth — scheduling runs in process context, by current itself. A lesson from our free Linux kernel development course
Intermediate
6.x Series
100% Free
This lesson on Linux CPU scheduler internals tackles one of the most widespread misconceptions in operating systems. Many developers — even experienced ones — picture a special “scheduler thread” sitting inside the kernel, waking up periodically to shuffle tasks around. That mental model is simply wrong on Linux. In this part of our free Linux kernel development course, you will see who really executes the scheduling code, why scheduling can never happen in interrupt or atomic context, and how this single “golden rule” shapes everyday kernel programming decisions such as choosing between GFP_KERNEL and GFP_ATOMIC.
What You Will Learn
- Why there is no dedicated “scheduler thread” in a monolithic kernel like Linux
- What
currentmeans and how it runs the scheduler on its own behalf - The golden rule: scheduling code must never run in atomic or interrupt context
- How this rule dictates memory-allocation flags inside interrupt handlers
- How the core
schedule()function walks the modular scheduling classes to pick the next task
Prerequisites
You should have completed Part 1 of this series (querying and setting scheduling policy), and you should know what process context and interrupt context mean at a basic level. All of it is covered earlier in this free embedded systems course, so feel free to jump back if needed.
The Big Misconception in Linux CPU Scheduler Internals
Textbooks often draw the scheduler as a separate box that “runs the system,” so it feels natural to assume some kernel entity — a daemon, a thread, a mysterious background process — periodically executes and re-arranges tasks. On Linux, no such entity exists. In a monolithic operating system, scheduling is performed by the very threads being scheduled. Whichever thread is currently executing kernel code on a CPU is the one that, at well-defined moments, calls into the scheduler and possibly switches itself out in favor of another thread.
The kernel refers to the currently executing thread on a CPU through the current macro, which yields a pointer to that thread’s task_struct. So the accurate statement is: the scheduling code is always run by the process context that is currently executing kernel code — by current itself.
|
|
This is the single most clarifying idea in Linux CPU scheduler internals. Think of it like a relay race with no referee: the runner holding the baton (the CPU) is the one who decides — at the designated hand-off points — to pass the baton to the next runner. There is no separate official who takes the baton away.
Inside schedule(): The Heart of Linux CPU Scheduler Internals
The core scheduling logic is anchored in the kernel function schedule() (with the heavy lifting in its internal helper, __schedule(), in kernel/sched/core.c). When invoked, it iterates over the modular scheduling classes in strict priority order — stop, deadline, real-time, fair (EEVDF since kernel 6.6), and idle — asking each class to nominate its best runnable task. The first class that has a runnable task wins, and if that task differs from current, a context switch follows.
| current calls schedule() |
→ | Ask each class in order: deadline? real-time? fair (EEVDF)? idle? |
→ | context switch to the chosen task (if different from current) |
The Golden Rule of Linux CPU Scheduler Internals: No Scheduling in Atomic Context
Here is a rule every kernel and driver developer must internalize: scheduling code must never run in any kind of atomic or interrupt context. Equivalently, interrupt-context code must be guaranteed non-blocking. A hardware interrupt handler has no process identity of its own — it borrows the CPU from whatever thread was running — so it has no legitimate “self” to switch out. Blocking there would freeze an innocent, unrelated thread and can deadlock the system. The kernel enforces this: sleeping in atomic context triggers the infamous “scheduling while atomic” / “BUG: sleeping function called from invalid context” diagnostics.
Also note the complementary detail: while the scheduler core itself runs, kernel preemption is disabled on that CPU — the scheduler must not be preempted mid-decision.
Why You Cannot Use GFP_KERNEL in an Interrupt Handler
This rule has a very practical consequence you will meet constantly in driver work. The kernel memory allocator, when called with the GFP_KERNEL flag, is allowed to block — if memory is tight it may sleep while the kernel reclaims pages. Sleeping means invoking the scheduler, which is exactly what interrupt context forbids. Hence:
/* In process context (e.g., a driver's read() method): OK */
buf = kmalloc(len, GFP_KERNEL); /* may sleep - fine here */
/* In an interrupt handler / softirq / while holding a spinlock: */
buf = kmalloc(len, GFP_ATOMIC); /* never sleeps - required here */
GFP_ATOMIC instructs the memory-management code to never block: it either satisfies the allocation immediately (dipping into emergency reserves if needed) or fails, returning NULL. Your code must handle that failure gracefully.
| Context | May it sleep / call schedule()? | Allocation flag |
|---|---|---|
| Process context (syscall path, kthread) | Yes | GFP_KERNEL |
| Hardirq / softirq / tasklet | Never | GFP_ATOMIC |
| Holding a spinlock | Never | GFP_ATOMIC |
| Holding a mutex (process context) | Yes | GFP_KERNEL |
Seeing It Yourself: current in a Tiny Module
The following original snippet (write it as its own module, insert, and check dmesg) shows that even module init code runs in the context of a perfectly ordinary process — usually insmod itself:
#include <linux/module.h>
#include <linux/sched.h>
static int __init whoami_init(void)
{
pr_info("running in context of: %s (PID %d)\n",
current->comm, task_pid_nr(current));
pr_info("in_interrupt? %s\n",
in_interrupt() ? "yes" : "no");
return 0;
}
static void __exit whoami_exit(void)
{
pr_info("exit in context of: %s (PID %d)\n",
current->comm, task_pid_nr(current));
}
module_init(whoami_init);
module_exit(whoami_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Demonstrates process context via current");
The output names insmod/rmmod as the executing context — living proof that kernel code, including the scheduler, is executed by regular process contexts and not by a phantom kernel entity.
Common Mistakes and Troubleshooting
- Calling a sleeping function under a spinlock: classic driver bug; results in “BUG: scheduling while atomic.” Move the sleeping work outside the lock or switch to a mutex if the context allows.
- Using msleep()/mutex_lock() in an IRQ handler: both may schedule — forbidden. Defer such work to a threaded IRQ handler or workqueue (which run in process context).
- Assuming GFP_ATOMIC always succeeds: it fails more readily than GFP_KERNEL. Always check for NULL.
- Believing a “scheduler process” can be reniced: there is nothing to renice; scheduling overhead is accounted to the threads themselves.
Best Practices
- Before writing any kernel code, ask: what context will this run in? Let the answer drive locking and allocation choices.
- Keep hardirq handlers minimal; push real work to threaded IRQs or workqueues where sleeping is legal.
- Use
might_sleep()in your own sleep-capable helpers so misuse is caught early in debug kernels (CONFIG_DEBUG_ATOMIC_SLEEP).
Key Takeaways
- In Linux CPU scheduler internals there is no dedicated scheduler thread;
currentruns the scheduler code itself. schedule()walks the scheduling classes in priority order to pick the next task.- Golden rule: never invoke the scheduler — directly or via a sleeping function — in atomic or interrupt context.
- That is precisely why interrupt-context allocations use
GFP_ATOMIC, notGFP_KERNEL. - Kernel preemption is disabled while the scheduler core executes.
Frequently Asked Questions
Q1. If no scheduler thread exists, what are kthreads like ksoftirqd?
They are kernel threads that handle deferred work (softirq overflow, in this case) — helpers for specific subsystems. None of them “is” the scheduler; when they run, they too invoke scheduling code as ordinary process contexts.
Q2. What exactly is current?
A per-CPU macro returning a pointer to the task_struct of the thread presently executing on that CPU. On x86-64 it is implemented via a per-CPU variable; on ARM64 it uses a dedicated register (SP_EL0).
Q3. Why is blocking in interrupt context so dangerous?
The handler borrowed the CPU from an arbitrary thread. Blocking would suspend that unrelated victim, and because the handler has no schedulable identity, there may be no way to ever resume correctly — leading to hangs or deadlocks.
Q4. Can I ever allocate memory in an interrupt handler?
Yes, with GFP_ATOMIC — a non-blocking allocation. Prefer pre-allocating buffers at init time when possible; atomic reserves are limited.
Q5. Does the EEVDF switch change any of this?
No. EEVDF (kernel 6.6+) replaces only the fair class’s picking algorithm. Who runs the scheduler, and the atomic-context rule, are architectural facts that remain unchanged.
Q6. Where is the authoritative documentation?
The kernel scheduler docs and the in-tree memory allocation guide at docs.kernel.org/core-api/memory-allocation.
Conclusion
You now hold the correct mental model of Linux CPU scheduler internals: scheduling is a service performed by running threads on themselves, never by a hidden referee, and never from atomic or interrupt context. This insight quietly governs half of kernel programming — locking choices, allocation flags, IRQ design. In Part 3 of this free Linux device drivers course track, we answer the follow-up question: if no thread is dedicated to scheduling, when does the scheduler actually get to run? The timer interrupt plays a clever supporting role — see you there.
Free Linux Kernel Development Course — Keep Going
EmbeddedPathashala offers a completely free Linux kernel development course, free Linux device drivers course, and free embedded systems course for everyone.

2 Comments