« Previous Lecture | Next Lecture »
If you have followed this free Linux kernel development course so far, you already know how to create a kernel thread, put it to sleep the race-free way, and clean it up correctly. But a kernel thread that simply runs under the default Linux kernel thread scheduling policy is still competing for CPU time exactly like every other ordinary task on the system. For latency-sensitive driver work — servicing an interrupt, feeding a hardware FIFO, or running a real-time control loop — that default behavior is not good enough. In this lecture of our free Linux device drivers course, we look at how to change the scheduling policy and priority of a kernel thread, why the kernel deliberately hides the raw API from most driver authors, and how the interface has actually simplified since kernel 5.9.
Why the Default Scheduler Isn’t Enough for Some Kernel Threads
By default, every kernel thread you spawn is placed under SCHED_NORMAL (technically
SCHED_OTHER), the same completely-fair scheduling class that ordinary user processes use. That’s the
right choice for background housekeeping work — log flushing, periodic cleanup, deferred bookkeeping — because it
plays fairly with everything else on the box. But some kernel threads exist specifically to react quickly: a
thread that drains a hardware ring buffer before it overflows, or one that must run within microseconds of an
event being posted, cannot afford to sit in a fair-share run queue behind a dozen other tasks. For that class of
work, the kernel exposes a small, deliberately GPL-only API to change a thread’s scheduling class and priority.
A Quick Tour of Linux Scheduling Policies
Fair-share, priority via nice (-20 to 19). Default for
almost every task and kthread.
Real-time, no time-slicing. Runs until it blocks or a higher / equal priority RT thread is ready. Priority 1-99.
Same as FIFO but time-sliced against threads at the same priority level, so nobody starves siblings.
Below-normal treatment for throughput-oriented or filler work that should never disturb interactive tasks.
SCHED_FIFO and SCHED_RR are the two policies you will reach for on a kernel thread that
needs deterministic, near-immediate scheduling. Both take a static priority from 1 (lowest real-time priority) to
99 (highest), and both will always preempt any SCHED_NORMAL task regardless of that task’s nice
value.
The Classic API: sched_setscheduler_nocheck()
Historically, driver code that wanted to bump a kernel thread’s priority called
sched_setscheduler_nocheck() directly, passing a policy constant and a struct sched_param
with the desired priority:
struct sched_param sp = { .sched_priority = 50 };
int ret = sched_setscheduler_nocheck(my_task, SCHED_FIFO, &sp);
if (ret)
pr_err("failed to set SCHED_FIFO priority: %d\n", ret);
The _nocheck suffix means the kernel skips the permission checks that a user-space caller would face
(there is no CAP_SYS_NICE concept for code already running in kernel context). The function is
exported strictly via EXPORT_SYMBOL_GPL(), so it is only callable from GPL-licensed modules — this is
intentional, since arbitrarily letting closed-source drivers seize real-time priority would be a stability and
security concern for the rest of the system.
The Modern Way: sched_set_fifo() and Friends (Kernel 5.9+)
Because dozens of in-tree drivers were repeating the same four-line pattern above — and frequently picking priority values inconsistently — the scheduler maintainers added a small set of convenience helpers that are now the recommended way to do this on any current kernel 6.x tree:
sched_set_fifo(struct task_struct *p)— setsSCHED_FIFOat a sane mid-range priority (MAX_RT_PRIO / 2, effectively 50)sched_set_fifo_low(struct task_struct *p)— setsSCHED_FIFOat priority 1, useful when you only need to rank above ordinary tasks, not fight other real-time threadssched_set_normal(struct task_struct *p, int nice)— restores a thread back toSCHED_NORMALwith a chosen nice value
These wrappers still call sched_setscheduler_nocheck() under the hood and are still
EXPORT_SYMBOL_GPL(), but they remove the temptation to invent an arbitrary priority number for every
new driver, and they read far more clearly at the call site.
Hands-On: A Kernel Thread with Real-Time Priority
Here is an original, minimal character-misc driver that spawns a kernel thread and immediately raises it to
SCHED_FIFO using the modern helper, built and tested against kernel 6.x:
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/sched.h>
struct rt_demo {
struct task_struct *thread;
};
static struct rt_demo demo;
static int rt_worker_fn(void *data)
{
pr_info("rt_worker: starting, policy=%d prio=%d\n",
current->policy, current->rt_priority);
while (!kthread_should_stop()) {
/* time-critical work would go here */
schedule_timeout_interruptible(msecs_to_jiffies(20));
}
pr_info("rt_worker: stopping cleanly\n");
return 0;
}
static int __init rt_demo_init(void)
{
demo.thread = kthread_run(rt_worker_fn, NULL, "rt_demo_worker");
if (IS_ERR(demo.thread))
return PTR_ERR(demo.thread);
/* Kernel 5.9+ convenience helper: SCHED_FIFO at priority 50 */
sched_set_fifo(demo.thread);
pr_info("rt_demo: loaded, worker running as SCHED_FIFO\n");
return 0;
}
static void __exit rt_demo_exit(void)
{
kthread_stop(demo.thread);
}
module_init(rt_demo_init);
module_exit(rt_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Demo: kernel thread with SCHED_FIFO priority");
Once loaded, you can confirm the effective policy and priority from user space with the chrt
utility:
$ sudo chrt -p $(pgrep rt_demo_worker)
pid 4821's current scheduling policy: SCHED_FIFO
pid 4821's current scheduling priority: 50
Where This Really Matters: Threaded Interrupt Handlers
You do not always call this API yourself. When you register a threaded interrupt handler with
request_threaded_irq(), the kernel internally creates a kernel thread for your handler and, on
request, can raise it to real-time priority so the hardware bottom-half runs with predictable latency instead of
waiting behind ordinary tasks. Understanding sched_setscheduler_nocheck() and the
sched_set_fifo*() family is what lets you read and reason about that internal machinery instead of
treating it as a black box.
Common Pitfalls
- GPL-only symbol — a proprietary module cannot call these functions; the linker will refuse to resolve them.
- Priority inflation — giving every kthread a high SCHED_FIFO priority defeats the purpose; reserve it for genuinely latency-critical work.
- Starving the system — a real-time thread stuck in a tight loop without yielding can trigger soft-lockup warnings or make the machine feel unresponsive, since SCHED_FIFO threads only give up the CPU voluntarily or when preempted by an equal/higher priority RT thread.
- Forgetting to lower priority on exit — if a thread’s role changes at runtime, drop it back
to
SCHED_NORMALwithsched_set_normal()rather than leaving it real-time indefinitely.
Frequently Asked Questions
Every kthread starts under SCHED_NORMAL (SCHED_OTHER), the same fair-share policy used by ordinary user-space processes, unless you explicitly change it.
Both accept a static real-time priority from 1 (lowest) to 99 (highest). Higher numbers always preempt lower ones and both classes always preempt SCHED_NORMAL tasks.
It is exported via EXPORT_SYMBOL_GPL(), a deliberate kernel policy decision that restricts scheduler-priority control to code the kernel community can audit.
Prefer the sched_set_fifo() / sched_set_fifo_low() / sched_set_normal() helpers added in kernel 5.9 — they wrap the same call with sane, consistent priority defaults.
SCHED_FIFO runs a thread until it blocks or a higher-priority thread preempts it, with no time-slicing. SCHED_RR behaves the same way but time-slices threads that share the same priority level.
Yes, if it never blocks or yields. A tight, non-sleeping SCHED_FIFO loop can starve lower-priority work and trigger soft-lockup detection, so real-time kthreads should always include a blocking or sleeping point.
Use chrt -p <pid>, which prints both the current policy and priority for any thread,
including kernel threads.
Yes — on a PREEMPT_RT kernel, SCHED_FIFO/SCHED_RR priorities are honored far more strictly throughout the kernel, so priority choices made with this API become even more meaningful for real-time behavior.
More free lectures on kernel threads, timers, workqueues, and device drivers are on the way.
Previous Lecture Next Lecture
2 Comments