Linux Kernel Thread Scheduling Policy and Priority Explained-Linux Device Driver Training Online

« Previous Lecture  |  Next Lecture »

Linux Kernel Thread Scheduling Policy and Priority Explained
SCHED_FIFO, SCHED_RR, and sched_set_fifo() on modern Linux Kernel 6.x

If you have followed this free Linux kernel development course so far, you already know how to create a kernel thread, put it to sleep the race-free way, and clean it up correctly. But a kernel thread that simply runs under the default Linux kernel thread scheduling policy is still competing for CPU time exactly like every other ordinary task on the system. For latency-sensitive driver work — servicing an interrupt, feeding a hardware FIFO, or running a real-time control loop — that default behavior is not good enough. In this lecture of our free Linux device drivers course, we look at how to change the scheduling policy and priority of a kernel thread, why the kernel deliberately hides the raw API from most driver authors, and how the interface has actually simplified since kernel 5.9.

free linux kernel development course free linux device drivers course free embedded systems course kernel thread scheduling SCHED_FIFO SCHED_RR
What You Will Learn
Linux scheduling policies (NORMAL, FIFO, RR, BATCH, IDLE) Static real-time priority vs dynamic nice value sched_setscheduler_nocheck() internals Modern sched_set_fifo() / sched_set_normal() helpers Setting priority on a kernel thread you created Why the raw API is GPL-only Priority pitfalls: starvation and soft lockups
Prerequisites
Comfortable creating kthreads with kthread_create() / kthread_run() Understand kthread_should_stop() based cleanup Linux kernel 6.x build environment (headers + module tools) Basic C and pointer handling

Why the Default Scheduler Isn’t Enough for Some Kernel Threads

By default, every kernel thread you spawn is placed under SCHED_NORMAL (technically SCHED_OTHER), the same completely-fair scheduling class that ordinary user processes use. That’s the right choice for background housekeeping work — log flushing, periodic cleanup, deferred bookkeeping — because it plays fairly with everything else on the box. But some kernel threads exist specifically to react quickly: a thread that drains a hardware ring buffer before it overflows, or one that must run within microseconds of an event being posted, cannot afford to sit in a fair-share run queue behind a dozen other tasks. For that class of work, the kernel exposes a small, deliberately GPL-only API to change a thread’s scheduling class and priority.

A Quick Tour of Linux Scheduling Policies

Linux Scheduling Policies at a Glance
SCHED_NORMAL

Fair-share, priority via nice (-20 to 19). Default for almost every task and kthread.

SCHED_FIFO

Real-time, no time-slicing. Runs until it blocks or a higher / equal priority RT thread is ready. Priority 1-99.

SCHED_RR

Same as FIFO but time-sliced against threads at the same priority level, so nobody starves siblings.

SCHED_BATCH / IDLE

Below-normal treatment for throughput-oriented or filler work that should never disturb interactive tasks.

SCHED_FIFO and SCHED_RR are the two policies you will reach for on a kernel thread that needs deterministic, near-immediate scheduling. Both take a static priority from 1 (lowest real-time priority) to 99 (highest), and both will always preempt any SCHED_NORMAL task regardless of that task’s nice value.

The Classic API: sched_setscheduler_nocheck()

Historically, driver code that wanted to bump a kernel thread’s priority called sched_setscheduler_nocheck() directly, passing a policy constant and a struct sched_param with the desired priority:

struct sched_param sp = { .sched_priority = 50 };
int ret = sched_setscheduler_nocheck(my_task, SCHED_FIFO, &sp);
if (ret)
        pr_err("failed to set SCHED_FIFO priority: %d\n", ret);

The _nocheck suffix means the kernel skips the permission checks that a user-space caller would face (there is no CAP_SYS_NICE concept for code already running in kernel context). The function is exported strictly via EXPORT_SYMBOL_GPL(), so it is only callable from GPL-licensed modules — this is intentional, since arbitrarily letting closed-source drivers seize real-time priority would be a stability and security concern for the rest of the system.

The Modern Way: sched_set_fifo() and Friends (Kernel 5.9+)

Because dozens of in-tree drivers were repeating the same four-line pattern above — and frequently picking priority values inconsistently — the scheduler maintainers added a small set of convenience helpers that are now the recommended way to do this on any current kernel 6.x tree:

  • sched_set_fifo(struct task_struct *p) — sets SCHED_FIFO at a sane mid-range priority (MAX_RT_PRIO / 2, effectively 50)
  • sched_set_fifo_low(struct task_struct *p) — sets SCHED_FIFO at priority 1, useful when you only need to rank above ordinary tasks, not fight other real-time threads
  • sched_set_normal(struct task_struct *p, int nice) — restores a thread back to SCHED_NORMAL with a chosen nice value

These wrappers still call sched_setscheduler_nocheck() under the hood and are still EXPORT_SYMBOL_GPL(), but they remove the temptation to invent an arbitrary priority number for every new driver, and they read far more clearly at the call site.

Hands-On: A Kernel Thread with Real-Time Priority

Here is an original, minimal character-misc driver that spawns a kernel thread and immediately raises it to SCHED_FIFO using the modern helper, built and tested against kernel 6.x:

#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/sched.h>

struct rt_demo {
        struct task_struct *thread;
};

static struct rt_demo demo;

static int rt_worker_fn(void *data)
{
        pr_info("rt_worker: starting, policy=%d prio=%d\n",
                current->policy, current->rt_priority);

        while (!kthread_should_stop()) {
                /* time-critical work would go here */
                schedule_timeout_interruptible(msecs_to_jiffies(20));
        }

        pr_info("rt_worker: stopping cleanly\n");
        return 0;
}

static int __init rt_demo_init(void)
{
        demo.thread = kthread_run(rt_worker_fn, NULL, "rt_demo_worker");
        if (IS_ERR(demo.thread))
                return PTR_ERR(demo.thread);

        /* Kernel 5.9+ convenience helper: SCHED_FIFO at priority 50 */
        sched_set_fifo(demo.thread);

        pr_info("rt_demo: loaded, worker running as SCHED_FIFO\n");
        return 0;
}

static void __exit rt_demo_exit(void)
{
        kthread_stop(demo.thread);
}

module_init(rt_demo_init);
module_exit(rt_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Demo: kernel thread with SCHED_FIFO priority");

Once loaded, you can confirm the effective policy and priority from user space with the chrt utility:

$ sudo chrt -p $(pgrep rt_demo_worker)
pid 4821's current scheduling policy: SCHED_FIFO
pid 4821's current scheduling priority: 50

Where This Really Matters: Threaded Interrupt Handlers

You do not always call this API yourself. When you register a threaded interrupt handler with request_threaded_irq(), the kernel internally creates a kernel thread for your handler and, on request, can raise it to real-time priority so the hardware bottom-half runs with predictable latency instead of waiting behind ordinary tasks. Understanding sched_setscheduler_nocheck() and the sched_set_fifo*() family is what lets you read and reason about that internal machinery instead of treating it as a black box.

Common Pitfalls

  • GPL-only symbol — a proprietary module cannot call these functions; the linker will refuse to resolve them.
  • Priority inflation — giving every kthread a high SCHED_FIFO priority defeats the purpose; reserve it for genuinely latency-critical work.
  • Starving the system — a real-time thread stuck in a tight loop without yielding can trigger soft-lockup warnings or make the machine feel unresponsive, since SCHED_FIFO threads only give up the CPU voluntarily or when preempted by an equal/higher priority RT thread.
  • Forgetting to lower priority on exit — if a thread’s role changes at runtime, drop it back to SCHED_NORMAL with sched_set_normal() rather than leaving it real-time indefinitely.

Frequently Asked Questions

Q1. What is the default scheduling policy for a new kernel thread?

Every kthread starts under SCHED_NORMAL (SCHED_OTHER), the same fair-share policy used by ordinary user-space processes, unless you explicitly change it.

Q2. What priority range do SCHED_FIFO and SCHED_RR use?

Both accept a static real-time priority from 1 (lowest) to 99 (highest). Higher numbers always preempt lower ones and both classes always preempt SCHED_NORMAL tasks.

Q3. Why is sched_setscheduler_nocheck() only available to GPL modules?

It is exported via EXPORT_SYMBOL_GPL(), a deliberate kernel policy decision that restricts scheduler-priority control to code the kernel community can audit.

Q4. Should I use sched_setscheduler_nocheck() directly in new driver code?

Prefer the sched_set_fifo() / sched_set_fifo_low() / sched_set_normal() helpers added in kernel 5.9 — they wrap the same call with sane, consistent priority defaults.

Q5. What is the difference between SCHED_FIFO and SCHED_RR?

SCHED_FIFO runs a thread until it blocks or a higher-priority thread preempts it, with no time-slicing. SCHED_RR behaves the same way but time-slices threads that share the same priority level.

Q6. Can a real-time kernel thread hang the whole system?

Yes, if it never blocks or yields. A tight, non-sleeping SCHED_FIFO loop can starve lower-priority work and trigger soft-lockup detection, so real-time kthreads should always include a blocking or sleeping point.

Q7. How do I check the scheduling policy of a running kernel thread from user space?

Use chrt -p <pid>, which prints both the current policy and priority for any thread, including kernel threads.

Q8. Does this API relate to PREEMPT_RT?

Yes — on a PREEMPT_RT kernel, SCHED_FIFO/SCHED_RR priorities are honored far more strictly throughout the kernel, so priority choices made with this API become even more meaningful for real-time behavior.

Continue the Free Linux Kernel Development Course

More free lectures on kernel threads, timers, workqueues, and device drivers are on the way.

Previous Lecture Next Lecture

« Previous Lecture  |  Next Lecture »

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *