Linux Scheduling Classes and the EEVDF Scheduler: How the Modern Kernel Picks the Next Task
From CFS to EEVDF to sched_ext — a lecture from our free Linux kernel development course
Intermediate
6.6 → 7.x
100% Free
In the previous lecture of this free Linux kernel development course we studied scheduling policies. Now we go one level deeper: how does the kernel organize those policies internally, and what algorithm actually picks the next thread? This is where Linux scheduling classes and the EEVDF scheduler come in. Most older books and tutorials still describe CFS, but the kernel replaced it — EEVDF became the fair-class engine in kernel 6.6, and by the 6.14+ and 7.x kernels the CFS code path is history. This lecture teaches you the current design, making it a genuinely up-to-date resource for anyone following a free Linux device drivers course or free embedded systems course.
What You Will Learn
Why CFS was replaced
EEVDF: lag, eligibility, virtual deadlines
vruntime concept
sched_ext (BPF schedulers)
Inspecting the scheduler from /proc and /sys
Prerequisites
You should know the Linux scheduling policies (SCHED_NORMAL, SCHED_FIFO, SCHED_RR, SCHED_DEADLINE) from the previous lecture of this free Linux kernel development course, and be comfortable running shell commands.
Scheduling Classes: The Kernel’s Chain of Command
Internally, the kernel does not treat all policies with one algorithm. It groups them into scheduling classes, arranged in a strict hierarchy. When a CPU needs a new task, the core scheduler asks each class in order: “do you have something runnable?” The first class that answers yes wins — lower classes are not even consulted.
| Order | Class | Policies Handled | Purpose |
|---|---|---|---|
| 1 (asked first) | stop | (kernel internal) | Per-CPU machinery: migration, hotplug. Unstoppable. |
| 2 | dl (deadline) | SCHED_DEADLINE | Earliest-Deadline-First with admission control. |
| 3 | rt (real-time) | SCHED_FIFO, SCHED_RR | Fixed-priority soft real-time (1–99). |
| 4 | fair | SCHED_NORMAL, SCHED_BATCH | EEVDF fairness engine — runs almost everything. |
| 5 | ext | SCHED_EXT | Custom BPF-defined schedulers (kernel 6.12+). |
| 6 (asked last) | idle | (idle task) | Runs when nothing else wants the CPU; enters power saving. |
A thread belongs to exactly one class at any point in time, determined by its policy. This chain-of-command design explains everything from the previous lecture: real-time beats normal simply because the rt class is asked before the fair class.
A Quick Word on vruntime: The Idea That Survived
Both the old CFS and the new EEVDF scheduler are built on the same beautiful idea: virtual runtime (vruntime). The kernel tracks how long each fair-class thread has actually run, then scales that time by the thread’s weight (derived from its nice value). A high-priority thread’s clock ticks slowly; a low-priority thread’s clock ticks fast.
Intuition: imagine every thread carries a personal stopwatch. Fairness simply means keeping everyone’s stopwatch reading approximately equal. A thread whose stopwatch is behind everyone else deserves the CPU.
CFS stopped there: it always picked the smallest vruntime, using a red-black tree ordered by vruntime. That guaranteed long-term fairness — but it had a blind spot.
Why CFS Was Replaced: The Latency Problem
Pure fairness ignores urgency. A video-conferencing thread and a background compiler thread with equal nice values received identical treatment under CFS, even though the first one needs the CPU in small, frequent, on-time doses and the second one just needs total throughput. CFS papered over this with a pile of heuristics (wakeup boosts, latency-nice experiments), which grew fragile and hard to maintain.
The EEVDF scheduler (Earliest Eligible Virtual Deadline First), based on a 1995 academic algorithm and implemented by Peter Zijlstra, replaces those heuristics with two clean concepts: eligibility and virtual deadlines. It merged in kernel 6.6 and fully replaced the CFS code path in later releases.
How the EEVDF Scheduler Works: Lag, Eligibility, Deadlines
The EEVDF scheduler makes its decision in two steps:
Step 1 — Eligibility (fairness check). For each thread, the scheduler computes lag = the ideal CPU time the thread should have received minus what it actually received. Positive lag means the system owes the thread time; negative lag means the thread has already taken more than its fair share. Only threads with lag ≥ 0 are eligible to run. This preserves fairness exactly like CFS did.
Step 2 — Virtual deadline (urgency check). Among eligible threads, each gets a virtual deadline: roughly vruntime + its requested timeslice. The thread with the earliest virtual deadline runs first. Here is the elegant part: a thread that asks for a shorter slice gets an earlier deadline, so latency-sensitive tasks are naturally serviced quickly — they get the CPU often but briefly, while throughput tasks get it rarely but for long stretches. Same total fairness, better responsiveness.
| All runnable fair-class threads |
| ↓ compute lag = owed time − received time |
| Keep only ELIGIBLE threads (lag ≥ 0) |
| ↓ virtual deadline = vruntime + requested slice |
| RUN the eligible thread with the EARLIEST virtual deadline |
Two more modern details worth knowing:
- Sleeping tasks and lag decay: a task cannot cheat by sleeping briefly to erase its negative lag. The kernel keeps it marked on the queue (“deferred dequeue”) and lets its lag decay gradually over virtual time.
- Custom slices: applications can request a specific timeslice through
sched_setattr(), giving latency-sensitive programs a first-class, heuristic-free mechanism.
// Requesting a short custom slice with sched_setattr (kernel 6.6+)
#define _GNU_SOURCE
#include <linux/sched.h>
#include <linux/sched/types.h>
#include <sys/syscall.h>
#include <unistd.h>
int main(void)
{
struct sched_attr attr = {
.size = sizeof(attr),
.sched_policy = SCHED_OTHER,
.sched_runtime = 2000000, /* hint: 2 ms slice */
};
/* glibc has no wrapper; call it directly */
syscall(SYS_sched_setattr, 0, &attr, 0);
/* latency-sensitive loop here */
return 0;
}
sched_ext: Write Your Own Linux CPU Scheduler in BPF
The newest member of the hierarchy is the ext class (kernel 6.12+). sched_ext lets you load an entire scheduling algorithm as a BPF program at runtime — no kernel rebuild, no reboot. If your BPF scheduler misbehaves (stalls a task, crashes), the kernel automatically ejects it and falls back to the default scheduler, so the system stays safe. Recent kernels have even added sub-scheduler support, moving toward different schedulers per cgroup.
Why should an embedded engineer care? Specialized workloads — gaming handhelds, telco appliances, ML inference boxes — often benefit from scheduling rules a general-purpose algorithm cannot assume. sched_ext turned scheduler development from a years-long kernel-community effort into a weekend experiment.
# Check whether a BPF scheduler is currently active
cat /sys/kernel/sched_ext/state
# Kernel config options involved
# CONFIG_SCHED_CLASS_EXT=y
# CONFIG_BPF_SYSCALL=y
# Emergency escape hatch: eject the BPF scheduler via SysRq
echo S > /proc/sysrq-trigger
Inspecting the Scheduler on a Live System
Practical commands you can try right now — this is a hands-on free Linux kernel development course, after all:
# Per-thread scheduler statistics (vruntime, switches, policy)
cat /proc/self/sched | head -25
# System-wide scheduler state dump (needs CONFIG_SCHED_DEBUG)
sudo cat /sys/kernel/debug/sched/debug | less
# Tunable: base slice used for virtual deadlines (nanoseconds)
cat /sys/kernel/debug/sched/base_slice_ns
# Count context switches per second, system wide
vmstat 1 5 # watch the "cs" column
In /proc/<pid>/sched you will see fields like se.vruntime, se.deadline, and counts of voluntary vs involuntary context switches — each one maps directly to the EEVDF concepts you just learned.
Real-World Use Cases
- Media and conferencing apps benefit from EEVDF’s deadline behavior: short-slice requests keep audio callbacks on time without any real-time privilege.
- Embedded gateways mix a SCHED_DEADLINE sampling loop (dl class), a FIFO communication thread (rt class), and everything else in the fair class — the class hierarchy guarantees the ordering.
- Fleet operators ship sched_ext BPF schedulers tuned to their exact workload, updating scheduling policy like an app instead of like a kernel.
Common Mistakes and Troubleshooting
- Studying only CFS material: interviewers now ask about EEVDF. Know lag, eligibility, and virtual deadlines — and know that CFS is gone from current kernels.
- Assuming classes share the CPU proportionally: they do not. The hierarchy is absolute; the fair class only runs when dl and rt classes are empty.
- Reading vruntime as wall-clock time: it is weighted virtual time; comparing vruntime across threads of different nice values without the weight context misleads you.
- Expecting sched_ext on old kernels: you need 6.12+ with
CONFIG_SCHED_CLASS_EXTenabled.
Key Takeaways
- The kernel organizes policies into classes asked in strict order: stop → dl → rt → fair → ext → idle.
- The fair class engine is now the EEVDF scheduler (kernel 6.6+), not CFS.
- EEVDF = fairness via lag/eligibility + responsiveness via earliest virtual deadline.
- Shorter requested slices earn earlier deadlines — latency-sensitive tasks win without heuristics.
- sched_ext (6.12+) lets you load custom BPF schedulers safely at runtime.
Interview Questions and FAQ
Q1. What replaced CFS in the Linux kernel and when?
The EEVDF (Earliest Eligible Virtual Deadline First) scheduler, merged in kernel 6.6 as the new fair-class engine, with CFS fully retired in subsequent releases.
Q2. What is lag in the EEVDF scheduler?
Lag is the difference between the CPU time a task ideally deserved and what it actually received. Positive lag makes a task eligible to run; negative lag temporarily excludes it, preserving fairness.
Q3. How does EEVDF improve latency compared to CFS?
Among eligible tasks it runs the one with the earliest virtual deadline. Tasks requesting shorter slices receive earlier deadlines, so interactive work is serviced promptly without special-case heuristics.
Q4. In what order does the kernel consult scheduling classes?
stop, deadline (dl), real-time (rt), fair, ext, idle. The first class with a runnable task supplies the next thread; lower classes are skipped entirely.
Q5. What is sched_ext and why is it significant?
A scheduling class (kernel 6.12+) whose behavior is defined by loadable BPF programs. It enables safe, runtime-swappable custom schedulers, with the kernel auto-reverting to defaults if the BPF scheduler errs.
Q6. Can a fair-class thread ever run while an rt-class thread is runnable on the same CPU?
Effectively no — except for the small RT-throttling reserve the kernel keeps (see sched_rt_runtime_us) so the system never fully starves non-RT work.
Q7. Where can you observe a thread’s vruntime and deadline on a live system?
In /proc/<pid>/sched, and system-wide in /sys/kernel/debug/sched/debug on kernels built with scheduler debugging.
References
- EEVDF documentation: docs.kernel.org/scheduler/sched-eevdf
- sched_ext documentation: docs.kernel.org/scheduler/sched-ext
- man pages:
sched(7),sched_setattr(2)
Keep Going — It’s All Free
Next lecture: visualizing which thread runs on which CPU using perf.

2 Comments