← Previous Lecture | Next Lecture →
If you have ever measured Linux kernel delay and sleep accuracy inside your own module, you have probably noticed something odd: a 10 millisecond delay does not take exactly 10 milliseconds, and a 10 millisecond sleep almost always takes longer than 10 milliseconds. This is not a bug in your driver — it is how the standard Linux timing subsystem behaves by design. In this lecture, part of our free embedded systems course and free Linux kernel development course, we explain exactly why this happens, how hrtimers help, and how the mainline PREEMPT_RT work (merged in kernel 6.12) changes the picture for real-time workloads.
Before this lecture, you should be comfortable with:
Quick Recap: Delay vs Sleep APIs
The Linux kernel gives driver authors two different families of timing APIs, and picking the wrong one for the situation is one of the most common beginner mistakes.
| Family | Examples | Behavior | Typical Use |
|---|---|---|---|
| Busy-wait delays | ndelay(), udelay(), mdelay() | CPU spins in a loop; atomic context safe | Very short, sub-microsecond to low-millisecond waits |
| Sleeping delays | usleep_range(), msleep(), ssleep() | Task is taken off the CPU and rescheduled later | Waits longer than roughly 10 microseconds, in process context |
Why *delay() Functions Often Finish Early
Busy-wait delay functions like udelay() and mdelay() do not use a hardware timer interrupt to know when to stop. Instead, they run a calibrated loop a fixed number of times, based on a value called loops_per_jiffy, which is measured once when the kernel boots. Because this value is only an estimate, and because CPU frequency scaling, cache state, and interrupt overhead can all change how fast the loop actually runs, the measured elapsed time can end up slightly less than what you asked for. On most general-purpose x86_64 and ARM systems today, delay routines are actually implemented on top of a hardware cycle counter (such as the TSC), which makes them fairly close to accurate — but “close” is not “guaranteed,” and small early-return behavior is still officially expected and documented in the kernel source.
Why *sleep() Functions Often Take Longer Than Requested
Sleeping APIs behave the opposite way. When a task calls msleep(), the kernel schedules it to wake up at some future point and hands the CPU to something else in the meantime. That “future point” is usually rounded up to the next available timer tick, and even once the timer fires, the scheduler still has to pick your task and context-switch back to it — which takes extra time, especially if other higher-priority tasks are runnable. On a standard non-real-time kernel, none of these steps are guaranteed to happen at an exact moment, so sleeps almost always run a little longer than requested, never shorter.
Understanding Jiffies, HZ, and Kernel Ticks
A jiffy is the basic unit of kernel timekeeping, and its length depends on the kernel’s configured CONFIG_HZ value. With HZ set to 250, one jiffy is 4 milliseconds; with HZ set to 1000, one jiffy is 1 millisecond. Legacy timer-based APIs like msleep() are still influenced by this tick granularity, which is one reason sleeps tend to round upward. Most distribution kernels today ship with HZ values of 250 or 1000, and many are also built “tickless” (CONFIG_NO_HZ), meaning the timer interrupt is skipped entirely when the CPU is idle — this saves power but adds its own small scheduling latency when a sleeping task needs to be woken up.
High-Resolution Timers (hrtimers) in Modern Kernels
To get much tighter timing than jiffy-based APIs allow, the kernel provides the hrtimer subsystem, which can schedule callbacks with nanosecond-level precision instead of being locked to the tick rate. usleep_range() is built on hrtimers, which is exactly why it is recommended for sleeps in the 10 microsecond to 20 millisecond range — it lets the scheduler pick the most power-efficient wake-up point within the range you give it, rather than forcing an exact but expensive wake-up. If your driver genuinely needs precise, repeating callbacks (rather than a one-shot sleep), you can use the kernel’s hrtimer API directly instead of the simpler delay/sleep wrappers.
Measuring Timing Accuracy: A Kernel 6.x Demo Module
Below is an original, minimal kernel module you can build and load yourself on kernel 6.x to see this behavior firsthand. It times a busy-wait delay and a scheduling sleep using ktime_get_ns() and prints the requested vs actual duration to the kernel log.
// timing_demo.c — original demo module for kernel 6.x
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/init.h>
#include <linux/delay.h>
#include <linux/ktime.h>
#define REQUESTED_US 10000UL /* 10 ms requested */
static void measure_call(const char *label, void (*fn)(void))
{
u64 start_ns, end_ns, elapsed_ns;
start_ns = ktime_get_ns();
fn();
end_ns = ktime_get_ns();
elapsed_ns = end_ns - start_ns;
pr_info("timing_demo: %s -> requested %lu us, actual %llu us\n",
label, REQUESTED_US / 1000, elapsed_ns / 1000);
}
static void do_busy_delay(void)
{
mdelay(REQUESTED_US / 1000);
}
static void do_sleep_delay(void)
{
msleep(REQUESTED_US / 1000);
}
static int __init timing_demo_init(void)
{
pr_info("timing_demo: module loaded on kernel %s\n", UTS_RELEASE);
measure_call("mdelay (busy-wait)", do_busy_delay);
measure_call("msleep (scheduling sleep)", do_sleep_delay);
return 0;
}
static void __exit timing_demo_exit(void)
{
pr_info("timing_demo: module unloaded\n");
}
module_init(timing_demo_init);
module_exit(timing_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Original demo module measuring delay/sleep timing accuracy");
Load it with sudo insmod timing_demo.ko and check the results with:
dmesg | tail -5
Run it a few times in a row — you will typically see the busy-wait figure land close to (or just under) 10,000 microseconds, while the sleep figure comes in noticeably above 10,000 microseconds, especially under any system load.
Real-Time Scheduling: SCHED_FIFO and SCHED_RR
If a user-space application needs more consistent wake-up timing, one option is to raise its scheduling priority using the real-time policies SCHED_FIFO or SCHED_RR, with priorities ranging from 1 to 99. A real-time-scheduled thread preempts ordinary (SCHED_OTHER) tasks whenever it becomes runnable, which reduces — but does not eliminate — scheduling latency. This is typically combined with POSIX high-resolution timer calls such as timer_create() and timer_settime() in user space.
PREEMPT_RT: Linux as a Real-Time OS (Mainline Since Kernel 6.12)
For years, turning Linux into a hard real-time OS meant applying the separate, out-of-tree PREEMPT_RT patch set on top of your kernel source. That changed with kernel 6.12, when the core PREEMPT_RT support was finally merged into mainline Linux after roughly two decades of development. On a kernel with CONFIG_PREEMPT_RT enabled, most kernel code — including spinlock-protected sections and interrupt handlers — becomes preemptible, which dramatically reduces the worst-case scheduling jitter that causes sleeps to run late. This makes PREEMPT_RT a realistic option today for time-critical embedded workloads such as motor control, industrial automation, and audio processing, without needing to track and rebase a separate patch series.
Best Practices for Timing-Sensitive Kernel Code
Common Mistakes and Troubleshooting
| Mistake | Why It’s a Problem | Fix |
|---|---|---|
| Using mdelay() for a 50 ms wait | Wastes CPU cycles busy-spinning instead of scheduling other work | Use msleep() or usleep_range() in process context |
| Using msleep() inside an interrupt handler | Sleeping is not allowed in atomic/interrupt context | Use udelay()/mdelay(), or defer the work to a workqueue |
| Assuming sleeps are exact | Leads to subtle timing bugs in protocol or sensor drivers | Always budget extra margin above the minimum required delay |
Performance Considerations
Busy-wait delays consume 100% of a CPU core for their entire duration, which is fine for microsecond-scale waits but wasteful and power-hungry for anything longer. Scheduling sleeps free the CPU for other work but cost extra latency from context switching and tick alignment. Choosing the correct API for the expected wait length is itself a performance decision, not just a correctness one.
Security Considerations
Timing behavior can leak information in security-sensitive code — for example, cryptographic routines whose execution time varies with secret data can be vulnerable to timing side-channel attacks. If you are writing kernel code that handles sensitive data, avoid data-dependent delay/sleep durations, and prefer constant-time comparison and processing techniques instead of relying on timing APIs to mask behavior.
Summary / Key Takeaways
Conclusion
Timing in a general-purpose Linux kernel is best-effort, not exact — and once you understand why busy-wait delays tend to finish early while scheduling sleeps tend to run late, you can pick the right API and design your driver around that reality instead of being surprised by it. With hrtimers and, on kernel 6.12+, mainline PREEMPT_RT, Linux now offers a genuine spectrum from “good enough” timing all the way up to real hard real-time behavior when your project needs it.
FAQ
Q1. Why does udelay() sometimes finish before the requested time?
Because it relies on a pre-calibrated loop count rather than a hardware deadline, and factors like interrupt overhead and CPU frequency scaling can make the loop run faster than expected.
Q2. Why does msleep() always seem to sleep longer than asked?
Because the wake-up is aligned to the next available timer tick and then has to wait for the scheduler to run the task again, both of which add extra time on top of the requested duration.
Q3. What is the difference between a jiffy and an hrtimer?
A jiffy is a fixed tick interval set by the kernel’s HZ configuration, while hrtimers can fire at nanosecond precision independent of the tick rate.
Q4. When should I use usleep_range() instead of msleep()?
For waits roughly between 10 microseconds and 20 milliseconds, since usleep_range() uses hrtimers and lets the scheduler choose a more power-efficient wake-up point within your given range.
Q5. Can I use *delay() or *sleep() inside an interrupt handler?
Busy-wait *delay() functions are safe in atomic/interrupt context for very short waits, but *sleep() functions are not, because sleeping requires a schedulable process context.
Q6. What changed with PREEMPT_RT in kernel 6.12?
The long-maintained out-of-tree PREEMPT_RT patch set was merged into the mainline kernel, so real-time preemption is now available with a standard kernel configuration option instead of a separate patch series.
Q7. Does a higher HZ value make sleeps more accurate?
It reduces the maximum rounding error from tick alignment, but scheduling and context-switch overhead still apply, so sleeps remain best-effort rather than exact.
Q8. Is this timing behavior specific to old kernels?
No — this is fundamental to how a general-purpose, non-real-time kernel schedules work, and it still applies on current kernel 6.x releases, though hrtimers and PREEMPT_RT give you more tools to manage it.
More free lectures on kernel timers, threads, and workqueues are coming — bookmark this series and keep learning.

2 Comments