Linux Kernel Delay and Sleep Accuracy Explained (Kernel 6.x)-Free Linux Device Driver Course Online

Linux Kernel Delay and Sleep Accuracy Explained (Kernel 6.x) | Free Linux Kernel Development Course

← Previous Lecture  |  Next Lecture →

Why Linux Kernel Delays and Sleeps Are Never Perfectly Accurate
A kernel 6.x guide to delay/sleep timing accuracy, hrtimers, and PREEMPT_RT — part of our free Linux kernel development course
⏱️ Timing Deep-Dive
🧩 Beginner Friendly
🐧 Kernel 6.x Ready

If you have ever measured Linux kernel delay and sleep accuracy inside your own module, you have probably noticed something odd: a 10 millisecond delay does not take exactly 10 milliseconds, and a 10 millisecond sleep almost always takes longer than 10 milliseconds. This is not a bug in your driver — it is how the standard Linux timing subsystem behaves by design. In this lecture, part of our free embedded systems course and free Linux kernel development course, we explain exactly why this happens, how hrtimers help, and how the mainline PREEMPT_RT work (merged in kernel 6.12) changes the picture for real-time workloads.

What You Will Learn
Delay vs sleep timing behavior Jiffies and HZ hrtimers SCHED_FIFO / SCHED_RR PREEMPT_RT in kernel 6.x Writing a timing-measurement module
Prerequisites

Before this lecture, you should be comfortable with:

Writing and loading a basic kernel module The *delay() and *sleep() kernel APIs Reading kernel logs with dmesg

Quick Recap: Delay vs Sleep APIs

The Linux kernel gives driver authors two different families of timing APIs, and picking the wrong one for the situation is one of the most common beginner mistakes.

FamilyExamplesBehaviorTypical Use
Busy-wait delaysndelay(), udelay(), mdelay()CPU spins in a loop; atomic context safeVery short, sub-microsecond to low-millisecond waits
Sleeping delaysusleep_range(), msleep(), ssleep()Task is taken off the CPU and rescheduled laterWaits longer than roughly 10 microseconds, in process context

Why *delay() Functions Often Finish Early

Busy-wait delay functions like udelay() and mdelay() do not use a hardware timer interrupt to know when to stop. Instead, they run a calibrated loop a fixed number of times, based on a value called loops_per_jiffy, which is measured once when the kernel boots. Because this value is only an estimate, and because CPU frequency scaling, cache state, and interrupt overhead can all change how fast the loop actually runs, the measured elapsed time can end up slightly less than what you asked for. On most general-purpose x86_64 and ARM systems today, delay routines are actually implemented on top of a hardware cycle counter (such as the TSC), which makes them fairly close to accurate — but “close” is not “guaranteed,” and small early-return behavior is still officially expected and documented in the kernel source.

Requested vs Actual Delay/Sleep Timing (Illustrative)
udelay(10000) — busy-wait delay
requested: 10,000 ns
actual: ~9,000–9,900 ns (often early)
msleep(10) — scheduling sleep
requested: 10 ms
actual: ~11–20 ms (often late)

Why *sleep() Functions Often Take Longer Than Requested

Sleeping APIs behave the opposite way. When a task calls msleep(), the kernel schedules it to wake up at some future point and hands the CPU to something else in the meantime. That “future point” is usually rounded up to the next available timer tick, and even once the timer fires, the scheduler still has to pick your task and context-switch back to it — which takes extra time, especially if other higher-priority tasks are runnable. On a standard non-real-time kernel, none of these steps are guaranteed to happen at an exact moment, so sleeps almost always run a little longer than requested, never shorter.

Understanding Jiffies, HZ, and Kernel Ticks

A jiffy is the basic unit of kernel timekeeping, and its length depends on the kernel’s configured CONFIG_HZ value. With HZ set to 250, one jiffy is 4 milliseconds; with HZ set to 1000, one jiffy is 1 millisecond. Legacy timer-based APIs like msleep() are still influenced by this tick granularity, which is one reason sleeps tend to round upward. Most distribution kernels today ship with HZ values of 250 or 1000, and many are also built “tickless” (CONFIG_NO_HZ), meaning the timer interrupt is skipped entirely when the CPU is idle — this saves power but adds its own small scheduling latency when a sleeping task needs to be woken up.

High-Resolution Timers (hrtimers) in Modern Kernels

To get much tighter timing than jiffy-based APIs allow, the kernel provides the hrtimer subsystem, which can schedule callbacks with nanosecond-level precision instead of being locked to the tick rate. usleep_range() is built on hrtimers, which is exactly why it is recommended for sleeps in the 10 microsecond to 20 millisecond range — it lets the scheduler pick the most power-efficient wake-up point within the range you give it, rather than forcing an exact but expensive wake-up. If your driver genuinely needs precise, repeating callbacks (rather than a one-shot sleep), you can use the kernel’s hrtimer API directly instead of the simpler delay/sleep wrappers.

Measuring Timing Accuracy: A Kernel 6.x Demo Module

Below is an original, minimal kernel module you can build and load yourself on kernel 6.x to see this behavior firsthand. It times a busy-wait delay and a scheduling sleep using ktime_get_ns() and prints the requested vs actual duration to the kernel log.

// timing_demo.c — original demo module for kernel 6.x
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/init.h>
#include <linux/delay.h>
#include <linux/ktime.h>

#define REQUESTED_US 10000UL   /* 10 ms requested */

static void measure_call(const char *label, void (*fn)(void))
{
        u64 start_ns, end_ns, elapsed_ns;

        start_ns = ktime_get_ns();
        fn();
        end_ns = ktime_get_ns();
        elapsed_ns = end_ns - start_ns;

        pr_info("timing_demo: %s -> requested %lu us, actual %llu us\n",
                label, REQUESTED_US / 1000, elapsed_ns / 1000);
}

static void do_busy_delay(void)
{
        mdelay(REQUESTED_US / 1000);
}

static void do_sleep_delay(void)
{
        msleep(REQUESTED_US / 1000);
}

static int __init timing_demo_init(void)
{
        pr_info("timing_demo: module loaded on kernel %s\n", UTS_RELEASE);
        measure_call("mdelay (busy-wait)", do_busy_delay);
        measure_call("msleep (scheduling sleep)", do_sleep_delay);
        return 0;
}

static void __exit timing_demo_exit(void)
{
        pr_info("timing_demo: module unloaded\n");
}

module_init(timing_demo_init);
module_exit(timing_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("Original demo module measuring delay/sleep timing accuracy");

Load it with sudo insmod timing_demo.ko and check the results with:

dmesg | tail -5

Run it a few times in a row — you will typically see the busy-wait figure land close to (or just under) 10,000 microseconds, while the sleep figure comes in noticeably above 10,000 microseconds, especially under any system load.

Real-Time Scheduling: SCHED_FIFO and SCHED_RR

If a user-space application needs more consistent wake-up timing, one option is to raise its scheduling priority using the real-time policies SCHED_FIFO or SCHED_RR, with priorities ranging from 1 to 99. A real-time-scheduled thread preempts ordinary (SCHED_OTHER) tasks whenever it becomes runnable, which reduces — but does not eliminate — scheduling latency. This is typically combined with POSIX high-resolution timer calls such as timer_create() and timer_settime() in user space.

PREEMPT_RT: Linux as a Real-Time OS (Mainline Since Kernel 6.12)

For years, turning Linux into a hard real-time OS meant applying the separate, out-of-tree PREEMPT_RT patch set on top of your kernel source. That changed with kernel 6.12, when the core PREEMPT_RT support was finally merged into mainline Linux after roughly two decades of development. On a kernel with CONFIG_PREEMPT_RT enabled, most kernel code — including spinlock-protected sections and interrupt handlers — becomes preemptible, which dramatically reduces the worst-case scheduling jitter that causes sleeps to run late. This makes PREEMPT_RT a realistic option today for time-critical embedded workloads such as motor control, industrial automation, and audio processing, without needing to track and rebase a separate patch series.

Best Practices for Timing-Sensitive Kernel Code

Prefer usleep_range() over msleep() for waits under 20 ms Never sleep or busy-wait for long periods inside a spinlock or interrupt handler Use hrtimers directly for repeating, precise callbacks Measure with ktime_get_ns(), not jiffies, for sub-millisecond accuracy Consider PREEMPT_RT only when your workload genuinely needs bounded latency

Common Mistakes and Troubleshooting

MistakeWhy It’s a ProblemFix
Using mdelay() for a 50 ms waitWastes CPU cycles busy-spinning instead of scheduling other workUse msleep() or usleep_range() in process context
Using msleep() inside an interrupt handlerSleeping is not allowed in atomic/interrupt contextUse udelay()/mdelay(), or defer the work to a workqueue
Assuming sleeps are exactLeads to subtle timing bugs in protocol or sensor driversAlways budget extra margin above the minimum required delay

Performance Considerations

Busy-wait delays consume 100% of a CPU core for their entire duration, which is fine for microsecond-scale waits but wasteful and power-hungry for anything longer. Scheduling sleeps free the CPU for other work but cost extra latency from context switching and tick alignment. Choosing the correct API for the expected wait length is itself a performance decision, not just a correctness one.

Security Considerations

Timing behavior can leak information in security-sensitive code — for example, cryptographic routines whose execution time varies with secret data can be vulnerable to timing side-channel attacks. If you are writing kernel code that handles sensitive data, avoid data-dependent delay/sleep durations, and prefer constant-time comparison and processing techniques instead of relying on timing APIs to mask behavior.

Summary / Key Takeaways

*delay() functions can finish slightly early *sleep() functions almost always run slightly late hrtimers give nanosecond-level precision beyond jiffy granularity PREEMPT_RT is now mainline as of kernel 6.12 Pick the right API based on context and wait length
This lecture is part of our free Linux kernel development course, free Linux device drivers course, and free embedded systems course.
free linux kernel development course free linux device drivers course free embedded systems course linux kernel delay and sleep accuracy

Conclusion

Timing in a general-purpose Linux kernel is best-effort, not exact — and once you understand why busy-wait delays tend to finish early while scheduling sleeps tend to run late, you can pick the right API and design your driver around that reality instead of being surprised by it. With hrtimers and, on kernel 6.12+, mainline PREEMPT_RT, Linux now offers a genuine spectrum from “good enough” timing all the way up to real hard real-time behavior when your project needs it.

FAQ

Q1. Why does udelay() sometimes finish before the requested time?
Because it relies on a pre-calibrated loop count rather than a hardware deadline, and factors like interrupt overhead and CPU frequency scaling can make the loop run faster than expected.

Q2. Why does msleep() always seem to sleep longer than asked?
Because the wake-up is aligned to the next available timer tick and then has to wait for the scheduler to run the task again, both of which add extra time on top of the requested duration.

Q3. What is the difference between a jiffy and an hrtimer?
A jiffy is a fixed tick interval set by the kernel’s HZ configuration, while hrtimers can fire at nanosecond precision independent of the tick rate.

Q4. When should I use usleep_range() instead of msleep()?
For waits roughly between 10 microseconds and 20 milliseconds, since usleep_range() uses hrtimers and lets the scheduler choose a more power-efficient wake-up point within your given range.

Q5. Can I use *delay() or *sleep() inside an interrupt handler?
Busy-wait *delay() functions are safe in atomic/interrupt context for very short waits, but *sleep() functions are not, because sleeping requires a schedulable process context.

Q6. What changed with PREEMPT_RT in kernel 6.12?
The long-maintained out-of-tree PREEMPT_RT patch set was merged into the mainline kernel, so real-time preemption is now available with a standard kernel configuration option instead of a separate patch series.

Q7. Does a higher HZ value make sleeps more accurate?
It reduces the maximum rounding error from tick alignment, but scheduling and context-switch overhead still apply, so sleeps remain best-effort rather than exact.

Q8. Is this timing behavior specific to old kernels?
No — this is fundamental to how a general-purpose, non-real-time kernel schedules work, and it still applies on current kernel 6.x releases, though hrtimers and PREEMPT_RT give you more tools to manage it.

Continue the Free Linux Kernel Development Course

More free lectures on kernel timers, threads, and workqueues are coming — bookmark this series and keep learning.

← Previous Lecture  |  Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *