Why Linux Interrupt Handlers Must Never Block or Sleep-Linux Device Driver Training in Hyderabad

Why Linux Interrupt Handlers Must Never Block or Sleep
A Beginner-Friendly Guide to Atomic Context Rules in Modern Linux Kernels (6.x)
Free Linux Kernel Programming Course
Free Linux Device Drivers Course
Free Embedded Systems Course

If you are learning Linux device driver development, one of the very first rules you must internalize is this: an interrupt handler must never block. In this lesson from our free Linux kernel programming course, we break down exactly what “blocking” means inside a Linux interrupt handler, why it causes kernel panics in production systems, and how modern kernel developers detect and avoid this class of bug on kernel 6.x. This tutorial is part of our free Linux device drivers course and free embedded systems course, and is written so that even someone new to kernel internals can follow along.

What You Will Learn

  • What “atomic context” and “interrupt context” mean in the Linux kernel
  • Why calling schedule() from an interrupt handler crashes the kernel
  • Which common kernel APIs are safe and unsafe to call from an ISR
  • How to use might_sleep() and lockdep-based debugging on modern kernels
  • How threaded IRQs (request_threaded_irq) solve most blocking problems
  • Practical, original code examples of a correct non-blocking interrupt handler

Prerequisites

  • Basic C programming knowledge
  • Familiarity with writing a simple Linux Kernel Module (LKM)
  • A Linux system running kernel 5.x or 6.x for testing (VM is fine)

What Does “Blocking” Mean in an Interrupt Handler?

Blocking means giving up the CPU and letting the scheduler pick another task to run. This normally happens through a call chain that eventually reaches schedule(). In ordinary process context, blocking is completely normal — a process waiting for disk I/O or a mutex simply goes to sleep and the CPU is handed to someone else.

An interrupt handler (also called a hardirq handler or ISR) is different. It interrupts whatever the CPU was doing, runs on borrowed time, and the kernel expects it to finish quickly and return. There is no “current process” context that makes sense to put to sleep, and the interrupt controller on that CPU core is effectively frozen until the handler returns. If your handler tries to sleep here, the kernel has nowhere sensible to switch to, and the result is a kernel oops or, on most production systems, a full kernel panic.

Process Context vs Interrupt Context
Property Process Context Interrupt Context
Can call schedule()? Yes No
Can sleep on a mutex? Yes No
Runs with preemption enabled? Yes No (during hardirq)
Typical duration Milliseconds+ Microseconds

Which Kernel Calls Are Unsafe Inside an Interrupt Handler?

A surprising number of common kernel functions can end up calling schedule() internally, even though it is not obvious from the function name. The table below lists the categories every new driver author should watch for.

Safe vs Unsafe Calls in Interrupt Context
Category Unsafe (may block) Safe alternative
Memory allocation kmalloc(…, GFP_KERNEL) kmalloc(…, GFP_ATOMIC)
User-space copy copy_to_user() / copy_from_user() Defer to a workqueue or threaded IRQ
Locking mutex_lock() spin_lock() / spin_lock_irqsave()
Waking a task wait_event() (blocks) wake_up() / wake_up_interruptible() (non-blocking)

Note: wake_up() and its variants are always safe to call from interrupt context. They only flip a task’s state from sleeping to runnable; they do not themselves invoke the scheduler on that CPU, so they never block.

Catching Blocking Bugs: might_sleep() and Lockdep on Kernel 6.x

Modern kernels give you excellent tooling to catch this class of bug during development rather than in production. The two most important tools are:

  • might_sleep() — a debug annotation placed inside functions that are allowed to sleep. If it is ever reached while atomic/interrupt context is active, the kernel prints a stack trace pointing straight at the offending call.
  • CONFIG_DEBUG_ATOMIC_SLEEP — a kernel config option (found under Kernel Hacking > Lock Debugging) that enables the checks used by might_sleep(). On kernel 6.x this integrates with the broader lockdep subsystem, so you also get warnings for related issues such as sleeping while holding a spinlock.

Always enable this option on your development and CI kernels. It has near-zero cost in terms of learning effort and catches bugs that would otherwise only appear under real hardware timing on a production device.

The Modern Fix: Threaded IRQs

Older drivers used to do all their work directly inside the hardirq handler and then had to carefully avoid every blocking call. The recommended approach on modern Linux (well established by kernel 6.x) is to use threaded interrupts via request_threaded_irq(). This splits your handler into two parts:

  • A short, fast, non-blocking primary handler that runs in true hardirq context and simply acknowledges the interrupt and decides whether to wake the thread.
  • A threaded handler that runs in normal process context, in its own kernel thread, where sleeping, allocating with GFP_KERNEL, and calling copy_to_user() are all perfectly safe.

Here is an original, simplified example showing the pattern (this is illustrative sample code, not taken from any book or reference):

#include <linux/interrupt.h>
#include <linux/gpio.h>
#include <linux/slab.h>

/* Primary handler: runs in real hardirq context.
 * Keep it minimal — no allocation, no locking that can sleep. */
static irqreturn_t sensor_irq_primary(int irq, void *dev_id)
{
    /* Just acknowledge and hand off to the thread */
    return IRQ_WAKE_THREAD;
}

/* Threaded handler: runs in process context.
 * Safe to allocate, sleep, or copy to user space here. */
static irqreturn_t sensor_irq_thread(int irq, void *dev_id)
{
    struct my_sensor_data *data;

    data = kzalloc(sizeof(*data), GFP_KERNEL); /* safe here */
    if (!data)
        return IRQ_NONE;

    /* read hardware, process data, wake up any waiters */
    wake_up_interruptible(&sensor_waitqueue);

    kfree(data);
    return IRQ_HANDLED;
}

static int sensor_probe(struct platform_device *pdev)
{
    int irq, ret;

    irq = platform_get_irq(pdev, 0);
    if (irq < 0)
        return irq;

    ret = request_threaded_irq(irq, sensor_irq_primary,
                                sensor_irq_thread,
                                IRQF_ONESHOT, "my-sensor", pdev);
    return ret;
}

Using IRQF_ONESHOT keeps the hardware IRQ line masked until the threaded handler completes, which is the correct behavior for most level-triggered peripherals such as I2C or SPI sensors with a dedicated interrupt pin.

Common Mistakes and Troubleshooting Tips

  • Mistake: Using GFP_KERNEL inside a plain (non-threaded) hardirq handler.
    Fix: Switch to GFP_ATOMIC or move the allocation into a threaded handler / workqueue.
  • Mistake: Calling copy_to_user() directly from the primary handler.
    Fix: Buffer the data and copy it out later from process context, or use a threaded IRQ.
  • Mistake: Taking a mutex_lock() in an ISR shared with process-context code.
    Fix: Use a spinlock with spin_lock_irqsave() on the process-context side to prevent deadlock.
  • Mistake: Forgetting to enable CONFIG_DEBUG_ATOMIC_SLEEP on your debug kernel.
    Fix: Always build your development kernel with kernel hacking / lock debugging options turned on.

Best Practices

  • Keep hardirq handlers as short as physically possible — acknowledge, read minimal status, and exit.
  • Prefer request_threaded_irq() over a raw handler for anything beyond trivial work.
  • Use spin_lock_irqsave() / spin_unlock_irqrestore() for data shared between IRQ and process context.
  • Test with CONFIG_DEBUG_ATOMIC_SLEEP and lockdep enabled before shipping any driver.

Performance Considerations

Every microsecond a hardirq handler spends is a microsecond during which that interrupt line (and often other work on the same CPU core) cannot be serviced. Long-running handlers increase interrupt latency system-wide, which is especially harmful on real-time or low-latency embedded targets. Threaded IRQs trade a small amount of scheduling latency for a large improvement in system responsiveness, which is why they are the default recommendation today.

Security Considerations

A handler that blocks or panics can be triggered by a malicious or malfunctioning peripheral to cause a denial-of-service against the entire system. Keeping handlers short and provably non-blocking is not just a stability requirement — it is also a hardening measure against interrupt-storm style attacks on embedded and IoT devices.

Summary / Key Takeaways

  • Interrupt handlers must never call anything that can invoke schedule().
  • Use GFP_ATOMIC, spinlocks, and wake_up() variants inside hardirq context.
  • Enable CONFIG_DEBUG_ATOMIC_SLEEP and might_sleep() checks during development.
  • Prefer request_threaded_irq() so most of your logic can safely run in process context.

Conclusion

Understanding why interrupt handlers cannot block is one of the foundational skills in Linux kernel and device driver programming. Once you internalize the difference between process context and interrupt context, and adopt threaded IRQs as your default pattern, an entire category of hard-to-reproduce kernel panics simply disappears from your driver code. This lesson is part of our ongoing free Linux kernel programming course — continue to the next lecture to learn how interrupt masking works across CPU cores.

Frequently Asked Questions

Q1. Why can’t I call schedule() from an interrupt handler?

Because interrupt context has no meaningful “current task” to switch away from cleanly, and the kernel does not support it — calling schedule() there causes an oops or panic.

Q2. Is GFP_ATOMIC always safe in an interrupt handler?

Yes, GFP_ATOMIC never sleeps, though it may fail more often under memory pressure since it cannot wait for memory to be freed.

Q3. What is the difference between a hardirq and a threaded IRQ handler?

A hardirq handler runs in true interrupt context with restrictions on blocking calls; a threaded handler runs in a dedicated kernel thread in process context, where blocking calls are safe.

Q4. Does wake_up() ever block?

No, wake_up() and its variants only mark a task as runnable; the actual context switch happens later via the scheduler.

Q5. How do I detect blocking bugs before they crash a production device?

Enable CONFIG_DEBUG_ATOMIC_SLEEP and lockdep on your development kernel; they print a stack trace the moment a sleeping call happens in atomic context.

Q6. Is this behavior different on real-time (PREEMPT_RT) kernels?

PREEMPT_RT converts most hardirq handlers into threaded handlers by default, but the same rule still applies to any code that genuinely must run in true atomic context, such as spinlock-protected sections.

Q7. Can I use IRQF_ONESHOT with any interrupt type?

It is primarily intended for level-triggered interrupts shared between a primary and threaded handler, such as those on I2C/SPI sensor lines; check your specific hardware’s interrupt trigger type before applying it.

Want more free Linux kernel and device driver lessons?
Explore the Free Course

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *