How to Interrupt Masking in the Linux Kernel Explained-Free Linux Device Driver Training in Hyderabad

Interrupt Masking in the Linux Kernel Explained
How IRQ Enable/Disable Works Across CPU Cores on Modern (6.x) Kernels
Free Linux Kernel Programming Course
Free Linux Device Drivers Course
Free Embedded Systems Course

Every hardware interrupt in Linux is controlled through a mask register on the interrupt controller. In this lesson of our free Linux kernel programming course we explain interrupt masking in the Linux kernel: what gets masked by default when your handler runs, why keeping interrupts enabled matters for system performance, and how this behaves on multicore systems running kernel 6.x. This tutorial continues our free Linux device drivers course and free embedded systems course series.

What You Will Learn

  • What interrupt masking means and which hardware controls it
  • What the kernel masks by default when your IRQ handler runs
  • How masking behaves differently across CPU cores on SMP systems
  • Why keeping interrupts unmasked as long as possible matters
  • How to safely disable/enable interrupts yourself using local_irq APIs
  • A practical, original code example of a short critical section

Prerequisites

  • Completion of, or familiarity with, the previous lesson on non-blocking interrupt handlers
  • Basic understanding of CPU interrupt lines and IRQ numbers
  • A Linux system running kernel 5.x or 6.x for experimentation

What Is Interrupt Masking?

The interrupt controller — a GIC (Generic Interrupt Controller) on most modern ARM SoCs, or an APIC on x86 — maintains a mask register per interrupt line. When a line is masked, the CPU simply ignores that hardware signal until it is unmasked again. Masking is how the kernel and drivers temporarily silence a peripheral’s ability to interrupt the CPU, for example while a shared data structure is being updated.

Not every interrupt can be masked. The Non-Maskable Interrupt (NMI) is deliberately designed to bypass masking, and is reserved for conditions the system must react to no matter what — hardware watchdog events or fatal error reporting are typical uses.

What Gets Masked by Default When Your Handler Runs

When IRQ number N fires and your registered hardirq handler starts executing, the Linux kernel automatically does two things for you:

  • Disables further delivery of that same IRQ number (N) on every CPU core in the system, so your handler is never re-entered by itself.
  • Disables all interrupts on the specific CPU core that is currently running your handler, so nothing else can preempt it mid-execution on that core.

This default behavior is what makes a hardirq handler inherently safe from being interrupted by another instance of itself. You do not need to manually mask the IRQ line yourself in the common case — the kernel already did it before calling you.

IRQ Masking Across a 4-Core System
CPU Core 0 CPU Core 1 CPU Core 2 CPU Core 3
Running IRQ 42
All local interrupts masked
Idle / other work
IRQ 42 masked here too
Idle / other work
IRQ 42 masked here too
Idle / other work
Other IRQs still enabled

Notice in the diagram above that only IRQ 42 itself is blocked system-wide, while other unrelated interrupts remain free to fire and run in parallel on cores 1 to 3. Core 0 is the only one with all local interrupts fully masked, and only for the short duration of the handler.

Why Keeping Interrupts Enabled Matters

Interrupts staying unmasked for as much of the time as possible is a genuine measure of operating system quality. If a peripheral’s interrupt line stays masked too long, that peripheral cannot signal the CPU, and from the user’s point of view the system feels laggy or unresponsive. A simple example: pressing and releasing a single keyboard key generates two separate hardware interrupts, and every one of those needs to be serviced promptly for input to feel instant.

This is also why taking a spinlock matters here: acquiring a spinlock with spin_lock_irqsave() disables local interrupts and preemption for the duration you hold it. The kernel-wide expectation is that such critical sections stay extremely short — ideally just a handful of instructions — precisely because everything else on that core is frozen while you hold the lock.

Practical Example: A Short, Correct Critical Section

Below is an original example showing the recommended pattern for protecting a small piece of shared state between a hardirq handler and process context using spin_lock_irqsave():

#include <linux/spinlock.h>
#include <linux/interrupt.h>

static DEFINE_SPINLOCK(counter_lock);
static unsigned long event_counter;

/* Called from process context (e.g. a sysfs read) */
static unsigned long read_event_counter(void)
{
    unsigned long flags, value;

    spin_lock_irqsave(&counter_lock, flags);
    value = event_counter;
    spin_unlock_irqrestore(&counter_lock, flags);

    return value;
}

/* Hardirq handler: keep the critical section as short as possible */
static irqreturn_t counter_irq_handler(int irq, void *dev_id)
{
    spin_lock(&counter_lock);   /* local IRQs already disabled here */
    event_counter++;
    spin_unlock(&counter_lock);

    return IRQ_HANDLED;
}

Notice that inside the handler we use the plain spin_lock(), not spin_lock_irqsave(), because local interrupts on this core are already masked by the kernel while the handler runs. On the process-context side, however, we must use spin_lock_irqsave() because an interrupt could legitimately fire on that core at any moment and try to take the same lock, which would deadlock without the “irqsave” variant.

Common Mistakes and Troubleshooting Tips

  • Mistake: Using plain spin_lock() in process-context code that shares data with an IRQ handler.
    Fix: Always use spin_lock_irqsave() / spin_unlock_irqrestore() on the process-context side.
  • Mistake: Holding a spinlock for an extended computation.
    Fix: Copy the minimal data you need out of the critical section and process it afterward with interrupts enabled.
  • Mistake: Assuming masking IRQ N also blocks unrelated IRQ M.
    Fix: Remember only the specific IRQ number is masked system-wide; other lines remain active on other cores.
  • Mistake: Expecting an NMI to respect your masking.
    Fix: Design any logic depending on NMI behavior separately, since NMIs are explicitly non-maskable.

Best Practices

  • Keep every interrupt-disabled critical section as short as humanly possible.
  • Use spin_lock_irqsave() on the process-context side of any IRQ-shared data.
  • Let the kernel’s default per-IRQ masking do its job instead of manually masking lines unless you have a specific reason.
  • Profile interrupt latency on real hardware if you suspect a driver is masking interrupts for too long.

Performance Considerations

Because masking freezes an entire CPU core’s ability to respond to any other interrupt, long masked sections directly translate into worse system-wide interrupt latency, not just latency for the one device involved. On multicore hardware this is somewhat mitigated since unrelated interrupts can still be serviced on other cores in parallel, but any driver that holds interrupts disabled for too long on a busy core will be felt across the whole system.

Security Considerations

A driver bug that disables interrupts for an unbounded amount of time (for example, an infinite loop inside a spin_lock_irqsave() section) can be used to stall an entire CPU core, effectively creating a local denial-of-service condition. Bounding every interrupt-disabled section and testing under load is an important hardening step for any driver handling untrusted or externally triggerable input.

Summary / Key Takeaways

  • Interrupt masking is controlled per line by the interrupt controller (GIC/APIC).
  • The kernel automatically masks the firing IRQ system-wide and all IRQs on the local core while your hardirq handler runs.
  • NMIs deliberately bypass masking and are reserved for critical events.
  • Keep interrupt-disabled critical sections extremely short to preserve system responsiveness.

Conclusion

Interrupt masking is the mechanism that keeps a running interrupt handler safe from re-entrancy, but it comes at the cost of temporarily freezing responsiveness on that CPU core. Understanding exactly what gets masked, and for how long, lets you write drivers that are both correct and fast. This wraps up this pair of lessons on interrupt handling fundamentals in our free Linux kernel programming course.

Frequently Asked Questions

Q1. Does masking IRQ 42 also block other interrupts on the same core?

Yes — while your hardirq handler for IRQ 42 runs, all local interrupts on that specific core are disabled, not just IRQ 42, though IRQ 42 itself is additionally masked on every core.

Q2. Can an NMI interrupt my hardirq handler?

Yes, by design NMIs are non-maskable and can preempt even a running interrupt handler.

Q3. Why do I need spin_lock_irqsave() instead of spin_lock() in process context?

Because process-context code can be interrupted by the very IRQ handler it shares data with; irqsave disables local interrupts before taking the lock to prevent a deadlock.

Q4. Is manual IRQ masking with disable_irq() ever needed?

Occasionally, for coordinating with code outside the automatic per-handler masking window, but it should be used sparingly since it affects the entire IRQ line, not just one core.

Q5. Does this behave the same on single-core embedded systems?

The concept is the same, but there is only one core to mask, so a masked interrupt effectively halts all other interrupt processing on the entire system until the handler completes.

Q6. How long is “too long” for a masked critical section?

There is no universal number, but the guiding principle is microseconds, not milliseconds — if your section is doing real work like memory allocation or I/O, it belongs outside the masked region.

Want more free Linux kernel and device driver lessons?
Explore the Free Course

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *