Spinlocks vs Hardware Interrupts: The Self-Deadlock Nobody Warns You About-Free Linux Kernel Development Course

Spinlocks vs Hardware Interrupts: The Self-Deadlock Nobody Warns You About
Free Linux Kernel Development Course — Kernel Synchronization Series
spin_lock_irq linux kernel free linux kernel development course free linux device drivers course free embedded systems course

If you’ve been following this free Linux kernel development course, you already know that a plain spin_lock() protects a critical section against another CPU core touching the same data. But there’s a scenario where a plain spinlock doesn’t just fail to protect you — it actively locks up your entire CPU core. This happens the moment a hardware interrupt handler tries to grab the same lock that a process-context function is already holding. That’s the problem spin_lock_irq() exists to solve, and it’s exactly what we’ll build and test in this lecture.

What You Will Learn

  • Why a hardirq handler and a “normal” driver function racing on the same spinlock can deadlock a single CPU core
  • How the same scenario behaves differently on a single-core (UP) system versus a multicore (SMP) system
  • How spin_lock_irq() / spin_unlock_irq() close this gap by disabling local interrupts
  • A working, original kernel 6.x driver example showing the buggy version and the fixed version side by side

Prerequisites

  • Comfortable with basic spinlock usage: spin_lock() / spin_unlock()
  • Basic idea of what a hardware interrupt handler (hardirq) is
  • A kernel 6.x build environment (kernel headers installed, ability to build and load a module)

The Setup: One Lock, Two Callers

Picture a character device driver with one shared, mutable structure. Two different code paths touch it:

  • A read() method, which runs in normal process context whenever user space calls read() on the device
  • An interrupt handler, registered with request_irq(), which fires whenever the underlying hardware peripheral raises an interrupt

Both paths take the same spinlock before touching the shared structure. On the surface that looks correct — and if the interrupt happens to fire while the read() method isn’t holding the lock, everything works fine. The trouble starts when the timing isn’t so cooperative.

Timeline: Interrupt Fires While the Lock Is Already Held
read() method: spin_lock() ── [critical section running] ── spin_unlock()
hardware IRQ fires here ↑ handler calls spin_lock() → spins forever, same core

On a single-core (UP) system this is fatal. The interrupt handler runs on the same core that is currently executing the read() method’s critical section. The read() method can’t finish and release the lock, because the CPU has jumped into the interrupt handler and won’t return until the handler completes. The handler can’t finish either, because it’s stuck spinning on a lock that will never be released. That’s a genuine self-deadlock — the driver has locked up its own CPU core.

On a multicore (SMP) system the story is a little better, but only a little. If the interrupt lands on a different core than the one running the read() method, the interrupt handler will spin on that other core until the read() method finishes and unlocks — wasteful, but not fatal. The real danger is that interrupts are not guaranteed to land on any particular core, so you can’t rely on SMP behaviour to save you. Any production driver has to assume the worst case: the same core.

The Fix: spin_lock_irq() / spin_unlock_irq()

The kernel provides an interrupt-aware variant of the spinlock API specifically for this situation:

#include <linux/spinlock.h>

void spin_lock_irq(spinlock_t *lock);
void spin_unlock_irq(spinlock_t *lock);

spin_lock_irq() does two things in one call: it disables hardware interrupts on the local CPU core, and then it takes the spinlock. Because interrupts are masked on that core for the duration of the critical section, the interrupt handler simply cannot run on that core until spin_unlock_irq() re-enables interrupts and releases the lock. The self-deadlock scenario becomes structurally impossible — not just unlikely, impossible.

It’s worth being precise about what “disables interrupts” means here: it affects only the local core the code is currently running on. If your system is SMP and the interrupt happens to fire on a different core, ordinary spinlock behaviour still applies there — that core spins until the lock is released, exactly as before.

Original Demo: Buggy vs Fixed, Kernel 6.x

The example below is an original teaching driver — not copied from any book or existing project. It’s deliberately simplified: instead of real hardware, we use a periodic hrtimer callback to stand in for a hardware interrupt handler, so you can reproduce the scenario on any machine without needing actual interrupt-capable hardware. The logic is identical to a real IRQ handler racing a process-context function.

/* ep_irq_race_demo.c — teaching module, kernel 6.x
 * Simulates: a "read path" holding a spinlock while a
 * timer callback (standing in for a hardirq handler) tries
 * to take the same lock.
 */
#include <linux/module.h>
#include <linux/spinlock.h>
#include <linux/hrtimer.h>
#include <linux/ktime.h>
#include <linux/delay.h>

static DEFINE_SPINLOCK(ep_slock);
static struct hrtimer ep_htimer;
static int ep_shared_counter;

/* Stand-in for the hardirq handler */
static enum hrtimer_restart ep_timer_cb(struct hrtimer *t)
{
    /* BUGGY version: plain spin_lock(), no irq masking */
    spin_lock(&ep_slock);
    ep_shared_counter++;
    spin_unlock(&ep_slock);

    hrtimer_forward_now(t, ms_to_ktime(5));
    return HRTIMER_RESTART;
}

/* Stand-in for the driver's read() method, holding the lock
 * long enough that the "interrupt" is likely to land inside it.
 */
static void ep_simulated_read_path(void)
{
    spin_lock(&ep_slock);      /* BUGGY: should be spin_lock_irq() */
    ep_shared_counter += 10;
    udelay(50);                /* simulate work inside critical section */
    spin_unlock(&ep_slock);
}

static int __init ep_irq_race_init(void)
{
    hrtimer_init(&ep_htimer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
    ep_htimer.function = ep_timer_cb;
    hrtimer_start(&ep_htimer, ms_to_ktime(5), HRTIMER_MODE_REL);

    ep_simulated_read_path();
    pr_info("ep_irq_race_demo: loaded (buggy variant)\n");
    return 0;
}

static void __exit ep_irq_race_exit(void)
{
    hrtimer_cancel(&ep_htimer);
    pr_info("ep_irq_race_demo: unloaded, counter=%d\n", ep_shared_counter);
}

module_init(ep_irq_race_init);
module_exit(ep_irq_race_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala: spinlock vs interrupt race demo");

The fixed version changes exactly one thing — the lock/unlock pair used by the process-context path — while the timer callback is left as-is (in a real driver, this would be your genuine interrupt handler, which should also generally avoid taking non-irq-safe locks that a process context could be holding):

/* CORRECT: process-context path now masks local interrupts */
static void ep_simulated_read_path(void)
{
    spin_lock_irq(&ep_slock);   /* disables local IRQs + takes lock */
    ep_shared_counter += 10;
    udelay(50);
    spin_unlock_irq(&ep_slock); /* releases lock + re-enables IRQs */
}

With this change, while ep_simulated_read_path() holds the lock, hardware interrupts cannot fire on that core at all — so the “interrupt” (our hrtimer callback, in this simplified demo) simply cannot interrupt the critical section on that core. It either runs to completion beforehand or gets deferred until immediately after spin_unlock_irq() re-enables interrupts.

Buggy vs Fixed: Quick Comparison

Aspectspin_lock() (buggy)spin_lock_irq() (fixed)
Local interrupts during critical sectionStay enabledDisabled
Same-core hardirq handler taking the same lockSelf-deadlockCannot run until unlock
Different-core hardirq handler (SMP)Spins, then proceeds normallySpins, then proceeds normally
Safe to use from a process-context path that a hardirq also locks?NoYes

Frequently Asked Questions

Q1. Does spin_lock_irq() disable interrupts on every CPU core, or just one?
Only the local core — the one currently executing the code that calls it. Interrupts on other cores are unaffected.

Q2. Is it safe to call spin_lock_irq() from inside an interrupt handler itself?
It’s unusual and generally discouraged; interrupt handlers typically use the plain spin_lock()/spin_unlock() pair since interrupts are already effectively disabled for hardirq context on that core.

Q3. What happens to kernel preemption when spin_lock_irq() runs?
Disabling hardware interrupts on the local core has the side effect of disabling kernel preemption there too, for as long as the lock is held.

Q4. Do I always need the _irq variant, even on a single-core system?
Yes. The self-deadlock scenario applies to UP systems just as much as SMP — arguably more directly, since there’s only one core for the interrupt to land on.

Q5. Does spin_lock_irq() also block softirqs and tasklets?
Yes — disabling hardware interrupts on the local core also prevents softirqs and tasklets from running there, since they’re driven from the interrupt/exit path.

Q6. Is there a downside to using spin_lock_irq() everywhere “just to be safe”?
Yes — every microsecond spent with local interrupts disabled is time-sensitive interrupts on that core cannot be serviced. Use the plain spin_lock() when you know for certain no interrupt handler touches the same lock, and reserve the _irq variant for genuine interrupt-racing cases.

Q7. What’s the difference between this and spin_lock_irqsave()?
spin_lock_irq() unconditionally enables interrupts again on unlock. spin_lock_irqsave() preserves whatever the previous interrupt mask state was and restores exactly that. We cover this important distinction, and why it matters, in the next lecture.

Coming Up Next

spin_lock_irq() assumes it’s safe to fully re-enable all interrupts when it unlocks — but what if another part of the kernel had deliberately masked some interrupts before your code ever ran? In the next lecture we look at exactly that trap, and the API pair that avoids it: spin_lock_irqsave() and spin_unlock_irqrestore().

Continue the Free Linux Kernel Development Course

More kernel synchronization, device driver, and embedded Linux lectures at EmbeddedPathashala.

Explore the Course

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *