← Previous Lecture Next Lecture →
In this lecture of our free linux device drivers course we look at a subtle spinlock interrupt handler race condition that catches many first-time driver authors off guard: your critical section is protected by a spinlock, and yet a race still happens — because the thing racing with you isn’t another thread at all. It’s your own hardware interrupt handler.
Prerequisites
- Comfortable with basic spinlock usage (spin_lock/spin_unlock) from earlier lectures
- Basic idea of what a hardware interrupt handler is in a Linux driver
- Kernel 6.x build environment for the code examples
What You Will Learn
- Why hardware interrupts are different from every other kind of concurrency we’ve discussed so far
- Three scenarios for classifying whether an interrupt handler is actually dangerous to your driver
- Why a mutex can never be used inside a hardirq handler
- Why a plain
spin_lock()is not safe when the same lock is also taken inside an interrupt handler - A first look at why
spin_lock_irqsave()exists (full API detail is the topic of the next lecture)
The New Kind of Concurrency: Hardware Interrupts
So far in this series, the competing execution paths we’ve protected against were other threads or processes, running on the same CPU or a different one, cooperatively scheduled by the kernel. A hardware interrupt is a different animal entirely: when the device asserts its interrupt line, the CPU stops whatever it was doing — instantly, without asking permission — and jumps straight into your interrupt handler. It doesn’t matter if your driver’s read or write method was in the middle of a critical section; the interrupt simply preempts it.
This raises an important design question every driver author must answer honestly: does the interrupt handler touch any of the same data that the rest of the driver touches?
Three Scenarios: Is There Actually a Race?
It helps to sort every driver into one of three buckets before reaching for a lock:
If your interrupt handler works entirely with variables local to itself, there is nothing to protect. It runs, finishes quickly, and control returns to whatever it interrupted. No shared state, no race.
The handler does touch global writeable data, but it’s a different piece of state than the one your read/write method is using. You still need a lock around the handler’s own critical section — just not the same lock, and not a conflict with the other method.
This is the dangerous case. Both the interrupt handler and, say, your read() method touch the same in-memory structure. A race is now not just possible but guaranteed to happen eventually, and both code paths must use the same lock to protect that structure — otherwise the lock protects nothing.
Rule One: No Mutexes in Interrupt Context
Once you’re in scenario three, the obvious question is: mutex or spinlock? The answer is immediate — a mutex can put the calling task to sleep while it waits, and hardware interrupt handlers run in atomic context, where sleeping is forbidden (this is exactly the “scheduling while atomic” bug we diagnosed in the previous lecture). That leaves only one option for protecting data shared with a hardirq handler: a spinlock.
Rule Two: The Same Lock, Everywhere
A lock only protects the data it actually guards at every access point. If your read() method takes spin_lock(&my_lock) before touching a shared counter, but the interrupt handler updates that same counter without ever taking my_lock, you have written code that looks safe and is not safe at all.
Where a Plain Spinlock Still Falls Short
Here’s the trap that catches even careful driver authors. Suppose a driver’s read() method does exactly the right thing on paper:
/* Simplified sketch - not yet interrupt-safe! */
static int ep_ctx_read(struct file *filp, char __user *buf,
size_t count, loff_t *off)
{
spin_lock(&ep_ctx.lock);
/* ... work with the shared ep_ctx fields ... */
spin_unlock(&ep_ctx.lock);
return count;
}
Now picture this exact sequence of events on a single CPU:
Notice what actually caused the deadlock: it wasn’t a second CPU racing in from outside. It was this exact CPU being interrupted mid-critical-section and then trying to reacquire the very lock it already holds. A plain spin_lock() has no defence against this at all, because it never disables the one thing that can preempt it on the same core — the hardware interrupt itself.
The Fix, in One Sentence (Details Next Lecture)
The kernel solves this with a family of interrupt-disabling spinlock APIs — spin_lock_irq() and spin_lock_irqsave() — which turn off interrupts on the local CPU for the duration of the critical section, so the scenario above simply cannot occur. We’ll cover their exact usage, and the difference between the two, in full detail in the next lecture. For now, the important takeaway is why they exist: plain spin_lock() is only safe when you are certain no interrupt handler on the same CPU can ever try to take the same lock.
| Situation | Safe Lock Choice |
|---|---|
| Two kernel threads, no interrupt handler involved | Plain spin_lock() / spin_unlock() |
| Data also touched by a hardirq handler on the same CPU | spin_lock_irqsave() / spin_unlock_irqrestore() |
| Data that may sleep while being processed | mutex (never inside a hardirq handler) |
We’ll unpack spin_lock_irq() vs. spin_lock_irqsave(), explain exactly what “saving flags” means, and build an original interrupt-safe driver example together.
FAQ
Q1. Can I ever use a mutex inside a hardware interrupt handler?
No, never. Hardirq handlers run in atomic context, and a mutex can put the caller to sleep, which is forbidden there.
Q2. If my interrupt handler only uses local variables, do I need any lock at all?
No. Local variables are private to that invocation and can’t be raced by definition.
Q3. Is a spinlock always enough to protect data shared with an interrupt handler?
Not by itself. If the handler can interrupt the same CPU that’s holding the lock, you also need the interrupt-disabling variant, spin_lock_irqsave(), which we cover next.
Q4. Why does the CPU let an interrupt preempt code inside a spinlock’s critical section?
Hardware interrupts are, by design, allowed to preempt almost anything so the system can respond quickly to real-world events. Software must explicitly ask to disable them when that’s unsafe.
Q5. Does this deadlock only happen on single-CPU systems?
No. It’s guaranteed on the same CPU that holds the lock, but on a multi-core system an interrupt on a different core can also spin indefinitely if it needs the same lock, so the same fix applies everywhere.
Q6. Is this the same kind of deadlock as calling schedule() under a spinlock?
They’re related but distinct bugs. Sleeping under a spinlock is a “scheduling while atomic” bug; a same-CPU interrupt reacquiring a held lock is a self-deadlock via interrupt preemption. Both stem from not fully respecting atomic-context rules.
Part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.
Browse the Full Course
2 Comments