Spinlocks and Interrupt Context in Linux Device Drivers
A hands-on lesson from our free Linux device drivers course, updated for kernel 6.12 and later
Locking is where most Linux device driver bugs are born, and interrupt handling is where locking gets genuinely tricky. If you’re following our free Linux device drivers course, this lecture picks up right where basic mutex and spinlock usage leaves off: what happens when a hardware interrupt fires while your driver is already holding a lock? This single question is the reason spinlocks exist, and understanding it properly is a core skill for anyone building a free embedded systems course style project or working on real kernel code.
Prerequisites
This lecture assumes you have already worked through the earlier modules of our free Linux kernel development course covering basic kernel modules, printk, and a first look at mutexes and spinlocks. You should be comfortable compiling and loading a kernel module on a Linux VM (kernel 6.6 LTS or newer is recommended, and everything here has been checked against 6.12).
Why a Mutex Cannot Protect an Interrupt Handler
A mutex is a sleeping lock. If the lock is busy, the calling thread is put to sleep by the scheduler and woken up later when the lock becomes free. That “going to sleep” step is the entire problem. Interrupt handlers run in a special execution context called atomic context (also referred to as interrupt context), where the CPU cannot be scheduled away. There is no “current process” in the usual sense while a hardware interrupt is being serviced, so there is nothing for the scheduler to switch to and back from.
If code inside an interrupt handler tries to take a mutex that is already held, the kernel would need to put the interrupt handler to sleep. That is simply not allowed. On a properly configured kernel this triggers a loud diagnostic; on a kernel without the right debug options it can silently corrupt system state or hang the machine. This is why the kernel restricts you to non-sleeping locks such as spinlocks whenever code might run in interrupt context.
The Race Condition: Process Context Meets Hardirq Context
Picture a driver that maintains a shared counter, updated both from a read() system call (process context) and from the device’s hardware interrupt handler (hardirq context). Both paths touch the same memory. If the interrupt fires at the exact moment the read() path is in the middle of updating that counter, you get a classic data race — unless both sides serialize their access with the same lock.
Here is that timeline shown as an inline HTML diagram rather than a text sketch, so you can see exactly when the danger window opens:
If the read() path used spin_lock() only, the interrupt handler would spin forever on the same CPU waiting for a lock that read() can never release — because read() itself is stuck, interrupted, and cannot resume until the interrupt handler finishes. This is the classic single-CPU deadlock case.
Notice the trap: if both sides simply call spin_lock() / spin_unlock(), and the interrupt happens to land on the very CPU that is inside the critical section, the system deadlocks. The interrupt handler spins waiting for a lock that the interrupted process context can never release, since it won’t run again until the interrupt handler returns. Solving this is the whole point of the irq-safe locking APIs.
spin_lock_irqsave(): The Correct API for This Case
The kernel gives you a family of spinlock variants specifically for code paths that are also touched by an interrupt handler on the same lock:
| API | Disables local IRQs? | When to use |
|---|---|---|
spin_lock() |
No | Data is never touched by a hardirq or softirq handler |
spin_lock_bh() |
No (disables softirqs) | Data is shared with a softirq/tasklet, not a hardirq |
spin_lock_irq() |
Yes, unconditionally | You know for certain IRQs were enabled before the call |
spin_lock_irqsave() |
Yes, and saves prior state | General-purpose, safe default when unsure of caller state |
spin_lock_irqsave() is the one to reach for by default, because it remembers whether interrupts were already disabled by the caller and restores that exact state on unlock, instead of blindly re-enabling interrupts that might have been off for a good reason.
#include <linux/spinlock.h>
static DEFINE_SPINLOCK(gctx_lock);
static int gCtx;
/* Called from the driver's read() file operation, process context */
static ssize_t mydrv_read(struct file *filp, char __user *buf,
size_t count, loff_t *pos)
{
unsigned long flags;
int local_val;
spin_lock_irqsave(&gctx_lock, flags);
local_val = gCtx;
spin_unlock_irqrestore(&gctx_lock, flags);
return simple_read_from_buffer(buf, count, pos, &local_val,
sizeof(local_val));
}
/* Hardware interrupt handler, atomic context */
static irqreturn_t mydrv_isr(int irq, void *dev_id)
{
unsigned long flags;
spin_lock_irqsave(&gctx_lock, flags);
gCtx++;
spin_unlock_irqrestore(&gctx_lock, flags);
return IRQ_HANDLED;
}
Because gctx_lock is taken with interrupts disabled on the local CPU in both paths, the two critical sections can never overlap on the same core, and the deadlock scenario from the diagram above is closed off.
Threaded Interrupts Change the Picture on Modern Kernels
One thing that has genuinely changed since older kernel textbooks were written: modern kernels lean heavily on threaded interrupt handlers (registered with request_threaded_irq()), and with a real-time (PREEMPT_RT) kernel — mainlined as of kernel 6.12 — the “hard” part of an interrupt handler is kept as short as physically possible, with the bulk of the work moved into a kernel thread that runs in process context. In that thread, spinlocks in a PREEMPT_RT kernel actually behave like sleeping locks internally, but from a driver author’s point of view you still write the same spin_lock_irqsave() / spin_unlock_irqrestore() calls; the correctness rules taught in this lecture remain unchanged, which is exactly why they are worth learning properly in a free linux kernel development course rather than memorizing as trivia.
Real-World Use Cases
- Network interface drivers — packet counters and ring-buffer indices updated both by the NAPI poll routine and the hardware IRQ.
- Character device drivers — a status register cache updated by both user-space triggered I/O and an interrupt-driven event.
- Sensor drivers (I2C/SPI over GPIO interrupts) — a data-ready flag toggled by a GPIO interrupt and read from a sysfs or read() handler.
Common Mistakes and Troubleshooting
spin_lock() in a function also called from an interrupt handler. This opens the exact single-CPU deadlock shown in the diagram above.
spin_lock_bh() over spin_lock_irqsave() — it is cheaper because it only disables softirqs, not all hardware interrupts.
Best Practices for Locking Around Interrupts
- Keep the critical section as short as possible; never do I/O, allocation, or logging inside it.
- Always pair
spin_lock_irqsave()withspin_unlock_irqrestore()using the exact sameflagsvariable. - Use the same lock on every code path that touches the shared data, including the interrupt handler.
- Prefer per-CPU data structures over shared globals when the algorithm allows it, to avoid contention entirely.
Performance Considerations
Disabling local interrupts is not free — while they are off, the CPU cannot respond to timer ticks or other devices, which increases interrupt latency system-wide. Keep the locked region short, and prefer spin_lock_bh() or plain spin_lock() whenever the stronger irqsave variant is not strictly required, since each level of protection carries a small additional cost.
Security Considerations
A missed lock around interrupt-shared data is not just a correctness bug; on some subsystems a corrupted shared structure can be leveraged for privilege escalation or a denial-of-service crash. Treat this class of race condition with the same seriousness as a memory-safety bug when auditing driver code.
Summary and Key Takeaways
- Interrupt handlers run in atomic context and can never sleep, so mutexes are off-limits there.
- Data shared between process context and a hardirq handler needs an irq-safe spinlock, not a plain one.
spin_lock_irqsave()/spin_unlock_irqrestore()is the safe general-purpose choice.- On modern PREEMPT_RT kernels the API stays the same even though the internal behaviour differs.
Conclusion
Getting interrupt-safe locking right is one of the clearest markers of a driver author who understands the kernel rather than one who is just copying patterns. Once you’re comfortable with the timeline in this lecture — process context, the interrupt firing mid-critical-section, and the irqsave APIs that close that window — you have the foundation needed for the next module of this free Linux device drivers course, where we look at what happens when someone gets this wrong and how to catch it with modern debugging tools.
Frequently Asked Questions
Q1. Can I use a mutex inside a threaded interrupt handler?
Yes. A threaded IRQ handler runs in process context (its own kernel thread), so mutexes are legal there — the restriction only applies to the “hard” top-half handler.
Q2. What’s the difference between spin_lock_irq() and spin_lock_irqsave()?
spin_lock_irq() always enables interrupts on unlock, assuming they were on before. spin_lock_irqsave() records the actual prior state and restores exactly that, making it safe to call from code whose caller state is unknown.
Q3. Does spin_lock_bh() disable hardware interrupts?
No, it only disables software interrupts (softirqs/tasklets) on the local CPU, leaving hardirqs enabled. Use it only when your data is never touched directly by a hardirq handler.
Q4. Is this locking behaviour different on a multi-core system?
spin_lock_irqsave() disables interrupts only on the local CPU and additionally spins if another CPU holds the lock, so it protects against both the single-CPU reentrancy problem and true multi-core races.
Q5. Why did older tutorials use spin_lock() everywhere without irqsave?
Many older examples targeted code that was never actually shared with an interrupt handler. The irqsave variant is only required when a hardirq handler can touch the same lock.
Q6. Do I need irqsave locking with PREEMPT_RT enabled?
Yes — the API and the correctness requirement are unchanged. What changes internally is how the kernel implements the primitive, not whether you need it.

2 Comments