Intermediate
6.x (mainline)
~14 minutes
If you have been following this free Linux kernel programming course, you already know that picking the right linux kernel spinlock types is one of the most practical skills a driver author can build. Across this chapter we built up spinlocks piece by piece — the plain form, the IRQ-disabling form, and the save/restore form. In this lecture we tie it all together into one clean decision guide, look briefly at what happens under the hood on x86_64 and ARM, and close out the “Kernel Synchronization Part 1” chapter with a recap you can bookmark and reuse.
This lecture deliberately avoids repeating old benchmark numbers from any particular book edition — kernel internals and typical usage patterns shift release to release, so instead we focus on the decision logic itself, which stays valid across kernel versions.
- How to choose between the three spinlock forms in real driver code
- Why spinlocks behave differently on a single-core (UP) system vs a multicore (SMP) system
- A simplified look at how spinlocks are implemented at the CPU instruction level
- What “local locks” are and why the real-time Linux project introduced them
- A full recap of every locking concept covered in this chapter
- Original practice exercises to test your understanding
This lecture assumes you have already gone through the earlier lectures in this chapter covering critical sections, atomicity, mutexes, deadlocks, and the basic spinlock API (spin_lock(), spin_lock_irq(), spin_lock_irqsave(), spin_lock_bh()). Comfort writing a basic loadable kernel module is also expected.
Choosing the Right Spinlock: A Three-Tier Decision Guide
Every spinlock decision in the kernel boils down to one question: can an interrupt handler or a softirq touch the same data as this critical section? Once you answer that, the correct spinlock variant almost picks itself. Here is the decision guide in one place:
| Situation | Recommended API | Relative Overhead |
|---|---|---|
| No interrupt handler ever touches this data | spin_lock() / spin_unlock() |
Lowest |
| A hardirq handler touches this data, and you don’t already know the interrupt state | spin_lock_irq() / spin_unlock_irq() |
Medium |
| The calling context’s interrupt state is unknown or could vary (safest general-purpose choice) | spin_lock_irqsave() / spin_unlock_irqrestore() |
Highest (but safest) |
Process context only
Disables local IRQs
Saves + restores IRQ state
A useful mental shortcut for this free Linux device drivers course: when in doubt, reach for spin_lock_irqsave() first. It is never wrong to be more cautious than necessary — it is only wrong to be less cautious than necessary.
A Simplified Look at Spinlock Internals
You never need to touch this layer directly, but understanding it removes a lot of mystery. A spinlock is, at its core, a single memory location that a CPU tries to flip from “free” to “held” using one indivisible instruction.
- x86_64: The kernel relies on an atomic compare-and-swap style instruction so that reading the lock’s current state and writing the new state happen as one uninterruptible step. No other core can see a half-finished update.
- ARM64: Instead of constantly re-reading memory in a tight loop, ARM cores can use a “wait for event” style instruction to sleep lightly until another core signals that the lock changed state, then wake up and retry. This is more power-efficient than a pure busy-loop.
Regardless of architecture, the golden rule stays the same: as a driver author, you only ever call the exposed spin_lock*() API. The architecture-specific instruction sequence is handled for you.
UP vs SMP: Why Spinlocks Behave Differently
On a genuinely single-core (UP) system, “spinning” while waiting for a lock makes no sense — there is no second core to be running the code that will release it. So on UP kernels, the spinlock primitives fall back to simply disabling interrupts and kernel preemption on that one core, which is enough to guarantee exclusive access.
On an SMP (multicore) system — which describes essentially every device you will target today, from a Raspberry Pi to a server — the spinning behaviour is real: a second core genuinely does busy-wait until the first core finishes and releases the lock.
spin_lock() → disable IRQs/preemption on the one CPU → no real spinning happens
spin_lock() → other cores busy-wait → owner finishes → unlock signals waiting cores
The good news: you never have to write different code for UP and SMP. You simply call the same spin_lock() family of functions, and the kernel’s build configuration decides what actually happens underneath.
A Modern Addition: Local Locks in Real-Time Linux
Kernel developers eventually added a lighter-weight concept called a local lock, originating from the real-time Linux (PREEMPT_RT) project. A local lock behaves like a per-CPU critical section marker rather than a true spinning lock, and it plays especially well with lock-debugging tools such as lockdep. On a fully real-time-preemptible kernel it changes behaviour slightly to remain safe under RT scheduling. As of recent kernel 6.x releases, PREEMPT_RT support has matured into the mainline tree, so local locks are worth knowing even if you are not building a hard real-time system — many core subsystems now use them internally.
Original Code Example: A Self-Documenting Lock Choice
Here is a small original driver snippet that shows all three spinlock forms living side by side in one context structure, with comments explaining why each one was chosen — a pattern worth copying into your own drivers.
#include <linux/module.h>
#include <linux/spinlock.h>
#include <linux/interrupt.h>
struct ep_demo_ctx {
spinlock_t counter_lock; /* touched only in process context */
spinlock_t queue_lock; /* touched by our hardirq handler too */
int counter;
int queue_depth;
};
static struct ep_demo_ctx ep_ctx;
/* Process-context only: plain spin_lock() is enough */
static void ep_bump_counter(void)
{
spin_lock(&ep_ctx.counter_lock);
ep_ctx.counter++;
spin_unlock(&ep_ctx.counter_lock);
}
/* Shared with a hardirq handler: must save/restore IRQ state */
static void ep_enqueue_item(void)
{
unsigned long flags;
spin_lock_irqsave(&ep_ctx.queue_lock, flags);
ep_ctx.queue_depth++;
spin_unlock_irqrestore(&ep_ctx.queue_lock, flags);
}
/* Our hardirq handler also updates queue_depth, so it uses the
* plain spin_lock() form here because IRQs are already disabled
* by the time a hardirq handler runs. */
static irqreturn_t ep_demo_irq_handler(int irq, void *dev_id)
{
spin_lock(&ep_ctx.queue_lock);
ep_ctx.queue_depth--;
spin_unlock(&ep_ctx.queue_lock);
return IRQ_HANDLED;
}
static int __init ep_demo_init(void)
{
spin_lock_init(&ep_ctx.counter_lock);
spin_lock_init(&ep_ctx.queue_lock);
pr_info("ep_demo: spinlock decision demo loaded\n");
return 0;
}
static void __exit ep_demo_exit(void)
{
pr_info("ep_demo: unloaded\n");
}
module_init(ep_demo_init);
module_exit(ep_demo_exit);
MODULE_LICENSE("GPL");
Notice that queue_lock uses spin_lock_irqsave() from process context but plain spin_lock() inside the hardirq handler itself — because a hardirq handler already runs with local interrupts disabled, so there is nothing left to save.
Common Mistakes to Avoid
| Mistake | Why It’s a Problem |
|---|---|
Using spin_lock_irq() when the interrupt state before locking is unknown |
It unconditionally enables interrupts on unlock, silently clobbering a previously-disabled state |
| Sleeping (allocating memory, calling functions that block) while holding any spinlock | Triggers a “scheduling while atomic” bug, covered earlier in this chapter |
| Assuming UP behaviour when testing only on a single-core VM | Real races only appear on genuine SMP hardware — always test on multicore before shipping |
Best Practices
- Default to
spin_lock_irqsave()unless you can clearly justify a lighter form. - Keep critical sections as short as possible — never do logging, allocation, or user-copy operations inside a spinlock.
- Document, right next to the lock declaration, exactly which contexts (process, hardirq, softirq) touch it.
- Enable
CONFIG_DEBUG_ATOMIC_SLEEPandCONFIG_PROVE_LOCKINGon your development kernel so mistakes surface immediately instead of in production.
Performance Considerations
Every step up the spinlock ladder — from plain, to IRQ-disabling, to save/restore — adds a small but real cost. On a hot path such as a network driver’s packet handler, that cost accumulates fast. The practical rule: use the strongest form for correctness during development, then revisit whether a lighter form is provably safe once the locking discipline is settled.
Security Considerations
Holding a spinlock for too long effectively creates a denial-of-service risk on other cores, since they will busy-wait the entire time. Untrusted input that can trigger unusually long critical sections (for example, a user-controlled loop bound inside a locked section) should always be validated before entering the lock.
Chapter Recap: Kernel Synchronization Part 1
This closes out the “Kernel Synchronization Part 1” chapter of our free Linux kernel development course. Here is everything we covered, in order:
What a critical section is, why exclusive execution matters, and the difference between an atomic operation and a race-prone read-modify-write sequence.
The four real triggers for data races: SMP, a preemptible kernel, blocking I/O, and hardware interrupts.
Sleep-capable locking for process context, interruptible and killable variants, and the mutex vs semaphore distinction including priority inversion.
Lock ordering discipline, self-deadlock, AB-BA chains, and how lockdep catches these at runtime.
The non-sleeping alternative to mutexes, and the three-tier family covered again in this lecture’s summary.
How to read the bug’s call trace, and why the same bug can be silent on a non-debug production kernel.
Practice Exercises
- Write a driver where a workqueue and a hardirq handler both touch the same counter. Decide, and justify, which spinlock form each context should use.
- Modify the code example above so that
queue_lockis instead shared with a tasklet rather than a hardirq handler. Which API changes, and why? - Enable
CONFIG_DEBUG_ATOMIC_SLEEPon a test kernel and deliberately callmsleep()inside a spinlock to observe the resulting call trace. - Explain, in your own words, why a UP kernel never truly “spins” on a spinlock.
Key Takeaways
- Pick the spinlock form based on what context shares the data — process-only, hardirq-shared with known state, or hardirq-shared with unknown state.
- Spinlocks never sleep; sleeping inside one is always a bug.
- UP systems fake spinlock behaviour via interrupt/preemption disabling; SMP systems truly spin.
- Local locks are a newer, lighter-weight primitive worth knowing as PREEMPT_RT matures in mainline kernels.
Conclusion
That wraps up Kernel Synchronization Part 1 in this free Linux kernel programming course. You now have a complete, reusable decision framework for choosing between mutexes and the three spinlock variants, you understand why UP and SMP systems behave differently under the same API, and you’ve seen how the discipline from this whole chapter fits together. The next chapter continues into further synchronization primitives used throughout real device drivers.
Frequently Asked Questions
Use spin_lock_irqsave() / spin_unlock_irqrestore(). It saves and restores the exact previous interrupt state, so it can never accidentally enable or disable interrupts incorrectly.
No. Spinlocks are strictly non-sleeping. Sleeping inside one causes a “scheduling while atomic” bug, which we covered earlier in this chapter.
No. On UP systems the kernel simply disables interrupts and preemption on that one core instead of spinning, since there is no second core to wait on.
A lighter-weight, per-CPU critical-section marker introduced by the real-time Linux project, useful for both RT kernels and for lock debugging on standard kernels.
spin_lock_irq() unconditionally enables interrupts on unlock, while spin_lock_irqsave() restores whatever interrupt state existed just before the lock was taken.
Yes — this entire free Linux kernel development course, including this free Linux device drivers course chapter, is available at no cost on EmbeddedPathashala.
More lectures on device drivers, memory management, and kernel synchronization are on the way.
Previous Lecture Next Lecture
2 Comments