← Previous Lecture | Next Lecture →
This lecture closes out the Kernel Synchronization Part 2 chapter of our free linux kernel programming course. Instead of new material, we tie together atomic operations, reader-writer locks, per-CPU data, lockdep debugging and memory barriers into one map, so you can see how each tool fits into a working free linux device drivers course project.
This recap assumes you have worked through the full Kernel Synchronization Part 2 chapter of this course: atomic integers, RMW bitops, reader-writer locks, false sharing, per-CPU variables, lockdep, lock debugging, and memory barriers.
The Chapter in One Diagram
Every mechanism in this chapter answers the same underlying question — “how do I let multiple CPUs touch shared state safely, without needlessly serializing everything?” — but each one trades off differently between simplicity, performance, and how much concurrency it actually allows.
Topic-by-Topic Recap
atomic_t and atomic64_t give you lock-free, race-free read-modify-write operations on a single integer, useful for counters and statistics. refcount_t is a stricter, saturating variant purpose-built for object lifetime reference counting, and it deliberately warns loudly on misuse such as incrementing from zero.
set_bit(), clear_bit(), and the test_and_set_bit() family let you manage individual flag bits inside a shared word without a spinlock, and they are guaranteed atomic across every CPU core, not just the local one.
Both allow many concurrent readers alongside a single exclusive writer. rwlock_t spins and is safe in atomic context; rw_semaphore sleeps and is only safe in process context. Both share a writer-starvation risk under heavy read pressure, which is why newer read-mostly designs increasingly reach for seqlock_t or RCU instead.
False sharing happens when unrelated variables land on the same CPU cache line, causing needless cache-coherency traffic between cores. Cache-line alignment fixes the symptom; per-CPU variables remove the sharing at the root by giving each core its own private copy, accessed through this_cpu_* operations or get_cpu_ptr()/put_cpu_ptr() for dynamic allocations.
lockdep proves locking correctness mathematically from recorded lock-ordering history, rather than depending on a bug actually occurring during a test run. CONFIG_DEBUG_ATOMIC_SLEEP, CONFIG_DEBUG_MUTEXES and CONFIG_LOCK_STAT round out the debug-kernel toolbox for catching scheduling-while-atomic bugs and lock contention hotspots.
rmb(), wmb() and mb() order memory operations for agents like DMA-capable hardware; the smp_ variants order operations only between CPU cores. READ_ONCE() and WRITE_ONCE() protect a single shared variable’s individual access from unwanted compiler optimization.
Decision Guide: Which Tool Should I Reach For?
| Situation | Recommended tool |
|---|---|
| A simple counter, no other fields involved | atomic_t / atomic64_t |
| Object lifetime, get/put reference counting | refcount_t |
| A handful of independent status flags | Atomic bitops (set_bit family) |
| Read-mostly data, occasional writer, process context | rw_semaphore |
| Read-mostly data, needed in atomic context | rwlock_t (or seqlock_t/RCU for heavy read load) |
| Multiple related fields updated together, may sleep | mutex |
| Multiple related fields, atomic/interrupt context | spinlock_t (with _irqsave as needed) |
| Per-core statistics, no cross-core reads needed | Per-CPU variables |
| Hardware/DMA write ordering, no mutual exclusion needed | wmb() / rmb() / mb() |
Practice Exercises
- Convert a plain
intopen-count variable in a misc driver torefcount_t, and add correct WARN-worthy misuse handling. - Take a driver context struct with four independent boolean flags and replace them with a single flag word managed via
set_bit()/test_and_set_bit(). - Benchmark a per-CPU counter against a single spinlock-protected counter under four concurrent kthreads, using
ktime_get_ns(). - Deliberately introduce an AB-BA deadlock between two mutexes in a test module, and use lockdep’s captured chains to identify the exact ordering violation.
- Build a two-thread producer/consumer flag handoff using
WRITE_ONCE()/READ_ONCE()pluswmb()/rmb(), and verify it under a stress-test loop.
Frequently Asked Questions
Do I need to memorize every API in this chapter?
No. The decision guide table above is meant to be the thing you actually keep handy. Recognizing which category a problem falls into matters more than recalling every function signature from memory.
Is RCU covered in this chapter?
RCU was mentioned as a modern alternative to writer-starved rwlock_t/rw_semaphore designs, but the full RCU API is deep enough to deserve its own dedicated chapter later in this course.
Why does the course prefer lockdep over just “being careful”?
Deadlock bugs are timing-dependent and can pass thousands of test runs before appearing in production. lockdep detects the possibility of a deadlock from recorded lock-ordering history, catching bugs that would otherwise hide until a rare unlucky interleaving occurs.
What is the single most common mistake this chapter warns about?
Reading multiple related fields of a shared structure without a lock and assuming the read is atomic as a whole — this “torn read” problem, plus its close cousin “scheduling while atomic,” come up repeatedly across the chapter’s demo drivers.
What comes after this chapter?
The next chapter moves into RCU (Read-Copy-Update) and further lock-free techniques, building directly on the atomic and per-CPU foundations from this chapter.
Continue your free linux kernel development course journey with the next chapter of this series.
Next Lecture Previous Lecture
1 Comment