Kernel Synchronization Part 2 Recap: Atomics, RCU-Adjacent Locks, Per-CPU Data & Lockdep (Kernel 6.x)-Free Linux Device Driver Course Online

Kernel Synchronization Part 2 Recap: Atomics, RCU-Adjacent Locks, Per-CPU Data & Lockdep (Kernel 6.x)

← Previous Lecture  |  Next Lecture →

Kernel Synchronization Part 2: Chapter Recap
Free Linux Kernel Programming Course — Kernel 6.x Edition

This lecture closes out the Kernel Synchronization Part 2 chapter of our free linux kernel programming course. Instead of new material, we tie together atomic operations, reader-writer locks, per-CPU data, lockdep debugging and memory barriers into one map, so you can see how each tool fits into a working free linux device drivers course project.

atomic_t / refcount_t rwlock_t / rw_semaphore per-CPU variables lockdep memory barriers
What You Will Learn
A one-page map of every locking tool covered A decision guide for picking the right tool Original practice exercises to self-test
Prerequisites

This recap assumes you have worked through the full Kernel Synchronization Part 2 chapter of this course: atomic integers, RMW bitops, reader-writer locks, false sharing, per-CPU variables, lockdep, lock debugging, and memory barriers.

The Chapter in One Diagram

Every mechanism in this chapter answers the same underlying question — “how do I let multiple CPUs touch shared state safely, without needlessly serializing everything?” — but each one trades off differently between simplicity, performance, and how much concurrency it actually allows.

From Cheapest to Most Powerful
1. Per-CPU variables → no sharing at all, near-zero cost
2. atomic_t / refcount_t → single-integer, lock-free RMW
3. RMW bitops (set_bit / test_and_set_bit) → lock-free flag words
4. rwlock_t / rw_semaphore → many readers, one writer
5. spinlock_t / mutex → full critical section, general purpose
6. Memory barriers → ordering without any lock at all

Topic-by-Topic Recap

Atomic Integers: atomic_t, atomic64_t, refcount_t

atomic_t and atomic64_t give you lock-free, race-free read-modify-write operations on a single integer, useful for counters and statistics. refcount_t is a stricter, saturating variant purpose-built for object lifetime reference counting, and it deliberately warns loudly on misuse such as incrementing from zero.

RMW Bitwise Operators

set_bit(), clear_bit(), and the test_and_set_bit() family let you manage individual flag bits inside a shared word without a spinlock, and they are guaranteed atomic across every CPU core, not just the local one.

Reader-Writer Locks: rwlock_t and rw_semaphore

Both allow many concurrent readers alongside a single exclusive writer. rwlock_t spins and is safe in atomic context; rw_semaphore sleeps and is only safe in process context. Both share a writer-starvation risk under heavy read pressure, which is why newer read-mostly designs increasingly reach for seqlock_t or RCU instead.

False Sharing and Per-CPU Variables

False sharing happens when unrelated variables land on the same CPU cache line, causing needless cache-coherency traffic between cores. Cache-line alignment fixes the symptom; per-CPU variables remove the sharing at the root by giving each core its own private copy, accessed through this_cpu_* operations or get_cpu_ptr()/put_cpu_ptr() for dynamic allocations.

Lockdep and Lock Debugging

lockdep proves locking correctness mathematically from recorded lock-ordering history, rather than depending on a bug actually occurring during a test run. CONFIG_DEBUG_ATOMIC_SLEEP, CONFIG_DEBUG_MUTEXES and CONFIG_LOCK_STAT round out the debug-kernel toolbox for catching scheduling-while-atomic bugs and lock contention hotspots.

Memory Barriers

rmb(), wmb() and mb() order memory operations for agents like DMA-capable hardware; the smp_ variants order operations only between CPU cores. READ_ONCE() and WRITE_ONCE() protect a single shared variable’s individual access from unwanted compiler optimization.

Decision Guide: Which Tool Should I Reach For?

SituationRecommended tool
A simple counter, no other fields involvedatomic_t / atomic64_t
Object lifetime, get/put reference countingrefcount_t
A handful of independent status flagsAtomic bitops (set_bit family)
Read-mostly data, occasional writer, process contextrw_semaphore
Read-mostly data, needed in atomic contextrwlock_t (or seqlock_t/RCU for heavy read load)
Multiple related fields updated together, may sleepmutex
Multiple related fields, atomic/interrupt contextspinlock_t (with _irqsave as needed)
Per-core statistics, no cross-core reads neededPer-CPU variables
Hardware/DMA write ordering, no mutual exclusion neededwmb() / rmb() / mb()

Practice Exercises

  1. Convert a plain int open-count variable in a misc driver to refcount_t, and add correct WARN-worthy misuse handling.
  2. Take a driver context struct with four independent boolean flags and replace them with a single flag word managed via set_bit()/test_and_set_bit().
  3. Benchmark a per-CPU counter against a single spinlock-protected counter under four concurrent kthreads, using ktime_get_ns().
  4. Deliberately introduce an AB-BA deadlock between two mutexes in a test module, and use lockdep’s captured chains to identify the exact ordering violation.
  5. Build a two-thread producer/consumer flag handoff using WRITE_ONCE()/READ_ONCE() plus wmb()/rmb(), and verify it under a stress-test loop.

Frequently Asked Questions

Do I need to memorize every API in this chapter?

No. The decision guide table above is meant to be the thing you actually keep handy. Recognizing which category a problem falls into matters more than recalling every function signature from memory.

Is RCU covered in this chapter?

RCU was mentioned as a modern alternative to writer-starved rwlock_t/rw_semaphore designs, but the full RCU API is deep enough to deserve its own dedicated chapter later in this course.

Why does the course prefer lockdep over just “being careful”?

Deadlock bugs are timing-dependent and can pass thousands of test runs before appearing in production. lockdep detects the possibility of a deadlock from recorded lock-ordering history, catching bugs that would otherwise hide until a rare unlucky interleaving occurs.

What is the single most common mistake this chapter warns about?

Reading multiple related fields of a shared structure without a lock and assuming the read is atomic as a whole — this “torn read” problem, plus its close cousin “scheduling while atomic,” come up repeatedly across the chapter’s demo drivers.

What comes after this chapter?

The next chapter moves into RCU (Read-Copy-Update) and further lock-free techniques, building directly on the atomic and per-CPU foundations from this chapter.

Continue your free linux kernel development course journey with the next chapter of this series.

Next Lecture Previous Lecture

← Previous Lecture  |  Next Lecture →

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *