← Previous Lecture | Next Lecture →
In this free linux kernel programming course lecture we look at memory barriers from a practical angle: why a driver author actually needs wmb(), and how READ_ONCE() and WRITE_ONCE() protect individual shared variables. This is a hands-on continuation of our earlier lecture on memory barrier fundamentals, and it is part of our free linux device drivers course covering kernel 6.x internals.
You should already be comfortable with the material from our earlier lecture on memory barrier fundamentals (rmb/wmb/mb and the smp_ variants), plus basic spinlock and atomic_t knowledge from this chapter.
Why a Single wmb() Call Matters
Modern CPUs and compilers are free to reorder independent memory operations for performance, as long as the reordering is invisible to a single thread of execution. That freedom becomes dangerous the moment a second party — another CPU core, or a piece of hardware such as a network or storage controller — is watching the same memory region and depends on a specific write order.
A very common real pattern in driver code looks like this: fill in several fields of a shared descriptor structure, and only after every field is correctly populated, flip a “this descriptor is now valid” flag. If the CPU or compiler is allowed to reorder the flag write ahead of the field writes, the hardware could read the flag as valid while the fields are still garbage.
An Original Demo: A Software Ring Descriptor
To keep this hands-on without depending on any specific NIC hardware, here is an original, simplified example: a fixed-size ring of descriptor slots that a producer kernel thread fills in, and a consumer kernel thread drains. Real DMA-capable hardware would sit in place of the consumer, but the ordering rule is identical, and this version runs on any machine, virtual or physical.
struct ep_ring_desc {
u32 addr;
u32 len;
u32 valid; /* 0 = empty, 1 = ready to consume */
};
#define EP_RING_SLOTS 8
static struct ep_ring_desc ep_ring[EP_RING_SLOTS];
static void ep_producer_fill(unsigned int slot, u32 addr, u32 len)
{
struct ep_ring_desc *d = &ep_ring[slot];
d->addr = addr;
d->len = len;
/* Ensure addr/len are visible before valid is set */
wmb();
WRITE_ONCE(d->valid, 1);
}
static bool ep_consumer_try_take(unsigned int slot,
u32 *addr, u32 *len)
{
struct ep_ring_desc *d = &ep_ring[slot];
if (!READ_ONCE(d->valid))
return false;
/* Pair the producer's wmb() with an rmb() here */
rmb();
*addr = d->addr;
*len = d->len;
d->valid = 0;
return true;
}
Notice the pairing: the producer’s wmb() only has meaning next to the consumer’s rmb(). A write barrier with nothing on the read side to pair against gives no correctness guarantee at all — barriers must always be reasoned about as matched pairs between the writer and the reader.
READ_ONCE() and WRITE_ONCE(): Protecting a Single Variable
Full barriers like wmb() and rmb() order multiple memory accesses relative to each other. READ_ONCE() and WRITE_ONCE() solve a narrower but equally important problem: making sure the compiler does not split, merge, cache in a register, or otherwise “optimize away” a single access to a variable that another thread or piece of hardware might touch concurrently.
Without WRITE_ONCE(d->valid, 1), a sufficiently aggressive compiler could, in theory, delay the store, combine it with a neighbouring store, or keep the value in a register instead of memory for a while. WRITE_ONCE() and READ_ONCE() close that gap for individual variables, and on kernel builds with KCSAN enabled they also mark the access as an intentional, race-checked access rather than an accidental data race.
| Tool | Protects | Typical use |
|---|---|---|
| wmb() / rmb() / mb() | Ordering between multiple memory operations | DMA descriptor setup, MMIO sequencing |
| READ_ONCE() / WRITE_ONCE() | A single shared variable’s access itself | Flags, counters read outside a lock |
| spinlock_t / mutex | A whole critical section, plus mutual exclusion | Multi-field structures needing atomicity |
Does volatile Replace Any of This?
A recurring question from students is whether marking a variable volatile makes locking or barriers unnecessary. It does not. volatile only tells the compiler to stop assuming it fully understands every access to that variable, so it will not cache the value in a register or eliminate what looks like a redundant read or write. That is genuinely useful for memory-mapped I/O registers, where a “redundant” read might actually trigger a hardware side effect.
What volatile does not give you is atomicity of read-modify-write sequences, and it does not, by itself, order that access against other unrelated memory operations the way an explicit barrier does. In modern kernel code, READ_ONCE()/WRITE_ONCE() are generally preferred over hand-rolled volatile variables, because they make the “this is a deliberately shared access” intent explicit at the call site.
Frequently Asked Questions
Do I need a memory barrier every time I touch shared data?
No. If a spinlock or mutex already protects the entire critical section, the lock and unlock operations already include the necessary barrier semantics. Explicit barriers are mainly needed for lock-free patterns, like the descriptor-valid-flag pattern shown above, or when talking to memory-mapped hardware.
What is the difference between wmb() and smp_wmb()?
wmb() is a full barrier meant to order accesses that are visible to non-CPU agents such as DMA-capable hardware. smp_wmb() only orders accesses between CPUs and can be a cheaper no-op on non-SMP kernel builds, since there is no other core to observe the reordering.
Can I use READ_ONCE()/WRITE_ONCE() instead of a spinlock?
Only for single, independent variables such as a flag or a counter that does not need to be consistent with other fields. As soon as two or more related fields must be updated or read together atomically, you need a lock, not just READ_ONCE()/WRITE_ONCE().
Why did the demo use a software ring instead of real DMA?
Real DMA descriptor formats are specific to each hardware vendor and often protected by that vendor’s driver source. The software ring in this lecture demonstrates the exact same ordering problem and fix in a form that is original, hardware-independent, and safe to run on any kernel 6.x system.
Is this pattern still relevant on kernel 6.x with modern hardware?
Yes. Memory ordering rules come from the CPU architecture and bus protocols, not from a particular kernel version, so wmb()/rmb() and READ_ONCE()/WRITE_ONCE() remain exactly as necessary on the newest kernels as they were a decade ago.
Continue the free linux kernel programming course with the next lecture in the Kernel Synchronization series.
Next Lecture Previous Lecture
2 Comments