← Previous Lecture | Next Lecture →
Linux kernel memory barriers exist because the hardware you are writing code for is lying to you, in a friendly and usually harmless way. CPUs reorder memory reads and writes for performance, compilers reorder them again at build time, and most of the time your program never notices because a single CPU always sees its own operations happen in the order it issued them. The trouble starts the moment two different CPUs, or a CPU and a peripheral device, need to agree on that order. This lecture introduces what memory barriers are, why they matter for kernel and driver code, and gives you a first, original code example so the concept stops being abstract.
This lecture assumes you already understand critical sections, spinlocks, and atomic operations from earlier lectures in this series. No prior memory-barrier knowledge is assumed — that is exactly what this lecture introduces.
Why Linux Kernel Memory Barriers Matter
CPUs and Compilers Both Reorder Your Code
When you write two independent stores in C, nothing in the language guarantees they reach memory in the order you wrote them. Modern CPU cores have store buffers and out-of-order execution engines that can commit writes to memory in whatever sequence is most efficient, as long as the single core issuing them still perceives its own program order. The compiler does something similar at build time, reordering or even eliminating instructions it judges to be equivalent from a single-threaded point of view. Both optimizations are safe on one core running one thread — the problem only appears the instant another observer, such as a second CPU core or a hardware peripheral, is watching that same memory.
CPU/compiler may do: write FLAG then write DATA
Another core sees FLAG set, reads DATA -> reads stale/garbage value
What a Memory Barrier Actually Does
A memory barrier is an instruction that tells both the CPU and the compiler: do not reorder memory operations across this point. It does not make anything faster, and it is not a lock — it does not provide mutual exclusion by itself. What it provides is ordering. Placed correctly alongside a flag or a lock, a barrier is what makes the flag or lock’s ordering guarantees actually hold at the hardware level, not just at the C source level.
The Core Barrier Primitives
The kernel exposes a small family of barrier macros, and picking the right one matters both for correctness and for performance, since barriers are not free.
| Macro | Orders | Typical Use |
|---|---|---|
barrier() | Compiler only, no CPU instruction | Stop the compiler reordering; single-CPU code |
rmb() | CPU read/read ordering, plus compiler | CPU <-> device MMIO ordering on reads |
wmb() | CPU write/write ordering, plus compiler | CPU <-> device MMIO ordering on writes |
mb() | Full CPU read+write ordering, plus compiler | General hardware-boundary ordering |
smp_rmb() | Read ordering, CPU-to-CPU only | SMP core-to-core read ordering, cheaper than rmb() on UP |
smp_wmb() | Write ordering, CPU-to-CPU only | SMP core-to-core write ordering |
smp_mb() | Full ordering, CPU-to-CPU only | General SMP core-to-core ordering |
Notice the split: the plain rmb()/wmb()/mb() family targets hardware boundaries — talking to a device over memory-mapped I/O — while the smp_ prefixed versions target ordering between CPU cores on a multiprocessor system, and compile down to nothing at all on a uniprocessor build. Using the wrong family is a common source of either a missed bug on your development board or unnecessary overhead in production.
An Original Example: The Flag-Without-a-Barrier Bug
The classic scenario where this bites people is a producer thread writing some data and then setting a “ready” flag, with a consumer thread on another core waiting for that flag before reading the data. Written naively, without any lock or barrier, this pattern can fail exactly the way the diagram above shows.
/* Original example — buggy version */
static int payload;
static int ready;
/* Producer, running on CPU 0 */
void ep_producer(int value)
{
payload = value;
ready = 1; /* may be observed BEFORE payload on another core */
}
/* Consumer, running on CPU 1 */
int ep_consumer(void)
{
while (!ready)
cpu_relax();
return payload; /* could read a stale value of payload here */
}
The fix is to insert explicit barriers on both sides, so the write to payload is guaranteed visible before ready is observed as set, and the read of ready is guaranteed to happen before the read of payload:
/* Original example — corrected version */
static int payload;
static int ready;
void ep_producer(int value)
{
payload = value;
smp_wmb(); /* payload write is visible before ready write */
ready = 1;
}
int ep_consumer(void)
{
while (!ready)
cpu_relax();
smp_rmb(); /* ready read happens before payload read */
return payload;
}
In real driver code you would usually reach for WRITE_ONCE()/READ_ONCE() alongside these barriers, or use an existing kernel primitive like smp_store_release()/smp_load_acquire() that bundles the barrier and the access together. The two-function example above is deliberately spelled out step by step so the ordering is visible, rather than hidden inside a single call.
CPU 1 (consumer): wait for ready == 1 -> smp_rmb() -> read payload
Barrier guarantees: payload write is globally visible before ready write is seen
When You Actually Need a Barrier
If you are already using a mutex, spinlock, or an atomic operation like the ones covered earlier in this course, you very often do not need an explicit barrier — locking primitives and full atomic read-modify-write operations already carry the necessary ordering guarantees as part of their implementation. Explicit barriers matter most in two situations: lock-free code that intentionally avoids locking for performance, such as some per-CPU or seqlock-style patterns, and driver code talking directly to hardware over memory-mapped I/O, where the “other observer” is a physical device rather than another CPU core.
Common Mistakes
- Using a plain
mb()/rmb()/wmb()for CPU-to-CPU ordering when the cheapersmp_variant was intended — this adds needless overhead on SMP systems and is redundant on UP kernels either way. - Assuming a barrier alone provides mutual exclusion — it only orders memory operations, it never prevents two cores from both entering a critical section.
- Forgetting the barrier on one side of a producer/consumer flag while adding it correctly on the other side, which silently reintroduces the same bug.
- Reaching for manual barriers when an existing primitive like a spinlock or
smp_store_release()/smp_load_acquire()would already be correct and clearer.
Best Practices
- Prefer existing kernel primitives (locks, atomics,
READ_ONCE()/WRITE_ONCE(), acquire/release helpers) over hand-rolled barrier pairs whenever one fits the pattern. - When you do need explicit barriers, always reason about and comment both sides of the pair — a barrier on only one side is not a fix.
- Use the
smp_variants for CPU-to-CPU ordering and reserve plainmb()/rmb()/wmb()for genuine hardware/device ordering. - Document why a barrier is there; unlike a lock, a barrier leaves no trace in the code that explains itself later.
Summary and Key Takeaways
- Both the CPU and the compiler are free to reorder memory operations as long as a single thread cannot tell the difference.
- A memory barrier enforces ordering, not mutual exclusion — it is not a substitute for a lock.
smp_*barriers order CPU-to-CPU visibility; plainmb()/rmb()/wmb()order CPU-to-device visibility over MMIO.- Producer/consumer flag patterns need a write barrier before setting the flag and a matching read barrier after checking it.
- Most kernel code should prefer existing locking and atomic primitives, which already embed the correct barriers, over hand-written barrier pairs.
Conclusion
This lecture only opens the door on linux kernel memory barriers — enough to recognize when reordering could bite you and to read barrier-using code without guessing. The topic goes considerably deeper, into acquire/release semantics, the kernel’s formal memory model, and architecture-specific ordering guarantees on ARM versus x86, all of which later lectures in this synchronization chapter will build on.
What is a memory barrier in the Linux kernel?
It is an instruction that prevents the CPU and compiler from reordering memory reads and writes across that point in the code, guaranteeing a specific order is visible to other observers such as another CPU core or a hardware device.
Does a memory barrier provide mutual exclusion like a lock?
No. A barrier only enforces ordering of memory operations. It does not prevent two CPUs from entering the same critical section at the same time — that is what locks are for.
What is the difference between mb() and smp_mb()?
mb() orders memory operations with respect to hardware, such as a memory-mapped device, and always compiles to a real barrier instruction. smp_mb() orders memory operations between CPU cores and compiles down to nothing on a uniprocessor kernel build.
Do I need explicit barriers if I already use a spinlock?
Usually not. Spinlocks, mutexes, and full atomic read-modify-write operations already carry the ordering guarantees you need. Explicit barriers are mainly for lock-free code and direct hardware/MMIO access.
Why does the producer/consumer example need barriers on both sides?
A write barrier on the producer side only guarantees the order in which the producer’s own writes become visible. A matching read barrier is needed on the consumer side to guarantee the consumer’s reads happen in the corresponding order. One-sided barriers do not fix the bug.
What is the modern alternative to hand-written barrier pairs?
Kernel primitives such as smp_store_release() and smp_load_acquire() bundle the correct barrier together with the memory access, which is usually clearer and less error-prone than writing the barrier and the access as two separate steps.
More lectures on memory barriers, acquire/release semantics, and lock-free patterns are on the way.
Previous Lecture Next Lecture
2 Comments