Linux Kernel Memory Barriers Explained: rmb(), wmb(), mb() and SMP Variants-Linux Kernel Development Course Online

Linux Kernel Memory Barriers Explained: rmb(), wmb(), mb() and SMP Variants
A free lesson from the Linux kernel development course — understand why memory barriers exist and how modern kernels use them in driver code.
Level
Intermediate/Advanced
Kernel
6.12+ (LTS)
Reading Time
13 min

Modern CPUs and compilers are allowed to reorder memory reads and writes for performance — and most of the time that’s completely safe. But in device driver code, where the CPU is coordinating with hardware or with other cores, that reordering can silently break correctness. This lesson, part of our free linux device drivers course, explains what a memory barrier is, when the kernel needs one, and how the API has evolved on modern kernels beyond the classic mb()/rmb()/wmb() trio.

In this lesson you’ll work with:
mb() / rmb() / wmb() smp_mb() READ_ONCE / WRITE_ONCE DMA ordering acquire/release semantics

What You Will Learn

  • Why compilers and CPUs reorder memory operations in the first place
  • The difference between a read barrier, a write barrier, and a full barrier
  • Why the kernel splits barriers into SMP-only and device-facing variants
  • How a real DMA descriptor write pattern uses a barrier to stay correct
  • Modern, safer alternatives to raw barriers for many common cases

Prerequisites

  • Basic C pointer and struct knowledge
  • Some exposure to multi-core CPU concepts (caches, cores)
  • Familiarity with basic device driver structure is helpful but not required

Why Reordering Happens

Both the compiler and the CPU are constantly looking for ways to execute your code faster without changing its observable single-threaded behavior. That freedom includes rearranging independent loads and stores, buffering writes in a store queue, and letting reads race ahead of writes to unrelated addresses. On a single core running one thread, none of this is visible. The moment a second core — or a hardware device sharing memory with the CPU — starts observing those same memory locations, reordering can expose a sequence of events that never matches what your source code said.

Why Order Matters: Two Cores, One Shared Flag
Core A
1. write data
2. write “ready” flag
Core B
1. read “ready” flag
2. read data

Without a barrier, Core B might see the flag set before the data write is visible — reading garbage.

The Three Classic Barrier Types

The kernel’s <asm/barrier.h> header exposes three fundamental barrier primitives:

BarrierGuarantee
rmb()All reads before it complete before any read after it
wmb()All writes before it become visible before any write after it
mb()Full ordering — both reads and writes before it are ordered against both reads and writes after it

These “full” barriers order memory as seen by anything sharing the bus — including memory-mapped I/O devices — which makes them essential for driver code, but also relatively expensive.

SMP-Only Barriers: The Cheaper Cousins

Not every ordering problem involves a hardware device. When you only need to order memory as seen by other CPU cores — not by a peripheral — the kernel provides SMP-scoped equivalents that compile down to nothing at all on uniprocessor builds:

smp_mb();
smp_rmb();
smp_wmb();

On a multi-core build these still emit the necessary hardware ordering instruction, but they skip the extra guarantees needed for device I/O. In current kernel code, smp_* variants are the default choice for pure CPU-to-CPU synchronization; the plain mb()/rmb()/wmb() family is reserved specifically for CPU-to-device ordering, such as setting up a DMA descriptor before a device reads it.

A Realistic DMA Ordering Example

Imagine a driver preparing a descriptor for a network card to read via DMA. The descriptor has an address field and a “valid” flag. If the device sees the valid flag set before the address field is actually written to memory, it will read garbage. A write barrier prevents exactly that:

struct my_dma_desc {
    u64 buffer_addr;
    u32 flags;
};

static void prepare_descriptor(struct my_dma_desc *desc, u64 addr)
{
    desc->buffer_addr = addr;

    /* Ensure the address write reaches memory before
     * the device-visible "valid" flag is set. */
    wmb();

    desc->flags = DESC_FLAG_VALID;
}

This pattern — write the payload, insert a barrier, then write the flag that signals the payload is ready — is the single most common use of memory barriers you’ll encounter across kernel driver code.

Modern Alternatives You Should Know

Recent kernels increasingly favor higher-level primitives that bundle a barrier with the memory access itself, reducing the chance of misplacing one:

PrimitiveWhen to reach for it
READ_ONCE() / WRITE_ONCE()Prevents the compiler from tearing, caching, or reordering a single shared variable access — no CPU-level ordering by itself
smp_load_acquire() / smp_store_release()Pairs a barrier directly with a load or store, commonly used to implement lock-free publish/subscribe patterns
Atomic ops with _acquire/_release suffixesGive you both an atomic update and the correct ordering in a single call

These don’t replace wmb()/rmb() for device-facing ordering, but for pure CPU-to-CPU synchronization they’re now generally preferred because the barrier can’t accidentally end up on the wrong side of the access it’s meant to protect.

Common Mistakes

  • Using smp_wmb() where a device is involved — SMP-only barriers don’t guarantee ordering as seen by a peripheral; use the full wmb() instead.
  • Assuming a lock automatically implies a full barrier everywhere it’s needed — spinlock/unlock do provide ordering guarantees, but for device-facing MMIO you still often need an explicit barrier.
  • Placing the barrier on the wrong side — a write barrier must sit between the writes it’s ordering, not after both of them.
  • Reaching for mb() everywhere “to be safe” — full barriers are expensive; use the narrowest one (rmb/wmb/smp_*) that actually matches your requirement.

Best Practices

  • Always consult the datasheet or SoC peripheral manual for your specific hardware before relying on driver-level barrier placement alone.
  • Prefer smp_store_release()/smp_load_acquire() pairs for CPU-only synchronization over raw barrier calls.
  • Comment every barrier in driver code explaining exactly what ordering it enforces — future maintainers (including you) will thank you.

Performance and Security Considerations

Barriers are not free — a full mb() can stall the pipeline while outstanding memory operations drain, so using the narrowest barrier that satisfies your correctness requirement matters for real-time and high-throughput code paths. From a security perspective, missing barriers around published shared state are a classic source of subtle race-condition bugs, so their placement deserves the same scrutiny as locking logic during code review.

Summary / Key Takeaways

  • Memory barriers stop the CPU and compiler from reordering operations across a point in your code where order actually matters.
  • rmb()/wmb()/mb() are for CPU-to-device ordering; smp_* variants are for CPU-to-CPU ordering only.
  • Modern code increasingly prefers acquire/release primitives that bundle the barrier with the access.

Conclusion

Memory barriers are one of those topics that feel abstract until you hit a bug that only reproduces on a multi-core system or with real hardware — at which point they become essential. Understanding the difference between device-facing and CPU-only barriers, and knowing when a higher-level primitive can do the job more safely, will save you from an entire category of intermittent, hard-to-reproduce driver bugs. This lesson is part of our ongoing free linux kernel development course covering synchronization from first principles through to production driver code.

Frequently Asked Questions

Q1. What’s the difference between mb() and smp_mb()?

mb() orders memory as seen by anything on the bus, including devices; smp_mb() only orders memory as seen by other CPU cores and compiles to nothing on uniprocessor kernels.

Q2. Do I need a barrier if I’m already using a spinlock?

Lock and unlock operations provide ordering guarantees for CPU-to-CPU visibility, but device-facing MMIO ordering often still needs an explicit barrier alongside the lock.

Q3. Are memory barriers only relevant for driver developers?

They matter most in driver and lock-free data structure code, but any developer writing multi-core synchronization primitives should understand them.

Q4. What does READ_ONCE() actually prevent?

It stops the compiler from caching a value in a register across iterations or reordering the access relative to other code — it does not add a CPU-level ordering guarantee by itself.

Q5. Why does the DMA API documentation specifically call out barriers?

Because DMA-capable devices access memory independently of the CPU, so the driver must guarantee its writes are actually visible in memory before telling the device the data is ready.

Q6. Is it safe to over-use mb() “just in case”?

It’s safe for correctness but costly for performance — full barriers stall the CPU pipeline, so unnecessary use directly hurts throughput and latency.

Q7. What replaced raw barriers in newer lock-free kernel code?

Acquire/release primitives such as smp_load_acquire() and smp_store_release(), which bundle the correct barrier directly into the memory access.

Continue the Free Linux Kernel Development Course

More lessons on synchronization, scheduling, and driver development — all free, all hands-on.

Browse the Course

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *