Mutex Lock in a Real Linux Device Driver: Fixing Dirty Reads (Kernel 6.x)-Free Linux Kernel Development Course

Mutex Lock in a Real Linux Device Driver: Fixing Dirty Reads (Kernel 6.x)
Free Linux Kernel Development Course — Kernel Synchronization Series

A mutex lock in a Linux device driver only does its job if every single access to shared data sits inside the lock/unlock pair — including the reads you use just to print a debug message. In this free Linux kernel development course lecture we take a small, original character device driver, wrap its shared driver-context data with a mutex, and then deliberately walk into a classic bug: reading protected fields after the mutex has already been released. You will see exactly why that is unsafe, how a dirty (torn) read happens on multicore kernel 6.x systems, and how to fix it correctly. We finish by introducing mutex_trylock(), the non-blocking sibling of mutex_lock().

What You Will Learn
Protecting a driver context struct with a mutex
Why init/exit code paths need no locking
Dirty reads and torn reads explained with a live example
Correct critical section boundaries
mutex_trylock() non-blocking variant
Kernel 6.x safe cleanup with mutex_destroy()
Prerequisites
Basic character/misc driver structure mutex_lock() / mutex_unlock() basics Critical sections and race conditions Kernel module build system (Kbuild)

If you have not yet covered why locking is needed at all, read the earlier lectures in this free linux kernel development course on race conditions and critical sections before continuing here.

A Small Driver Context Protected by One Mutex

Real drivers rarely protect a single bare integer. More commonly you have a small “driver context” structure holding several related fields — a byte counter, a small internal buffer, maybe a state flag — and you protect the whole structure with one mutex. That is the pattern we will build here, from scratch, for kernel 6.x.

struct ep_ctx {
    struct device *dev;
    struct mutex lock;      /* protects everything below */
    int access_count;       /* how many times device was opened */
    unsigned long bytes_served;
};

static struct ep_ctx *ep_ctx;

static int ep_open(struct inode *inode, struct file *filp)
{
    mutex_lock(&ep_ctx->lock);
    ep_ctx->access_count++;
    mutex_unlock(&ep_ctx->lock);

    dev_info(ep_ctx->dev, "device opened, access_count=%d\n",
             ep_ctx->access_count);   /* BUG: unprotected read, see below */

    return 0;
}

The increment itself is safely wrapped by mutex_lock() / mutex_unlock(). So far, so good. But look closely at the line right after mutex_unlock(): we read ep_ctx->access_count again, this time with no lock held at all.

Where the Critical Section Actually Ends
Thread A (open) Thread B (open, other CPU) —————- ————————– mutex_lock(&lock) access_count++ <– protected mutex_unlock(&lock) mutex_lock(&lock) dev_info(…, access_count) <– UNPROTECTED READ, races with B below access_count++ mutex_unlock(&lock) dev_info(…, access_count)

The moment mutex_unlock() runs, any other CPU is free to take the lock and start modifying access_count immediately. The dev_info() call on Thread A is now reading a field that Thread B may be mutating at the exact same instant. On most architectures a plain int read/write is atomic in isolation, but that is not the guarantee we actually need here: we need the printed value to correspond to a consistent state of the structure, and once you have more than one field involved (as in the next example), a torn read becomes a real possibility, not just a theoretical one.

A Genuine Dirty Read with Two Fields

The danger is easiest to see when the debug print touches two related fields at once. Extend the context with a second field, bytes_served, updated together with access_count inside the lock, but printed outside it:

static void ep_record_access(size_t n)
{
    mutex_lock(&ep_ctx->lock);
    ep_ctx->access_count++;
    ep_ctx->bytes_served += n;
    mutex_unlock(&ep_ctx->lock);

    /* BUG: both fields read here, with no lock held */
    dev_info(ep_ctx->dev, "count=%d bytes=%lu\n",
             ep_ctx->access_count, ep_ctx->bytes_served);
}

Because the two reads inside dev_info() are two separate memory accesses, another CPU can run its own locked update in between them. The log line can then print an access_count from call number 5 alongside a bytes_served total from call number 6 — two numbers that never actually existed together in that combination. That mismatch is exactly what “dirty read” or “torn read” means in kernel synchronization: not corrupted bits, but an inconsistent snapshot of related shared state.

The Fix: Keep the Read Inside the Critical Section

The rule is simple: anything you read for the purpose of reporting the current state of shared data must be read while still holding the lock. Copy the values you need into local variables before unlocking, then print the local copies:

static void ep_record_access(size_t n)
{
    int count;
    unsigned long total;

    mutex_lock(&ep_ctx->lock);
    ep_ctx->access_count++;
    ep_ctx->bytes_served += n;
    count = ep_ctx->access_count;      /* snapshot while locked */
    total = ep_ctx->bytes_served;      /* snapshot while locked */
    mutex_unlock(&ep_ctx->lock);

    dev_info(ep_ctx->dev, "count=%d bytes=%lu\n", count, total);
}

count and total are local variables living on this thread’s kernel-mode stack, so they need no protection — each thread gets its own private copy. Once they are captured inside the lock, the dev_info() call is free to run after mutex_unlock() without any risk, because it is no longer touching shared memory at all.

Approach Where shared fields are read Safe?
Print directly from struct after unlock Outside critical section No — dirty/torn read possible
Snapshot into locals, then print Snapshot inside, print outside Yes
Print directly from struct before unlock Inside critical section Yes, but lengthens the critical section

Both safe options are valid engineering trade-offs. Keeping the print inside the lock is simplest but makes the critical section slightly longer (a printk-family call is not free). Snapshotting into locals keeps the critical section short and is usually the better choice when the log statement does any expensive formatting.

Why init() and exit() Never Needed the Lock

You might wonder why the module’s init and exit functions can freely write to ep_ctx fields with no locking at all:

static int __init ep_init(void)
{
    ep_ctx = kzalloc(sizeof(*ep_ctx), GFP_KERNEL);
    if (!ep_ctx)
        return -ENOMEM;

    mutex_init(&ep_ctx->lock);
    ep_ctx->access_count = 0;      /* no lock needed here */
    ep_ctx->bytes_served = 0;      /* no lock needed here */

    return misc_register(&ep_miscdev);
}

static void __exit ep_exit(void)
{
    misc_deregister(&ep_miscdev);
    mutex_destroy(&ep_ctx->lock);
    kfree(ep_ctx);
}

The reason is concurrency, or rather the guaranteed absence of it. ep_init() runs exactly once, in exactly one process context (typically the insmod(8)/modprobe(8) process), before the device node even exists for anyone to open. ep_exit() runs exactly once too, after misc_deregister() has already made sure no new opens can start. With no second thread able to touch the structure at those two moments, there is nothing to protect against — a lock only matters when two or more contexts can race, and here they cannot.

mutex_trylock(): The Non-Blocking Variant

Every example so far uses mutex_lock(), which puts the calling thread to sleep if the lock is already held. Sometimes that is not what you want — you would rather find out immediately that the lock is busy, do something else, and try again later. That is exactly what mutex_trylock() is for: it never sleeps. It grabs the lock if free and returns non-zero, or returns zero immediately if the lock is currently held, without waiting at all.

if (mutex_trylock(&ep_ctx->lock)) {
    /* lock acquired immediately */
    ep_ctx->access_count++;
    mutex_unlock(&ep_ctx->lock);
} else {
    /* lock was busy - do not block, handle gracefully */
    dev_dbg(ep_ctx->dev, "busy, skipping this update\n");
}
mutex_lock() vs mutex_trylock()
mutex_lock(&lock) mutex_trylock(&lock) lock free? — yes –> acquire lock free? — yes –> acquire, return nonzero lock free? — no –> SLEEP lock free? — no –> return 0 immediately, no sleep (wait for owner) (caller decides what to do next)
Variant Blocks caller? Typical use
mutex_lock() Yes, uninterruptible sleep Normal case, caller can afford to wait
mutex_lock_interruptible() Yes, but signal-interruptible User-facing paths, must honour Ctrl-C
mutex_trylock() Never Polling loops, best-effort updates, avoiding priority inversion in latency-sensitive paths

On kernel 6.x, mutex_trylock() is still built on the same underlying fast-path/slow-path mutex implementation as mutex_lock(); it simply refuses to enter the slow, sleeping path when the fast path fails. Use it when blocking is unacceptable in that code path — for example inside some atomic-context adjacent helper, or when you would rather skip an optional statistics update than stall a hot path.

Common Mistakes

  • Printing shared fields after unlocking — the dirty read bug covered above.
  • Locking init()/exit() “just to be safe” — harmless but misleading; it suggests concurrency exists where none does, confusing future readers of the driver.
  • Treating mutex_trylock() failure as an error — a failed trylock simply means “busy right now”; design the fallback path deliberately instead of logging it as a fault.
  • Forgetting mutex_destroy() — every mutex_init() needs a matching mutex_destroy() in the cleanup path on kernel 6.x, or you will trip lock-debugging warnings on a debug kernel.

Best Practices

  • Snapshot shared fields into local variables before unlocking if you need to log or return them.
  • Keep the critical section as short as reasonably possible — move expensive work (formatting, logging, computation) outside the lock wherever it does not need the shared data.
  • Reach for mutex_trylock() only when blocking is genuinely unacceptable in that path; default to mutex_lock() everywhere else.
  • Always test locking changes on a debug kernel (CONFIG_PROVE_LOCKING, KCSAN) so lockdep can catch ordering and usage mistakes before they reach production.

Summary / Key Takeaways

  • A mutex only protects the exact statements between mutex_lock() and mutex_unlock() — nothing before or after.
  • Reading shared fields after unlocking, purely to log them, reintroduces the race you just fixed.
  • Snapshot shared values into local variables while still holding the lock; local variables never need protection.
  • init()/exit() code paths run in a single context by construction, so they need no locking.
  • mutex_trylock() gives you a non-blocking option when your code path cannot afford to sleep.

FAQ: Mutex Locks and Critical Sections in Linux Drivers

Is reading a single shared integer without a lock ever safe?

Only if that single read is not meant to be consistent with any other shared field, and you have verified the platform treats aligned word-size reads as atomic. Relying on this is fragile; explicit locking or READ_ONCE() is the safer default.

Does keeping a printk inside the lock slow the driver down?

It lengthens the critical section slightly because other CPUs must wait for the print to finish before acquiring the lock. For low-frequency debug logging this is usually acceptable; for hot paths, snapshot into locals instead.

What happens if I forget mutex_destroy() on kernel 6.x?

The mutex object is left in an undefined state when the memory backing it is freed. Lock-debugging kernels can flag this as a use-after-free-style issue, so always pair every mutex_init() with a mutex_destroy() in the matching cleanup path.

Can mutex_trylock() be used inside interrupt context?

No. Mutexes are sleeping locks and are not intended for atomic/interrupt context at all, even the non-blocking trylock variant. Use spinlocks for anything that must run in interrupt context.

Is a dirty read the same thing as data corruption?

Not exactly. The individual bits read are usually valid values that existed at some point, but the combination read together may never have existed as a single consistent state — that mismatch is what makes it “dirty” or “torn” rather than simply wrong.

Do I need a separate mutex per field, or one per structure?

One mutex per logically-related structure is the common and recommended pattern. Splitting locks per field adds complexity and lock-ordering risk without a clear benefit unless profiling shows real contention.

Why not just make access_count an atomic_t instead of using a mutex?

atomic_t is indeed the better tool for a single counter and avoids locking overhead entirely. Mutexes become necessary once you must update or read more than one related field together as one consistent unit, which an atomic type alone cannot guarantee.

Continue the Free Linux Kernel Development Course

More lectures on kernel synchronization, mutexes, spinlocks and atomic operations are on the way in this free Linux device drivers course.

Previous Lecture Next Lecture

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *