Linux Kernel Data Race: SMP, Preemption & Interrupt Concurrency-Linux Device Driver Training Online

← Previous Lecture    Next Lecture →

Linux Kernel Data Race: SMP, Preemption & Interrupt Concurrency
Free Linux Kernel Programming Course — Kernel Synchronization Part 1
📘 Kernel 6.x Updated
⏱ 18 Min Read
💻 Original Code Demos

A linux kernel data race is one of the hardest bugs to catch in device driver code, because it only shows up when timing lines up just wrong. In this free linux kernel development course lecture, we look at the three real-world situations that create a data race inside a driver: multicore SMP systems running the same code on two CPUs at once, a preemptible kernel switching processes mid-critical-section, and hardware interrupts firing on top of your driver’s process-context code. Every example here is written and tested against a current 6.x kernel, so you are learning patterns that still apply today, not decade-old APIs.

What You Will Learn

  • Multicore SMP data races
  • Preemptible kernel races
  • Blocking I/O races
  • Hardware interrupt races
  • Locking granularity
  • One-lock-per-structure rule

Prerequisites

You should already know what a critical section is and why a plain counter++ is not atomic on any CPU architecture. If you have not covered that yet, please read the previous lecture in this free linux device drivers course first, since this lecture builds directly on it.

1. Multicore SMP Systems and Data Races

Almost every embedded board shipping today, from a quad-core Raspberry Pi class SoC to an automotive-grade multicore SoC, is a Symmetric Multi-Processing (SMP) system. That single fact changes everything about how you must write driver code. If two user-space processes, say P1 and P2, open the same device file and both call read() at almost the same instant, the scheduler is free to place P1 on CPU 0 and P2 on CPU 2. Both cores now execute your driver’s read() callback simultaneously, in real parallel time, not just interleaved. If that callback touches a global or static writable buffer without a lock, you have two CPUs mutating the same memory at once — a genuine SMP data race.

Diagram: Two CPUs Racing on One Shared Buffer
CPU 0 — Process P1
1. enter read()
2. iterate shared_buf[]
3. write shared_buf[i]
⇄
CPU 2 — Process P2
1. enter read()
2. iterate shared_buf[]
3. write shared_buf[i]
Both CPUs execute step 3 at the same real-world instant — the buffer is corrupted, not just reordered.

Here is an original, minimal character-device snippet that shows the unsafe pattern on a modern kernel. Notice there is no locking at all around the shared array — this is deliberately broken code, written only to illustrate the race:

/* unsafe_read.c — deliberately unprotected shared buffer (kernel 6.x) */
#define ARRAY_LEN 64
static int shared_stats[ARRAY_LEN];   /* global, writable, NO lock */

static ssize_t unsafe_dev_read(struct file *filp, char __user *buf,
                                size_t count, loff_t *pos)
{
    int i;

    /* Two CPUs can be inside this loop at the same time */
    for (i = 0; i < ARRAY_LEN; i++) {
        shared_stats[i]++;            /* read-modify-write, unprotected */
    }

    return simple_read_from_buffer(buf, count, pos,
                                    shared_stats, sizeof(shared_stats));
}

The fix for this is always some form of mutual exclusion — a spinlock or mutex around the loop — which we begin covering properly starting next lecture. For now, the key takeaway is: on SMP, “simultaneous” is not a theoretical word, it is literal parallel execution.

2. Preemptible Kernels and Blocking I/O Data Races

Even a single-core (UP) embedded board is not safe from data races. Modern Linux kernels default to some form of preemption — controlled today through CONFIG_PREEMPT_DYNAMIC, which lets you pick none, voluntary, or full at boot time via the preempt= kernel command-line parameter, and newer kernels are also adding a lazy preemption model. With full or dynamic preemption, the kernel can switch away from process P1 in the *middle* of your critical section, purely based on scheduling priority, and run another process P2 that happens to enter the very same code path.

A second, closely related danger is a blocking call inside a critical section. Functions like wait_event_interruptible(), mutex_lock() on a contended mutex, or even a naive msleep() put the calling process to sleep, handing the CPU to whichever process the scheduler picks next. If that next process re-enters your unprotected code, the shared data changes underneath the sleeping process, and it wakes up to corrupted state.

Diagram: Preemption Mid-Critical-Section
t0 — P1 enters critical section, starts updating shared struct
t1 — scheduler preempts P1 mid-update (timer tick / higher-priority task ready)
t2 — P2 is scheduled, enters the SAME critical section, sees half-updated data
t3 — P1 resumes later — struct now inconsistent, driver misbehaves

Here is an original demo kernel module. It uses two kthreads that repeatedly update a shared struct with an artificial usleep_range() in the middle of the update, on purpose, to reliably reproduce a preemption-style race for learning purposes:

/* preempt_race_demo.c — reproduces a preemption-style race (kernel 6.x) */
struct shared_record {
    u32 seq;
    u32 checksum;
};
static struct shared_record rec;   /* no lock, deliberately unsafe */

static int worker_fn(void *arg)
{
    while (!kthread_should_stop()) {
        rec.seq++;
        /* blocking-style gap: lets the scheduler switch to the
         * other kthread right here, mid-update */
        usleep_range(50, 100);
        rec.checksum = rec.seq * 2;   /* may now be based on a
                                       * seq value the other thread
                                       * already changed again */
        cond_resched();
    }
    return 0;
}

Running two instances of worker_fn() as separate kthreads on the same core will show checksum != seq * 2 intermittently — that mismatch is your race, captured in a controlled, repeatable way.

3. Hardware Interrupts and Data Races

Hardware interrupts are the highest-priority code path on Linux — they preempt everything, including kernel code, on the CPU where they fire. If your driver’s interrupt handler (the hardirq top half, or a threaded IRQ / tasklet bottom half) touches the same global data as your process-context code, you have a process-context-vs-atomic-context race, and it is one of the most common driver bugs in the field.

A practical way to reproduce this without real hardware is to use a high-resolution timer (hrtimer), because its callback runs in soft-IRQ (atomic) context — the same execution environment as most interrupt bottom halves. Here is an original demo that races a kthread (process context) against an hrtimer callback (atomic context) on a shared counter:

Diagram: Process Context vs Atomic (Interrupt-like) Context
Process Context
kthread running worker_fn(), can sleep, can be preempted
⚠
Atomic Context
hrtimer callback, runs in soft-IRQ, cannot sleep, always preempts process context
Both paths touch shared_count at unpredictable moments — a classic top-half / bottom-half style race.
/* hrtimer_race_demo.c — process context vs atomic context race (kernel 6.x) */
static struct hrtimer race_timer;
static long shared_count;          /* unprotected on purpose */

static enum hrtimer_restart timer_cb(struct hrtimer *t)
{
    shared_count++;                 /* runs in atomic/softirq context */
    hrtimer_forward_now(t, ms_to_ktime(2));
    return HRTIMER_RESTART;
}

static int process_ctx_worker(void *arg)
{
    while (!kthread_should_stop()) {
        shared_count += 3;           /* runs in normal process context */
        cond_resched();
    }
    return 0;
}

/* In init: hrtimer_setup(&race_timer, timer_cb, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
 *          hrtimer_start(&race_timer, ms_to_ktime(2), HRTIMER_MODE_REL);
 *          kthread_run(process_ctx_worker, NULL, "race_worker");
 * In exit:  hrtimer_cancel(&race_timer); kthread_stop(worker_task);
 */

Because the hrtimer callback can interrupt process_ctx_worker() at any point in its read-modify-write of shared_count, increments get silently lost. That lost-update pattern is exactly what happens in real interrupt handlers that share data with process context, just without a timer standing in for real hardware.

4. Locking Guidelines: Granularity and One-Lock-Per-Structure

Once you can spot a critical section, the next skill is choosing how to lock it well. Two mistakes are equally common in real projects:

Mistake Consequence
Too few locks (one giant lock for everything) Heavy contention, poor scalability on SMP, one hot lock becomes a bottleneck
Too many locks, unclear ownership Hard to reason about, higher deadlock risk, easy to grab the wrong lock for a structure

The practical rule most kernel codebases follow is: one lock protects one shared data structure, and the lock is often embedded as a member field inside that very structure so its ownership is unambiguous to every future reader of the code — including you, six months from now.

/* one-lock-per-structure convention (kernel 6.x, spinlock_t shown for reference) */
struct device_stats {
    spinlock_t lock;      /* protects every field below it */
    u32 read_count;
    u32 write_count;
};

We will put this pattern to work with real spinlock and mutex APIs starting in the next lecture of this free linux kernel development course.

Topics covered:
free linux kernel development course free linux device drivers course free embedded systems course linux kernel data race kernel preemption

Frequently Asked Questions

Q1. What is a data race in the Linux kernel?
A data race happens when two or more execution contexts — processes, threads, or interrupt handlers — access the same writable memory at the same time without proper locking, and at least one access is a write.

Q2. Can a single-core (UP) embedded system have data races?
Yes. A preemptible kernel can switch processes mid-critical-section, and hardware interrupts can fire on top of process-context code, both of which create races even without a second CPU core.

Q3. Why does a blocking call inside a critical section matter?
A blocking call puts the current process to sleep and hands the CPU to another process, which may re-enter the same unprotected code path and change the shared data before the sleeping process resumes.

Q4. What is the difference between process context and atomic context?
Process context code can sleep and be preempted normally, while atomic context — such as a hardirq handler or an hrtimer/softirq callback — cannot sleep and always has priority over process context on the same CPU.

Q5. What does CONFIG_PREEMPT_DYNAMIC control?
It lets a single kernel build choose its preemption model — none, voluntary, or full — at boot time via the preempt= command-line parameter, instead of requiring a separate kernel build for each mode.

Q6. Why use one lock per data structure instead of one global lock?
A single global lock becomes a contention bottleneck on SMP systems, while embedding a lock inside each structure keeps ownership clear and limits contention to only the code touching that specific structure.

Q7. Is this course updated for current kernels?
Yes, every example in this free linux kernel development course targets current 6.x kernel APIs rather than older, removed interfaces.

← Previous Lecture    Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *