← Previous Lecture | Next Lecture →
If you’re serious about writing correct, production-grade device drivers, understanding the Linux kernel critical section concept is non-negotiable. Every driver you ship runs on multi-core hardware today, which means two or more CPUs can execute your driver’s code at exactly the same moment. If that code touches shared, writeable data without protection, you get silent data corruption — one of the hardest classes of bugs to debug. This lecture, part of our free Linux kernel development course, breaks down critical sections, exclusivity, and atomicity from first principles, using modern kernel 6.x terminology and tooling throughout.
- Basic C programming — pointers, structs, global variables
- A Linux machine (kernel 6.x) with kernel headers installed, for building modules
- Familiarity with loading/unloading kernel modules (insmod/rmmod) is helpful but not required
- No prior synchronization knowledge assumed — we start from zero
What Exactly Is a Critical Section?
A critical section is any piece of code that reads or writes shared, writeable data — data that more than one thread, kernel thread, interrupt handler, or CPU core could touch at the same time. In a device driver, this is typically a global variable, a static variable, or a field inside a shared data structure that your open(), read(), write(), or ioctl() callbacks all touch.
The rule is simple but absolute: a Linux kernel critical section must run exclusively — meaning exactly one thread of execution is allowed inside it at any given instant. Without this guarantee, two CPUs can interleave their reads and writes on the same memory location, and the final result depends on unpredictable timing. This is called a race condition, and it is one of the most common sources of kernel crashes and silent data corruption in real-world drivers.
Code that only touches local (stack) variables is completely safe to run in parallel — every thread has its own private stack, so there is nothing to share. The danger begins the instant your code reaches out to global or static memory.
A Simple Timeline View
Instead of a static picture, think of any function in your driver as passing through three phases. The middle phase is the one that needs protection:
Local variables only
✅ Safe in parallel
Shared/global data
⚠️ Critical Section
Local variables only
✅ Safe in parallel
Exclusive Execution vs Atomicity — Not the Same Thing
People new to kernel synchronization often conflate two related but distinct properties:
| Property | Meaning | When Required |
|---|---|---|
| Exclusive | Exactly one thread runs the critical section’s code at a time (serialized) | Always, whenever shared writeable data is touched |
| Atomic | The operation runs indivisibly, to completion, without being interrupted partway through | Only when the code can run in a non-blocking context, such as a hardirq, tasklet, or softirq |
In other words: exclusivity is about how many threads are allowed inside; atomicity is about whether the operation can be paused halfway. A critical section running in normal process context (like a typical read() system call) can usually sleep and doesn’t need to be atomic — it just needs to be exclusive. But code running inside an interrupt handler cannot sleep at all, so if it touches shared data, that access must be both exclusive and atomic.
Process Context vs Atomic Context
Where your code runs decides how strict your synchronization needs to be. This distinction becomes critical once you start choosing between a mutex and a spinlock in the next lecture, so it’s worth internalizing now.
| Context Type | Examples | Can It Sleep? | Atomicity Required? |
|---|---|---|---|
| Process context | open, read, write, ioctl, mmap, kernel thread body, workqueue handler | Yes | Not necessarily |
| Atomic / interrupt context | hardirq handler, softirq, tasklet | No | Yes — always |
A handy mental check: if the code path could legally call something that sleeps (like kmalloc(GFP_KERNEL, ...) or mutex_lock()), it’s process context. If it absolutely cannot block — such as inside a hardware interrupt handler — it’s atomic context, and only atomic-safe primitives like spinlocks or atomic operations are allowed there.
Why Are Some Operations Automatically Atomic?
Modern CPUs guarantee exactly two categories of operation as atomic without any extra locking:
- A single machine language instruction always runs to completion — it cannot be interrupted mid-instruction.
- A read or write of an aligned primitive variable that fits within the processor’s native word size (32-bit or 64-bit) is torn-free. Any thread reading it will see either the fully old value or the fully new value, never a corrupted half-written mix.
This second point is exactly why many old driver tutorials casually increment a plain int global counter and call it “safe.” A single aligned load or store is atomic — but an increment (read-modify-write) is not, because it’s actually three separate machine steps: read the value, add one, write it back. Two CPUs can both read the same old value before either writes back the new one, and one increment gets silently lost.
Kernel 6.x Update: READ_ONCE(), WRITE_ONCE(), and KCSAN
This is one area where the theory hasn’t changed, but kernel tooling has moved forward significantly. Modern kernel code (6.x) rarely accesses a shared variable with a bare read or write anymore. Instead, the kernel encourages:
READ_ONCE(x)andWRITE_ONCE(x, val)— these prevent the compiler from reordering, caching, or splitting the access in ways that would break the atomicity guarantee described above.atomic_t/atomic64_tand helpers likeatomic_inc(),atomic_read()— for counters that genuinely need safe read-modify-write semantics without a full lock.- KCSAN (Kernel Concurrency Sanitizer) — a build-time and runtime data-race detector available via
CONFIG_KCSANon modern kernels. It instruments memory accesses and reports unprotected concurrent reads/writes to the kernel log, catching exactly the kind of bug this lecture describes, automatically.
Hands-On: A Race Condition Demo Driver
Theory is easy to nod along to; seeing a race actually happen is what makes it click. Below is an original, minimal kernel module (tested against kernel 6.x kthread APIs) that spins up two kernel threads which both increment a plain shared int counter, with no protection at all.
#include <linux/init.h>
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/delay.h>
#define ITERATIONS 200000
static int shared_counter;
static struct task_struct *worker_a, *worker_b;
static int racer_thread(void *data)
{
const char *name = (const char *)data;
int i;
for (i = 0; i < ITERATIONS; i++) {
/* Deliberately UNPROTECTED read-modify-write */
shared_counter++;
}
pr_info("ep_race: thread %s finished\n", name);
return 0;
}
static int __init ep_race_init(void)
{
shared_counter = 0;
worker_a = kthread_run(racer_thread, "worker-A", "ep_race_a");
worker_b = kthread_run(racer_thread, "worker-B", "ep_race_b");
if (IS_ERR(worker_a) || IS_ERR(worker_b)) {
pr_err("ep_race: failed to start worker threads\n");
return -ENOMEM;
}
pr_info("ep_race: module loaded, expected final value = %d\n",
ITERATIONS * 2);
return 0;
}
static void __exit ep_race_exit(void)
{
pr_info("ep_race: final shared_counter = %d (expected %d)\n",
shared_counter, ITERATIONS * 2);
}
module_init(ep_race_init);
module_exit(ep_race_exit);
MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("Original demo: unprotected critical section race condition");
Load this module, wait for both threads to finish (check dmesg), then unload it. On a multi-core machine you will very often see the final shared_counter value come out lower than the expected ITERATIONS * 2. That gap is real, silently lost updates — the direct, observable cost of an unprotected Linux kernel critical section. We deliberately have not fixed this yet with a lock; that’s the subject of the next lecture in this series.
A Quick Preview: Why atomic_t Alone Fixes This Specific Case
For a simple counter like the one above, the kernel provides a ready-made atomic type instead of a full lock:
#include <linux/atomic.h>
static atomic_t safe_counter = ATOMIC_INIT(0);
static int racer_thread_safe(void *data)
{
int i;
for (i = 0; i < ITERATIONS; i++)
atomic_inc(&safe_counter);
return 0;
}
atomic_inc() performs the read-modify-write as a single indivisible hardware-backed operation, so no update is ever lost. This works well for a lone counter, but it doesn’t generalize to protecting a multi-field struct or a longer code path — that’s exactly why the kernel also gives us mutexes and spinlocks, which we cover next.
Common Mistakes When Reasoning About Critical Sections
- Assuming a single-statement increment is safe because “it’s just one line of C” — it usually compiles to multiple machine instructions.
- Protecting the variable but not the whole logical operation — e.g., locking only the read, not the subsequent write.
- Using a sleeping lock inside interrupt context — a mutex can sleep; hardirq handlers cannot afford to.
- Testing only on a single-core VM — races that are rare on one core can appear far more often on real multi-core hardware.
- Ignoring compiler reordering — without
READ_ONCE/WRITE_ONCE, the compiler is free to optimize accesses in surprising ways.
Best Practices
- Identify every global/static variable your driver touches and explicitly decide its locking strategy before writing the logic around it.
- Keep critical sections as short as possible — the less code inside, the smaller the performance cost of serialization.
- Prefer the narrowest primitive that solves your problem: a lone counter often just needs
atomic_t, not a full mutex. - Enable
CONFIG_KCSANin your test kernel builds to catch races you didn’t think of. - Document, in a comment, exactly which lock protects which field — future maintainers (including you) will thank you.
Performance and Security Considerations
Performance: Every exclusivity mechanism has a cost — even a fast spinlock adds cache-line bouncing between cores under contention. The performance goal isn’t “avoid locking,” it’s “lock the smallest possible critical section for the shortest possible time.”
Security: Unprotected critical sections aren’t just a correctness bug — a race window in a driver that manages permissions, buffers, or reference counts can sometimes be turned into a security vulnerability (a classic Time-Of-Check-Time-Of-Use style bug). Treat every shared-data race you find as a bug worth fixing immediately, not a cosmetic glitch.
Summary — Key Takeaways
- A Linux kernel critical section is any code touching shared writeable data.
- Critical sections must always run exclusively; they must additionally run atomically only when executing in a non-blocking atomic context (hardirq/softirq/tasklet).
- Only single instructions and aligned word-size reads/writes are automatically atomic on modern CPUs — a read-modify-write like
counter++is not. - Modern kernel 6.x code should prefer
READ_ONCE()/WRITE_ONCE()andatomic_tover bare shared-variable access, and KCSAN can catch what you miss.
Conclusion
Understanding the Linux kernel critical section concept is the foundation everything else in kernel synchronization builds on. Before you ever write mutex_lock() or spin_lock(), you need to be able to look at a function and immediately spot which lines touch shared state — that’s the actual skill. In the next lecture of this free Linux kernel development course, we’ll take this exact race condition demo driver and fix it properly, comparing mutexes and spinlocks and showing precisely when to reach for each one.
FAQ
Q1. What is a critical section in the Linux kernel?
It’s any code path that reads or writes data shared between multiple threads, CPUs, or interrupt contexts, and therefore must be protected from simultaneous execution.
Q2. Is incrementing a global integer in a kernel module thread-safe?
No. Even though the variable itself is small enough to be read/written atomically, the increment operation is a read-modify-write sequence of multiple instructions, which is not atomic on its own.
Q3. What’s the difference between exclusive execution and atomicity?
Exclusive means only one thread runs the code at a time. Atomic means the operation cannot be interrupted partway through. Exclusivity is always required for shared data; atomicity is only required in non-blocking atomic contexts.
Q4. What counts as atomic context in the kernel?
Hardware interrupt handlers, softirqs, and tasklets. Code running here cannot sleep and must use atomic-safe primitives only.
Q5. Why do modern kernels use READ_ONCE() and WRITE_ONCE()?
They prevent the compiler from reordering, merging, or splitting memory accesses in ways that would silently break the atomicity guarantees of aligned reads/writes.
Q6. What is KCSAN?
The Kernel Concurrency Sanitizer — a runtime tool available via CONFIG_KCSAN that detects unprotected concurrent data accesses (data races) in a running kernel and reports them in the kernel log.
Q7. Should I always use a lock instead of atomic_t?
No. For a single counter or flag, atomic_t is lighter-weight and sufficient. Locks are needed when you must keep multiple related fields consistent together.
Q8. Can a race condition in a driver be a security issue?
Yes. Races around permission checks, reference counts, or buffer state can sometimes be exploited, similar to classic time-of-check-to-time-of-use bugs.
Continue Your Free Linux Kernel Development Course
This lecture is part of EmbeddedPathashala’s free Linux kernel development course, covering kernel synchronization, device drivers, and embedded Linux from the ground up.
Next Lecture: Mutex vs Spinlock Back to Course Index
2 Comments