← Previous Lecture | Next Lecture →
This free linux kernel synchronization tutorial answers a question that
trips up almost every new device driver author: why can a plain i++ on a
shared integer corrupt data on a multi-core Linux system, even though the C statement
looks like a single, harmless operation? In this lecture, part of our ongoing
free linux kernel development course, we break the increment down into
the real machine instructions the CPU executes, show why the answer depends on both the
processor and the compiler, and use that insight to introduce the single most important
synchronization idea in the kernel: the lock. Everything here targets
current kernel 6.x internals and uses only original driver code, so you can follow along
on any recent distribution.
This lecture builds directly on the previous lecture in this
free linux device drivers course, which covered critical sections,
process context versus atomic context, and READ_ONCE()/WRITE_ONCE().
A working knowledge of loadable kernel modules, basic C, and how kernel threads are
created is assumed. You do not need any prior locking experience — that is exactly
what this lecture builds toward.
The Classic Trap: A Global Counter and i++
Picture a shared, global integer that several kernel threads increment concurrently.
A common assumption among new driver developers is that i++ is a single,
indivisible CPU operation, so it “must” be safe without any protection. That assumption
is wrong far more often than it is right, and understanding why is the foundation for
everything else in kernel synchronization.
At the C source level, i++ looks atomic. But the compiler does not translate
it into one instruction by default. On most architectures, incrementing a variable that
lives in memory (rather than purely in a CPU register) breaks down into three distinct
steps:
- Fetch — read the current value of
ifrom memory into a CPU register. - Increment — add 1 to the value inside the register.
- Store — write the updated register value back out to memory.
Between any two of these steps, the CPU’s scheduler can interrupt the current thread, a hardware interrupt can fire, or on a multi-core system another CPU can be executing the exact same three steps on the exact same variable at the exact same moment. If two threads both read the old value before either one writes the new value back, one increment is silently lost. This is the textbook definition of a race condition on a critical section.
Atomicity Depends on the ISA and the Compiler — Not On You
Whether i++ ends up as one instruction or three is determined by two things
that are completely outside the C source code:
- The processor’s Instruction Set Architecture (ISA). Some ISAs offer a single machine instruction capable of doing a read-modify-write on a memory location as one indivisible unit. Others do not, and require separate load, modify, and store instructions.
- The compiler and its optimization level. Even on an ISA that supports a single-instruction increment, the compiler has to actually choose to emit it. A build with no optimization frequently emits separate load/add/store instructions, while a higher optimization level may collapse them — and that behavior can change between compiler versions, compiler flags, and target architectures.
The practical consequence for kernel and driver developers is simple and important: you cannot look at a line of C and know whether it is atomic on the target hardware. Kernel code regularly runs on x86_64, multiple generations of ARM/ARM64, RISC-V, and other architectures, often built by different compiler versions with different optimization settings. The only safe engineering position is to assume every shared read-modify-write is unsafe unless you have explicitly made it safe — either with a proper locking primitive or with the kernel’s atomic integer/bitwise operators (which we introduce in the next lecture).
A Minimal Original Driver: Watching the Race Happen
The kernel module below is an original, minimal example built for kernel 6.x. It starts two kernel threads that each increment a shared, unprotected counter a large number of times. On a multi-core system you will typically see the final counter value fall short of the expected total, proving the lost-update problem is real and not just theoretical.
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/kthread.h>
#include <linux/delay.h>
#define LOOPS_PER_THREAD 2000000
static int shared_counter;
static struct task_struct *worker_a, *worker_b;
static int bump_counter(void *data)
{
int i;
for (i = 0; i < LOOPS_PER_THREAD; i++) {
shared_counter++; /* unprotected read-modify-write */
}
pr_info("ep_race: thread [%s] finished its loop\n", current->comm);
return 0;
}
static int __init ep_race_init(void)
{
shared_counter = 0;
worker_a = kthread_run(bump_counter, NULL, "ep_race_worker_a");
worker_b = kthread_run(bump_counter, NULL, "ep_race_worker_b");
if (IS_ERR(worker_a) || IS_ERR(worker_b)) {
pr_err("ep_race: failed to start worker threads\n");
return -ENOMEM;
}
pr_info("ep_race: expected final value = %d\n",
LOOPS_PER_THREAD * 2);
return 0;
}
static void __exit ep_race_exit(void)
{
msleep(500); /* give the threads a moment to finish on load/unload testing */
pr_info("ep_race: actual final value = %d\n", shared_counter);
}
module_init(ep_race_init);
module_exit(ep_race_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala: demonstrating i++ is not atomic");
Load this module on a multi-core VM or board, then unload it a second or two later and
check dmesg. The “expected” value is LOOPS_PER_THREAD * 2; the
“actual” value printed on unload is frequently smaller. That gap is every increment that
got silently overwritten by another CPU racing on the same critical section —
exactly the scenario described above.
Introducing the Lock: How We Fix It
The fix is conceptually simple, even though the kernel offers many different locking primitives (spinlocks, mutexes, semaphores, RCU, and more, which later lectures in this series cover in depth). Every one of them is built on the same core idea:
A lock guarantees that exactly one thread of execution can “hold” it at any given moment. Before entering a critical section, a thread must acquire the lock; when it leaves the critical section, it releases the lock. Any other thread that tries to acquire the lock while it is held must wait (or, depending on the primitive, be told the lock is busy). This turns a section of code that would otherwise run in parallel into a section that runs serialized — one thread at a time — which is exactly what removes the race.
Unprotected vs Protected: A Quick Comparison
| Aspect | Unprotected i++ | Locked critical section |
|---|---|---|
| Execution | Parallel, order not controlled | Serialized, one thread at a time |
| Result correctness | Depends on ISA/compiler — often wrong | Guaranteed correct |
| Portability across architectures | Not portable — behavior varies | Portable — correct on every architecture |
| Cost | None, but incorrect | Small overhead, but correct |
That small overhead is a price every serious kernel and device driver developer accepts
gladly, because a driver that occasionally corrupts shared state is far worse than one
that is marginally slower. In the next lecture of this free linux kernel
development course, we look at the kernel’s actual locking primitives, starting
with the lightweight spinlock_t, and then move on to the atomic integer
operators that let you avoid a full lock for simple counters altogether.
Frequently Asked Questions
1. Is i++ ever actually atomic on Linux?
Sometimes, depending on the architecture and compiler optimization level, but you cannot
rely on it. The only dependable approach is to protect it or use a dedicated atomic
operator.
2. What is a critical section in kernel programming?
A critical section is any piece of code that reads and/or writes shared, writeable state
that more than one thread of execution could touch at the same time.
3. Why does the compiler optimization level change atomicity?
Higher optimization levels can fold a load, add, and store into a single machine
instruction on ISAs that support it, while unoptimized builds usually keep them as
separate instructions.
4. Does this race condition only happen on multi-core systems?
Multi-core makes it far more likely, but even on a single core, interrupts, exceptions,
and preemption can interleave two read-modify-write sequences on the same variable.
5. What is the simplest fix for an unprotected shared counter?
For a plain integer counter, the kernel’s atomic integer type and operators (covered in
the next lecture) are usually simpler and cheaper than a full lock.
6. What does “serialized” mean in this context?
It means only one thread executes the critical section at a time; every other thread
that wants to enter must wait until the current holder releases the lock.
7. Do I need a lock even for a single-core embedded device?
Usually yes. Interrupts and preemption can still interleave a read-modify-write on a
single core, so the same “assume unsafe by default” rule applies.
8. Which locking primitive should I learn first?
Start with the spinlock, since it is the simplest and most widely used primitive for
short critical sections in the kernel; we cover it in the next lecture.
← Previous Lecture | Next Lecture →
Continue the Free Linux Kernel Development Course
More lectures on kernel locking primitives, atomic operators, and device driver synchronization are coming next in this free linux device drivers course.

2 Comments