Why i++ Is Not Atomic in the Linux Kernel-Free Linux Device Driver Training in Hyderabad

← Previous Lecture  |  Next Lecture →

Why i++ Is Not Atomic in the Linux Kernel
Free Linux Kernel Synchronization Tutorial — Kernel 6.x, Locks & Critical Sections
100% Free
Kernel 6.x Updated
Original Code Examples
free linux kernel development course free linux device drivers course free embedded systems course linux kernel synchronization tutorial kernel locking basics

This free linux kernel synchronization tutorial answers a question that trips up almost every new device driver author: why can a plain i++ on a shared integer corrupt data on a multi-core Linux system, even though the C statement looks like a single, harmless operation? In this lecture, part of our ongoing free linux kernel development course, we break the increment down into the real machine instructions the CPU executes, show why the answer depends on both the processor and the compiler, and use that insight to introduce the single most important synchronization idea in the kernel: the lock. Everything here targets current kernel 6.x internals and uses only original driver code, so you can follow along on any recent distribution.

What You Will Learn
Why i++ is three separate CPU operations How ISA + compiler decide atomicity Critical sections revisited The lock/unlock concept An original race-condition driver Why “assume unsafe by default” is the safe rule
Prerequisites

This lecture builds directly on the previous lecture in this free linux device drivers course, which covered critical sections, process context versus atomic context, and READ_ONCE()/WRITE_ONCE(). A working knowledge of loadable kernel modules, basic C, and how kernel threads are created is assumed. You do not need any prior locking experience — that is exactly what this lecture builds toward.

The Classic Trap: A Global Counter and i++

Picture a shared, global integer that several kernel threads increment concurrently. A common assumption among new driver developers is that i++ is a single, indivisible CPU operation, so it “must” be safe without any protection. That assumption is wrong far more often than it is right, and understanding why is the foundation for everything else in kernel synchronization.

At the C source level, i++ looks atomic. But the compiler does not translate it into one instruction by default. On most architectures, incrementing a variable that lives in memory (rather than purely in a CPU register) breaks down into three distinct steps:

  1. Fetch — read the current value of i from memory into a CPU register.
  2. Increment — add 1 to the value inside the register.
  3. Store — write the updated register value back out to memory.

Between any two of these steps, the CPU’s scheduler can interrupt the current thread, a hardware interrupt can fire, or on a multi-core system another CPU can be executing the exact same three steps on the exact same variable at the exact same moment. If two threads both read the old value before either one writes the new value back, one increment is silently lost. This is the textbook definition of a race condition on a critical section.

Atomicity Depends on the ISA and the Compiler — Not On You

Whether i++ ends up as one instruction or three is determined by two things that are completely outside the C source code:

  • The processor’s Instruction Set Architecture (ISA). Some ISAs offer a single machine instruction capable of doing a read-modify-write on a memory location as one indivisible unit. Others do not, and require separate load, modify, and store instructions.
  • The compiler and its optimization level. Even on an ISA that supports a single-instruction increment, the compiler has to actually choose to emit it. A build with no optimization frequently emits separate load/add/store instructions, while a higher optimization level may collapse them — and that behavior can change between compiler versions, compiler flags, and target architectures.

The practical consequence for kernel and driver developers is simple and important: you cannot look at a line of C and know whether it is atomic on the target hardware. Kernel code regularly runs on x86_64, multiple generations of ARM/ARM64, RISC-V, and other architectures, often built by different compiler versions with different optimization settings. The only safe engineering position is to assume every shared read-modify-write is unsafe unless you have explicitly made it safe — either with a proper locking primitive or with the kernel’s atomic integer/bitwise operators (which we introduce in the next lecture).

Why a Plain Increment Can Lose Updates
CPU 0: read i (value = 5)
CPU 1: read i (value = 5) <– both CPUs read the SAME old value
CPU 0: add 1 -> 6
CPU 1: add 1 -> 6
CPU 0: write i = 6
CPU 1: write i = 6 <– final value is 6, not 7. one increment is LOST

A Minimal Original Driver: Watching the Race Happen

The kernel module below is an original, minimal example built for kernel 6.x. It starts two kernel threads that each increment a shared, unprotected counter a large number of times. On a multi-core system you will typically see the final counter value fall short of the expected total, proving the lost-update problem is real and not just theoretical.

#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/kthread.h>
#include <linux/delay.h>

#define LOOPS_PER_THREAD  2000000

static int shared_counter;
static struct task_struct *worker_a, *worker_b;

static int bump_counter(void *data)
{
	int i;

	for (i = 0; i < LOOPS_PER_THREAD; i++) {
		shared_counter++;   /* unprotected read-modify-write */
	}

	pr_info("ep_race: thread [%s] finished its loop\n", current->comm);
	return 0;
}

static int __init ep_race_init(void)
{
	shared_counter = 0;

	worker_a = kthread_run(bump_counter, NULL, "ep_race_worker_a");
	worker_b = kthread_run(bump_counter, NULL, "ep_race_worker_b");

	if (IS_ERR(worker_a) || IS_ERR(worker_b)) {
		pr_err("ep_race: failed to start worker threads\n");
		return -ENOMEM;
	}

	pr_info("ep_race: expected final value = %d\n",
		LOOPS_PER_THREAD * 2);
	return 0;
}

static void __exit ep_race_exit(void)
{
	msleep(500);  /* give the threads a moment to finish on load/unload testing */
	pr_info("ep_race: actual final value   = %d\n", shared_counter);
}

module_init(ep_race_init);
module_exit(ep_race_exit);

MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala: demonstrating i++ is not atomic");

Load this module on a multi-core VM or board, then unload it a second or two later and check dmesg. The “expected” value is LOOPS_PER_THREAD * 2; the “actual” value printed on unload is frequently smaller. That gap is every increment that got silently overwritten by another CPU racing on the same critical section — exactly the scenario described above.

Introducing the Lock: How We Fix It

The fix is conceptually simple, even though the kernel offers many different locking primitives (spinlocks, mutexes, semaphores, RCU, and more, which later lectures in this series cover in depth). Every one of them is built on the same core idea:

A lock guarantees that exactly one thread of execution can “hold” it at any given moment. Before entering a critical section, a thread must acquire the lock; when it leaves the critical section, it releases the lock. Any other thread that tries to acquire the lock while it is held must wait (or, depending on the primitive, be told the lock is busy). This turns a section of code that would otherwise run in parallel into a section that runs serialized — one thread at a time — which is exactly what removes the race.

A Critical Section Protected by a Lock
Thread A: LOCK -> acquired
Thread A: read shared_counter, add 1, write shared_counter
Thread A: UNLOCK
Thread B: LOCK -> (had to wait) now acquired
Thread B: read shared_counter, add 1, write shared_counter
Thread B: UNLOCK
Result: every increment counted. no update lost.

Unprotected vs Protected: A Quick Comparison

Aspect Unprotected i++ Locked critical section
Execution Parallel, order not controlled Serialized, one thread at a time
Result correctness Depends on ISA/compiler — often wrong Guaranteed correct
Portability across architectures Not portable — behavior varies Portable — correct on every architecture
Cost None, but incorrect Small overhead, but correct

That small overhead is a price every serious kernel and device driver developer accepts gladly, because a driver that occasionally corrupts shared state is far worse than one that is marginally slower. In the next lecture of this free linux kernel development course, we look at the kernel’s actual locking primitives, starting with the lightweight spinlock_t, and then move on to the atomic integer operators that let you avoid a full lock for simple counters altogether.

Frequently Asked Questions

1. Is i++ ever actually atomic on Linux?
Sometimes, depending on the architecture and compiler optimization level, but you cannot rely on it. The only dependable approach is to protect it or use a dedicated atomic operator.

2. What is a critical section in kernel programming?
A critical section is any piece of code that reads and/or writes shared, writeable state that more than one thread of execution could touch at the same time.

3. Why does the compiler optimization level change atomicity?
Higher optimization levels can fold a load, add, and store into a single machine instruction on ISAs that support it, while unoptimized builds usually keep them as separate instructions.

4. Does this race condition only happen on multi-core systems?
Multi-core makes it far more likely, but even on a single core, interrupts, exceptions, and preemption can interleave two read-modify-write sequences on the same variable.

5. What is the simplest fix for an unprotected shared counter?
For a plain integer counter, the kernel’s atomic integer type and operators (covered in the next lecture) are usually simpler and cheaper than a full lock.

6. What does “serialized” mean in this context?
It means only one thread executes the critical section at a time; every other thread that wants to enter must wait until the current holder releases the lock.

7. Do I need a lock even for a single-core embedded device?
Usually yes. Interrupts and preemption can still interleave a read-modify-write on a single core, so the same “assume unsafe by default” rule applies.

8. Which locking primitive should I learn first?
Start with the spinlock, since it is the simplest and most widely used primitive for short critical sections in the kernel; we cover it in the next lecture.

← Previous Lecture  |  Next Lecture →

Continue the Free Linux Kernel Development Course

More lectures on kernel locking primitives, atomic operators, and device driver synchronization are coming next in this free linux device drivers course.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *