← Previous Lecture | Next Lecture →
If you are following our free Linux kernel development course, you already know that a critical section must be protected whenever more than one thread can touch the same shared writeable data. What we have not discussed yet is what protection actually costs you at runtime. In this lecture we look closely at linux kernel locking overhead, why a lock always creates a bottleneck no matter how cleverly it is written, and a subtle bug called a torn read that catches out even experienced driver developers. Everything here is rebuilt from scratch for kernel 6.x, with original driver code you will not find copied from any book or repository.
What You Will Learn
- Why a lock always turns parallel execution into serial execution
- How to picture lock contention as a bottleneck instead of memorizing the definition
- What a torn read (also called a dirty read) actually is, with a working demonstration
- Why “I am only reading the variable, so it must be safe” is a dangerous assumption
- How real-time scheduling (SCHED_FIFO) is used to get close to atomicity in user space
- The exact list of situations where you do not need explicit locking
- The four classic triggers that introduce concurrency concerns into kernel code
Prerequisites
- Comfortable building and loading a loadable kernel module (LKM) on kernel 6.x
- Understanding of what a critical section is and why it needs protection (covered in the previous lecture of this free linux kernel development course)
- Basic familiarity with kernel threads (kthread_create/kthread_run)
- A test VM or board running a recent 6.x kernel with CONFIG_SMP enabled
Locking Always Costs You Parallelism
Picture n threads all wanting to enter the same critical section at the same moment. Exactly one of them is allowed to become the “owner” and proceed; the remaining n-1 threads must wait. The moment the owner finishes and unlocks, one of the waiting threads becomes the new owner, and the cycle repeats until every thread has had its turn. Notice what this means: no matter how many CPUs your system has, the code inside the critical section runs one thread at a time, always. The rest of your program can run in glorious parallel on every core you own, but the moment execution enters the locked region, it collapses down to a single lane.
A useful mental model is a wide multi-lane highway that suddenly narrows into a single-lane toll booth. Every car (thread) can drive freely in its own lane until it reaches the booth. At the booth, cars are forced into single file, one after another, regardless of how many lanes fed into it. This is exactly what a lock does to your threads, and it is why linux kernel locking overhead grows quickly as contention increases — the busier the booth, the longer the queue.
This does not mean locking is bad. It means locking should be used only where it is genuinely required, and the amount of code inside the critical section (the “lock hold time”) should be kept as small as possible. A common design goal in real driver code is to minimize the use of global/shared writeable state in the first place, because less shared state means fewer critical sections, which means less serialization.
Torn Reads: The Bug Hiding in “Read-Only” Code
A newer kernel developer often assumes that reading a shared variable is always safe, and that only writers need to lock. This is false, and the failure mode has a name: a torn read (sometimes called a dirty read). It happens when the size of the data item is larger than what the CPU can load or store in a single, indivisible instruction. A 64-bit counter on a 32-bit machine is the textbook case — updating it actually takes two separate store instructions (the low word and the high word), not one.
If a reader thread loads that same 64-bit value at the exact moment a writer thread is halfway through updating it, the reader can end up with a value that never existed: half of the old number stitched to half of the new number. That is a torn read.
Below is an original demonstration module for kernel 6.x. It runs two kernel threads: one writer thread that repeatedly updates a shared 64-bit counter split into two 32-bit halves (to simulate a non-atomic update pattern on purpose), and one reader thread that reads it back without any protection. Do not run this on a production system — it exists purely to make the bug visible in your kernel log.
#include <linux/kthread.h>
#include <linux/delay.h>
#include <linux/module.h>
struct split_counter {
u32 low;
u32 high;
};
static struct split_counter shared_counter;
static struct task_struct *writer_thread, *reader_thread;
static int writer_fn(void *arg)
{
u64 value = 0;
while (!kthread_should_stop()) {
value++;
/* Intentionally unprotected two-step update - demonstrates the bug */
shared_counter.low = (u32)(value & 0xFFFFFFFF);
udelay(50); /* widen the window so the race is easy to observe */
shared_counter.high = (u32)(value >> 32);
usleep_range(1000, 2000);
}
return 0;
}
static int reader_fn(void *arg)
{
while (!kthread_should_stop()) {
u32 low = shared_counter.low;
u32 high = shared_counter.high;
u64 seen = ((u64)high << 32) | low;
pr_info("torn_read_demo: reader saw value = %llu\n", seen);
usleep_range(1000, 2000);
}
return 0;
}
static int __init torn_read_demo_init(void)
{
writer_thread = kthread_run(writer_fn, NULL, "torn_writer");
reader_thread = kthread_run(reader_fn, NULL, "torn_reader");
return 0;
}
static void __exit torn_read_demo_exit(void)
{
kthread_stop(writer_thread);
kthread_stop(reader_thread);
}
module_init(torn_read_demo_init);
module_exit(torn_read_demo_exit);
MODULE_LICENSE("GPL");
Watch dmesg while this module runs and you will occasionally see a value jump backwards or forwards in a way that plain incrementing cannot explain — that jump is the torn read in action. The fix, which we build in the next lecture, is a spinlock (or on modern kernels, an atomic64_t) wrapped around both the write side and the read side.
Why the Same Lock Must Protect Reads Too
A closely related mistake is protecting the writer with a lock but leaving the reader unprotected because “reading doesn’t change anything.” As the example above shows, the reader can still observe an inconsistent, half-updated value. Every access path to a shared writeable item — read or write — must go through the same lock, or the protection is meaningless.
Getting Close to Atomicity in User Space
In kernel space, the kernel gives you real tools (spinlocks, mutexes, atomic types) to make a block of code behave atomically. In user space, true atomicity across more than one CPU instruction is not something you can guarantee at all. The closest a user-space application can get is to make itself effectively un-preemptible: assign the thread the SCHED_FIFO real-time policy with priority 99. At that priority, nothing except hardware interrupts and exceptions can take the CPU away from it, so for practical purposes it runs “atomically” for as long as it chooses to run. This technique still shows up today in low-latency audio and real-time control applications, though on a modern PREEMPT_RT-enabled kernel you also want to be careful about how long you hold that priority, since starving other real-time tasks is its own class of bug.
#include <pthread.h>
#include <sched.h>
#include <stdio.h>
static void make_thread_real_time(void)
{
struct sched_param param;
param.sched_priority = 99;
if (pthread_setschedparam(pthread_self(), SCHED_FIFO, ¶m) != 0)
perror("pthread_setschedparam failed - are you running as root?");
}
In kernel space you rarely need this trick, because the kernel offers spinlocks that give you genuine, guaranteed atomicity for a critical section. We start working with spinlocks hands-on in the very next lecture of this course.
Key Rules for Protecting a Critical Section
| Rule | Why It Matters |
|---|---|
| Identify every critical section in your code base | You cannot protect what you have not noticed |
| Protect both reads and writes of shared data | Unprotected reads cause torn/dirty reads, as shown above |
| Always use the same lock for a given data item | Two different locks guarding one variable give you zero real protection |
| Keep the critical section as short as possible | Shorter lock-hold time means less serialization and less overhead |
| Design to minimize shared writeable state | Less shared state means fewer critical sections in the first place |
When You Do Not Need Explicit Protection
Locking is not required everywhere. There are three well-defined exceptions worth memorizing:
| Situation | Why It Is Safe Without a Lock |
|---|---|
| Local variables | Allocated on the private stack of the thread (or the local IRQ stack in interrupt context), so no other thread can see them |
| Code that can only ever run once, serially | A module’s module_init() and module_exit() functions run exactly once, on insmod/rmmod, and can never run concurrently with themselves |
| Data that is truly constant and read-only | Nothing ever writes to it, so there is no possibility of a race — but remember, the C const keyword alone does not guarantee this at the hardware or logical level |
Four Triggers That Introduce Concurrency Concerns
When you review kernel or driver code and ask “could this need locking?”, check for these four classic triggers:
- SMP systems — kernels built with
CONFIG_SMPcan genuinely run two threads on two different CPUs at the exact same instant - A preemptible kernel — even on a single CPU, the scheduler can switch out your thread mid-critical-section unless you disable preemption
- Blocking I/O — any call that can sleep gives the CPU to another thread, which may touch the same data before you resume
- Hardware interrupts — an interrupt handler can run on top of your code, on either an SMP or a uniprocessor (UP) system, and touch the same shared state
Common Mistakes and Troubleshooting Tips
- Mistake: locking only the writer and assuming the reader is safe. Fix: lock every access path to the shared item.
- Mistake: using two different locks (or no lock at all in one path) for the same variable. Fix: pick one lock per data item and use it everywhere that item is touched.
- Mistake: holding a lock across a long operation (large memcpy, logging, sleeping calls). Fix: copy what you need, release the lock, then do the slow work outside the critical section.
- Troubleshooting tip: if you see impossible values in a counter or struct, suspect a torn read before you suspect anything else — especially on 32-bit builds handling 64-bit data.
Best Practices
- Minimize the amount of shared writeable (global) state in your driver design
- Treat every field you plan to share across threads or interrupt context as a critical section candidate from day one
- Prefer smaller, well-scoped locks over one giant lock guarding unrelated data
- Document, next to the declaration of a shared variable, exactly which lock protects it
Performance Considerations
Lock hold time is the single biggest lever you control. The longer a critical section runs, the longer every other contending thread queues at the “toll booth.” On a busy multi-core system this shows up as CPUs spinning or sleeping while waiting instead of doing useful work, which directly lowers your effective throughput. Measuring lock contention (which we cover with kernel lock debugging tools later in this course) is the right way to find out whether your locking strategy is actually costing you performance.
Security Considerations
Torn or dirty reads are not just a correctness bug — in code that makes security-relevant decisions (permission checks, buffer length fields, reference counts) a torn read can be turned into a real vulnerability, since the “impossible” value read by an attacker-influenced path may bypass a check that assumed the value could only ever be fully old or fully new. Treat unprotected access to shared decision-making data as a security bug, not just a reliability one.
Summary / Key Takeaways
- Locking always serializes the critical section, no matter how many CPUs you have — this is unavoidable overhead, not a bug
- A torn/dirty read happens when a multi-instruction read races against a multi-instruction write on the same data item
- Reads need the same protection as writes; “read-only” access to shared data is not automatically safe
- User space can only approximate atomicity (SCHED_FIFO priority 99); kernel space can guarantee it with spinlocks
- Local variables, single-run init/exit code, and truly constant data do not need locking
- SMP, preemptible kernels, blocking I/O, and hardware interrupts are the four triggers to check for concurrency concerns
Conclusion
Understanding the cost of locking is just as important as knowing how to lock. Once you accept that every lock creates a bottleneck, you naturally start designing drivers with less shared state and shorter critical sections, which is exactly the mindset this free linux kernel development course is building toward. In the next lecture we finally get hands-on with the spinlock API itself — initialization, locking, unlocking, and the IRQ-safe variants you need for real driver code. If you are also working through our free linux device drivers course and free embedded systems course material, this concept of protected critical sections will come up again and again, so it is worth re-reading the torn read example until it clicks.
FAQ
Q1. Does locking always slow down my driver?
Locking only slows down the specific threads that are contending for the same lock at the same time. Code outside the critical section still runs fully in parallel.
Q2. Can a torn read happen on a 64-bit CPU too?
It is far less likely for a naturally aligned 64-bit value on a true 64-bit machine, but it can still occur with larger structs, misaligned data, or compiler-generated multi-instruction sequences, so protection is still recommended for any shared struct larger than a single machine word.
Q3. Is const in C enough to guarantee a variable is safe without locking?
No. const only stops your own code from writing through that pointer; it does not guarantee the underlying memory is never modified by another path, so you must independently confirm the data is truly immutable.
Q4. Why does module_init() not need locking?
Because insmod guarantees it runs exactly once, and rmmod guarantees module_exit() also runs exactly once, with no possibility of either running concurrently with itself.
Q5. What is the practical difference between a preemptible kernel and an SMP kernel when it comes to locking?
SMP means two threads can truly run at the same physical instant on different CPUs. A preemptible kernel means even a single CPU can switch your thread out mid-critical-section, so both need protection, just for different reasons.
Q6. Can I use SCHED_FIFO priority 99 instead of a spinlock in kernel code?
No, SCHED_FIFO is a user-space workaround for the fact that user space has no guaranteed atomicity primitive. Kernel code should use the real synchronization primitives such as spinlocks, mutexes, or atomic types instead.
Q7. What is the very next topic in this free Linux kernel development course?
The next lecture introduces the spinlock API itself — how to declare, initialize, lock, and unlock a spinlock, plus the IRQ-safe variants used when a critical section can also be touched from interrupt context.
Continue your free linux kernel development course journey with EmbeddedPathashala.

2 Comments