← Previous Lecture | Next Lecture →
If two CPU cores update two completely different, unrelated variables at the same time, you would expect zero slowdown — after all, there is no data race and no lock is needed. Yet on real multicore hardware this exact scenario can quietly cripple performance. The reason is cache coherency, and the symptom is called false sharing linux kernel developers must learn to recognize early, because it hides inside code that looks perfectly correct.
What You Will Learn
- How CPU L1/L2/L3 caches and cachelines actually work on modern multicore hardware
- What a cache coherency protocol does when two cores touch the same cacheline
- Why “cache ping-pong” happens even without any shared variable or lock
- How to detect and fix false sharing using cacheline padding and alignment macros
- How to find your CPU’s real cacheline size on kernel 6.x systems
Prerequisites
- Basic understanding of Linux kernel modules (insmod/rmmod)
- Familiarity with kernel threads (kthreads) — covered in earlier lectures of this course
- A general idea of what a critical section and a data race are
A Quick Recap: What Is a CPU Cacheline?
Every modern CPU core does not read memory one byte or one variable at a time. Instead, it pulls in a fixed-size block of RAM — the cacheline — into its L1, L2, and often a shared L3 cache. On almost every x86_64 and ARM64 server or desktop chip shipping today, that block is 64 bytes. If two small variables happen to sit within the same 64-byte block, the CPU treats them as a single unit for caching purposes, even though your C code treats them as two completely independent integers.
How Cache Coherency Protocols React
Both CPU0 and CPU1 are allowed to keep their own private copy of that cacheline in L1 — that is the entire point of caching. The problem starts the moment either core writes to its copy. The hardware’s cache coherency protocol (a MESI-family protocol on virtually every mainstream CPU) must guarantee that no core is ever reading a stale value, so it enforces a simple rule: only one core may hold a cacheline in a “modified” state at a time.
Now picture kthread T1 running on CPU0 doing a++ repeatedly, while kthread T2 running on CPU1 does b++ repeatedly — two unrelated counters, no shared logic, no lock required by correctness rules. But because a and b live in the same cacheline:
- T1 writes
aon CPU0 → CPU0’s cacheline copy becomes Modified; CPU1’s copy of that same line must be invalidated. - T2 immediately writes
bon CPU1 → CPU1 must re-fetch the line (now owned by CPU0), forcing CPU0 to flush it back to shared memory / L3 first. - The line bounces back and forth between the two cores on every single write, even though
aandbare never touched by the same thread.
This constant bouncing is called false sharing — “false” because the two threads are not really sharing any data; they are only sharing a cacheline by unlucky memory layout. The cost is real: every ping-pong round trip can cost hundreds of CPU cycles, compared to a handful of cycles for a genuine cache hit.
Finding Your Real Cacheline Size on Kernel 6.x
Never hardcode 64 bytes blindly in production code — always confirm it for the target hardware. On a running kernel 6.x system:
$ getconf LEVEL1_DCACHE_LINESIZE
64
$ cat /sys/devices/system/cpu/cpu0/cache/index0/coherency_line_size
64
Inside kernel code, the constant L1_CACHE_BYTES (defined per-architecture) and the helper macro ____cacheline_aligned_in_smp exist precisely so you never have to guess this number by hand.
Original Demo: Fixing False Sharing With Cacheline Padding
Below is an original kernel module, not taken from any book, that shows the fix. The naive struct packs two per-thread counters next to each other; the fixed struct forces each counter onto its own cacheline using ____cacheline_aligned_in_smp, which is the modern kernel 6.x-supported way to do this (no manual padding bytes needed).
#include <linux/module.h>
#include <linux/cache.h>
/* Naive layout: BOTH counters can land on the same 64-byte line */
struct ep_naive_counters {
unsigned long cnt_a;
unsigned long cnt_b;
};
/* Fixed layout: each counter forced onto its own cacheline */
struct ep_padded_counters {
unsigned long cnt_a ____cacheline_aligned_in_smp;
unsigned long cnt_b ____cacheline_aligned_in_smp;
};
static struct ep_padded_counters ep_counters;
static int __init ep_cache_align_demo_init(void)
{
pr_info("ep_cache_align_demo: sizeof(naive)=%zu sizeof(padded)=%zu\n",
sizeof(struct ep_naive_counters),
sizeof(struct ep_padded_counters));
return 0;
}
static void __exit ep_cache_align_demo_exit(void)
{
pr_info("ep_cache_align_demo: unloaded\n");
}
module_init(ep_cache_align_demo_init);
module_exit(ep_cache_align_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala false sharing padding demo");
Load it and check dmesg:
$ sudo insmod ep_cache_align_demo.ko
$ dmesg | tail -1
ep_cache_align_demo: sizeof(naive)=16 sizeof(padded)=128
The padded struct is deliberately larger — that is the trade-off. You spend extra memory (up to one full cacheline per hot field) to guarantee two counters written by different CPUs never share a line again.
Naive vs Padded Layout Comparison
| Aspect | Naive Struct | Cacheline-Padded Struct |
|---|---|---|
| Memory used | Minimal | One cacheline per hot field |
| Cache ping-pong under concurrent writes | Severe | Eliminated |
| Correctness (data race) | Same (no race either way) | Same |
| Best used for | Rarely-written or single-thread data | Hot per-CPU / per-thread counters |
Common Mistakes
- Assuming “no shared variable” means “no performance issue.” False sharing has nothing to do with correctness — it is a pure performance bug that never shows up as a crash or a wrong value.
- Padding everything by default. Cacheline padding wastes memory; only apply it to fields that are genuinely hot-written from different CPUs concurrently.
- Hardcoding a cacheline size of 64 instead of using
L1_CACHE_BYTES/____cacheline_aligned_in_smp, which stay correct across architectures.
Best Practices
- Profile before padding — tools like
perf c2ccan directly point out cacheline contention hot spots. - Group read-mostly fields together, and separate frequently-written fields onto their own cachelines.
- Prefer per-CPU variables (covered in the next lecture of this free Linux kernel development course) over manually padded shared structs whenever each CPU truly needs its own private counter.
FAQ
Q1. Is false sharing a bug or just a performance issue?
It is purely a performance issue. The program’s output is correct either way; only the speed suffers.
Q2. Does false sharing only happen in kernel code?
No — it affects any multithreaded program, user-space or kernel-space, wherever independent hot variables share a cacheline.
Q3. How large is a cacheline on ARM64?
Commonly 64 bytes as well, but always verify with coherency_line_size since some ARM implementations differ.
Q4. Does padding always help?
Only when the padded fields are actually written concurrently from different CPUs. Padding cold or read-only fields wastes memory for no gain.
Q5. Is ____cacheline_aligned_in_smp available on kernel 6.x?
Yes, it remains the standard kernel macro for SMP-aware cacheline alignment on current kernel 6.x trees.
Q6. How do I detect false sharing on a running system?
Use perf c2c record / perf c2c report, which is built specifically to surface cacheline contention between CPUs.
Q7. Is this the same as a data race?
No. A data race is a correctness problem needing a lock. False sharing can exist even when there is zero data race.
Conclusion
Cache coherency is what keeps every CPU core’s view of memory consistent, but that same guarantee is what makes false sharing possible: two unrelated, individually race-free variables can still bounce a cacheline between cores on every write. Understanding L1_CACHE_BYTES and ____cacheline_aligned_in_smp gives you the tools to fix it in real kernel code. In the next lecture of this free Linux kernel development course, we move to a technique that avoids the sharing problem entirely — per-CPU variables.

2 Comments