CPU Cache Effects and False Sharing in Linux Kernel Programming-Free Linux Device Driver Course Online

« Previous Lecture | Next Lecture »

CPU Cache Effects and False Sharing in Linux Kernel Programming

Free Linux Kernel Programming Course — Kernel Synchronization Part 2

Kernel 6.x Updated
Original Demo
Beginner Friendly

Correct locking is only half the story in concurrent Linux kernel code. Even a perfectly locked data structure can quietly kill performance if it triggers false sharing — a CPU cache effect where two unrelated variables happen to sit on the same cacheline and end up fighting each other across cores. This lecture in our free Linux kernel development course explains how CPU caches work, what false sharing is, and how to fix it with cacheline alignment, illustrated with an original kernel module.

false sharing linux kernel cpu cacheline alignment cache ping-pong free linux device drivers course

What You Will Learn

  • How multi-level CPU caches (L1/L2/L3) and cachelines actually work
  • What “false sharing” and “cache ping-pong” mean, with a concrete diagram
  • How to detect a false-sharing-prone layout in your own driver’s data structures
  • How to fix it using cacheline padding and alignment on modern kernel 6.x
  • An original kernel module demonstrating a naive layout versus a cache-friendly layout

Prerequisites

This lecture assumes familiarity with spinlocks and atomic operations from earlier in this free Linux kernel programming course. No prior computer-architecture background is required — we build the cache concepts up from scratch.

How CPU Caches Actually Work

Modern CPUs never fetch a single byte from RAM. Whenever the processor needs one byte at some address, the memory controller pulls in an entire cacheline — typically 64 bytes on most current x86_64 and ARM64 systems — into the CPU’s cache hierarchy (commonly L1, L2, and a shared L3). This is what makes sequential memory access fast: once the first byte of a cacheline is fetched, the next several accesses to nearby addresses are served straight out of cache.

Roughly speaking, an L1 cache hit costs a handful of nanoseconds, while a full round trip to RAM can cost anywhere from around 50 to 100 nanoseconds depending on the hardware — a difference of well over an order of magnitude. This is exactly why kernel developers are taught to keep frequently accessed structure members together near the top of a struct, so they land in the same cacheline.

A 64-Byte Cacheline Loaded From RAM
RAM address: [ 0][ 1][ 2] … [63] entire 64-byte line is pulled into L1/L2/L3 (bytes 0-63 are now all “hot” in cache)

What Is False Sharing?

False sharing happens when two variables that are logically unrelated — used by completely different threads or CPUs — happen to be placed close enough together in memory that they land on the same cacheline. Even though neither CPU is actually touching the other CPU’s variable, the cache-coherency hardware has no idea the two variables are “unrelated”: it only knows the cacheline as a whole was modified, so it must invalidate that line in every other core’s cache and reload it.

Imagine two counters, counter_a and counter_b, declared next to each other in a struct. CPU 0 increments counter_a in a tight loop; CPU 1 increments counter_b in its own tight loop. Neither core ever reads or writes the other’s counter — yet performance collapses, because every write from either core forces the cacheline out of the other core’s cache. This constant back-and-forth invalidation is called cache ping-pong.

False Sharing: Two Unrelated Counters, One Cacheline
struct { u64 counter_a; u64 counter_b; } cacheline invalidated in CPU1’s cache CPU1 writes counter_b –> cacheline invalidated in CPU0’s cache CPU0 writes counter_a –> cacheline invalidated in CPU1’s cache again … this “ping-pong” repeats on every single write, even though CPU0 and CPU1 never touch each other’s variable

Fixing False Sharing With Cacheline Padding

The fix is straightforward once you see the problem: force each hot, per-CPU or per-thread variable onto its own cacheline, using padding or alignment attributes. The kernel provides the ____cacheline_aligned_in_smp annotation (and the plain ____cacheline_aligned for non-SMP-sensitive cases) specifically for this purpose, alongside the standard L1_CACHE_BYTES constant.

ToolPurpose
L1_CACHE_BYTESThe cacheline size for the current architecture, usually 64
____cacheline_aligned_in_smpAligns a structure or field to a cacheline boundary, but only actually pads on SMP builds
Manual padding arrayAn explicit u8 pad[...] field for full control over layout

An Original Demo: Naive Layout vs Cacheline-Padded Layout

The module below defines two per-CPU-style counter structures side by side: one naive layout where the counters share a cacheline, and one padded layout where each counter gets its own cacheline. Two kernel threads increment their own counter in a tight loop for a fixed duration, and the module reports the total increments achieved by each layout through debugfs — on real SMP hardware the padded version consistently completes noticeably more iterations in the same time window.

#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/debugfs.h>
#include <linux/jiffies.h>
#include <linux/cache.h>

/* Naive: both counters likely share one cacheline */
struct ep_naive_counters {
	u64 counter_a;
	u64 counter_b;
};

/* Padded: each counter forced onto its own cacheline */
struct ep_padded_counters {
	u64 counter_a;
} ____cacheline_aligned_in_smp;

struct ep_padded_counters_b {
	u64 counter_b;
} ____cacheline_aligned_in_smp;

static struct ep_naive_counters ep_naive;
static struct ep_padded_counters ep_padded_a;
static struct ep_padded_counters_b ep_padded_b;

static struct task_struct *ep_thr_naive_a, *ep_thr_naive_b;
static struct task_struct *ep_thr_padded_a, *ep_thr_padded_b;

#define EP_RUN_JIFFIES (HZ / 2)

static int ep_bump_naive_a(void *data)
{
	unsigned long end = jiffies + EP_RUN_JIFFIES;

	while (time_before(jiffies, end) && !kthread_should_stop())
		ep_naive.counter_a++;
	return 0;
}

static int ep_bump_naive_b(void *data)
{
	unsigned long end = jiffies + EP_RUN_JIFFIES;

	while (time_before(jiffies, end) && !kthread_should_stop())
		ep_naive.counter_b++;
	return 0;
}

static int ep_bump_padded_a(void *data)
{
	unsigned long end = jiffies + EP_RUN_JIFFIES;

	while (time_before(jiffies, end) && !kthread_should_stop())
		ep_padded_a.counter_a++;
	return 0;
}

static int ep_bump_padded_b(void *data)
{
	unsigned long end = jiffies + EP_RUN_JIFFIES;

	while (time_before(jiffies, end) && !kthread_should_stop())
		ep_padded_b.counter_b++;
	return 0;
}

static int ep_report_show(struct seq_file *s, void *v)
{
	seq_printf(s, "naive:  a=%llu b=%llu (total=%llu)\n",
		   ep_naive.counter_a, ep_naive.counter_b,
		   ep_naive.counter_a + ep_naive.counter_b);
	seq_printf(s, "padded: a=%llu b=%llu (total=%llu)\n",
		   ep_padded_a.counter_a, ep_padded_b.counter_b,
		   ep_padded_a.counter_a + ep_padded_b.counter_b);
	return 0;
}
DEFINE_SHOW_ATTRIBUTE(ep_report);

static struct dentry *ep_dbgdir;

static int __init ep_false_sharing_demo_init(void)
{
	ep_dbgdir = debugfs_create_dir("ep_false_sharing_demo", NULL);
	debugfs_create_file("report", 0444, ep_dbgdir, NULL, &ep_report_fops);

	ep_thr_naive_a  = kthread_run(ep_bump_naive_a, NULL, "ep_naive_a");
	ep_thr_naive_b  = kthread_run(ep_bump_naive_b, NULL, "ep_naive_b");
	ep_thr_padded_a = kthread_run(ep_bump_padded_a, NULL, "ep_padded_a");
	ep_thr_padded_b = kthread_run(ep_bump_padded_b, NULL, "ep_padded_b");

	pr_info("ep_false_sharing_demo: loaded, threads racing for %d jiffies\n",
		EP_RUN_JIFFIES);
	return 0;
}

static void __exit ep_false_sharing_demo_exit(void)
{
	kthread_stop(ep_thr_naive_a);
	kthread_stop(ep_thr_naive_b);
	kthread_stop(ep_thr_padded_a);
	kthread_stop(ep_thr_padded_b);
	debugfs_remove_recursive(ep_dbgdir);
	pr_info("ep_false_sharing_demo: unloaded\n");
}

module_init(ep_false_sharing_demo_init);
module_exit(ep_false_sharing_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original false sharing benchmark demo");

Read /sys/kernel/debug/ep_false_sharing_demo/report right after loading the module (while the four threads are still finishing their half-second run) to compare the total increments achieved by the naive, cacheline-sharing layout versus the padded, cacheline-isolated layout.

Real-World Use Cases

  • Per-CPU statistics counters in high-throughput drivers and network paths
  • Lock structures placed adjacent to frequently-written unrelated fields inside a device’s private data structure
  • Ring-buffer producer and consumer index fields, which are classic false-sharing hotspots if placed next to each other

Performance Considerations

Padding trades memory for speed — every padded field can consume up to a full cacheline (64 bytes) even though the field itself might only be 4 or 8 bytes. This is almost always worth it for hot, frequently-written, multi-core-contended fields, but blindly padding every structure member in your driver wastes memory and cache capacity for no benefit. Reserve cacheline alignment for fields you have actually identified as contended across cores.

Common Mistakes

MistakeWhy it’s a problemFix
Declaring unrelated per-CPU-style counters adjacent in a structCauses false sharing and cache ping-pong under concurrent accessPad each hot field to its own cacheline with ____cacheline_aligned_in_smp
Padding every field “just in case”Wastes memory and cache footprint without measurable benefitProfile first, pad only fields shown to be contended
Assuming false sharing is a correctness bugIt is purely a performance issue — the code is still correct, just slowTreat it as a profiling/optimization task, not a locking bug

Best Practices

  • Group logically related, frequently-accessed fields together near the top of a structure.
  • Isolate independently-written, per-core or per-thread hot fields onto separate cachelines.
  • Use L1_CACHE_BYTES and the ____cacheline_aligned_in_smp annotation rather than hardcoding a padding size.
  • Measure before and after — false sharing effects are only visible under real concurrent load on real SMP hardware.

Summary

CPU caches load data in fixed-size cachelines, which is normally a huge performance win — but when unrelated variables written by different cores share a cacheline, the cache-coherency protocol forces constant invalidation between cores, a phenomenon called false sharing or cache ping-pong. The fix is to align hot, independently-written fields onto separate cachelines using the kernel’s cacheline-alignment tools.

Frequently Asked Questions

What is a cacheline in simple terms?

It is the fixed-size chunk of memory, typically 64 bytes, that the CPU always fetches from RAM as a unit, even if you only asked for a single byte.

Is false sharing a bug that produces wrong results?

No. The code still produces correct results; false sharing is purely a performance problem caused by unnecessary cache invalidation traffic between cores.

How do I know if my driver’s data structure suffers from false sharing?

Profile it under real concurrent load on multi-core hardware and look for unexpectedly poor scaling as more cores are added, then inspect which fields different cores are writing.

Does padding every structure field prevent false sharing entirely?

It would, but at a significant memory cost. Pad only the fields you have identified as genuinely hot and contended across cores.

Is ____cacheline_aligned_in_smp different on a single-CPU (UP) kernel build?

Yes — on a non-SMP build the padding is compiled out entirely, since there is only one CPU core and no cache-coherency traffic to avoid.

Conclusion

Understanding CPU cache behaviour and false sharing rounds out the performance side of kernel synchronization, complementing the correctness-focused locking primitives covered earlier in this free Linux kernel development course. Getting your locking right and your data layout cache-friendly together is what separates a driver that merely works from one that scales well on modern multi-core hardware.

« Previous Lecture | Next Lecture »

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *