← Previous Lecture | Next Lecture →
Every lock you add to the Linux kernel is, in effect, a checkpoint where CPUs must queue up. On a system with a handful of cores that cost is small; on a modern server with a few hundred cores it can dominate everything. Per-CPU variables linux kernel developers rely on are one of the kernel’s core lock-free techniques for sidestepping this bottleneck entirely, by giving every CPU its own private copy of the data instead of making them fight over a shared one.
What You Will Learn
- Why locking does not scale well on high core-count systems
- What lock contention and lock proliferation mean in practice
- What a per-CPU variable is and how it eliminates the critical section
- How to statically declare a per-CPU variable with
DEFINE_PER_CPU - The difference between
DEFINE_PER_CPUandDECLARE_PER_CPU
Prerequisites
- Understanding of critical sections and why shared data needs protection (earlier lectures)
- Basic familiarity with spinlocks and mutexes
- Comfortable building and loading a simple kernel module
Why Locking Alone Does Not Scale
A lock is effectively a single-lane checkpoint: whichever CPU wants to touch the protected data has to wait its turn, one at a time, no matter how many cores the machine has. On a two- or four-core embedded board this rarely matters. But scale that same idea up to a server with a few hundred cores, and every thread wanting the same lock funnels through that one narrow point — the more cores you add, the worse the queue gets. This is usually described as “locking does not scale.”
The obvious fix — splitting one big lock into many smaller, finer-grained locks — helps reduce contention between any two threads, but it does not scale forever either. A kernel already manages many thousands of locks internally, and the more locks that exist, the higher the chance that two code paths acquire them in a different order and hit a subtle deadlock.
What Exactly Is a Per-CPU Variable?
A per-CPU (PCP) variable is not one variable at all — under the hood it is an array with one private element per live CPU on the system. On a 4-core system, a per-CPU int is really 4 separate integers, one that belongs exclusively to CPU0, one to CPU1, and so on. A thread running on CPU2 only ever touches CPU2’s private slot.
Because no two CPUs ever write to the same slot, the entire critical section disappears — there is nothing left to protect, and therefore no lock is needed. This is conceptually the same idea as a local variable on a thread’s private stack, or POSIX Thread-Local Storage (TLS) in user space: private-per-execution-context data never needs a lock.
Statically Declaring a Per-CPU Variable
The Linux kernel offers two macros for compile-time (static) per-CPU allocation, both from <linux/percpu.h>:
| Macro | Purpose |
|---|---|
DEFINE_PER_CPU(type, name) | Allocates and zero-initializes the per-CPU variable — use this in the .c file that owns it |
DECLARE_PER_CPU(type, name) | Only declares an extern reference to a per-CPU variable defined elsewhere — use this in a header shared across files |
Original Demo: Declaring a Per-CPU Packet Counter
Here is a small, original kernel module (not copied from any book) that statically defines a per-CPU counter and reports how many live CPUs the system currently has — the exact number of private slots that counter owns.
#include <linux/module.h>
#include <linux/percpu.h>
#include <linux/cpumask.h>
/* One private "pkt_count" slot per live CPU - no lock required */
DEFINE_PER_CPU(int, pkt_count);
static int __init ep_percpu_demo_init(void)
{
int cpu;
pr_info("ep_percpu_demo: live CPUs = %d\n", num_online_cpus());
for_each_online_cpu(cpu) {
pr_info("ep_percpu_demo: initial pkt_count on CPU%d = %d\n",
cpu, per_cpu(pkt_count, cpu));
}
return 0;
}
static void __exit ep_percpu_demo_exit(void)
{
pr_info("ep_percpu_demo: unloaded\n");
}
module_init(ep_percpu_demo_init);
module_exit(ep_percpu_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala static per-CPU variable demo");
Build and load it on a 4-core kernel 6.x system:
$ sudo insmod ep_percpu_demo.ko
$ dmesg | tail -5
ep_percpu_demo: live CPUs = 4
ep_percpu_demo: initial pkt_count on CPU0 = 0
ep_percpu_demo: initial pkt_count on CPU1 = 0
ep_percpu_demo: initial pkt_count on CPU2 = 0
ep_percpu_demo: initial pkt_count on CPU3 = 0
Note: This lecture only covers static allocation and inspection with per_cpu() for a fixed, known CPU number. The kernel-safe way to read and modify the currently running CPU’s own slot — using helpers like get_cpu_var()/put_cpu_var() and this_cpu_* operations — is covered in the next lecture of this free Linux kernel development course.
Per-CPU Variables vs Locking vs TLS
| Technique | Needs a lock? | Memory cost | Best for |
|---|---|---|---|
| Shared variable + spinlock/mutex | Yes | Low | Data genuinely shared/read by all CPUs |
| Per-CPU variable | No | One copy per live CPU | Independent counters/statistics per CPU |
User-space TLS (__thread) | No | One copy per thread | Per-thread private data in user space |
Common Mistakes
- Using per-CPU variables for large data structures. Since the item is duplicated once per CPU, large structs waste significant memory on systems with hundreds of cores.
- Reaching into a per-CPU variable directly by pointer instead of using the kernel-provided accessor macros — always use the official helper API, never raw pointer arithmetic.
- Forgetting that per-CPU does not mean per-thread. If a thread migrates to a different CPU mid-access, it can end up reading a different slot than intended unless preemption is properly disabled — the accessor macros exist to prevent exactly this bug.
Best Practices
- Reach for per-CPU variables for small, frequently-updated statistics (packet counts, hit/miss counters, latency samples).
- Use
DECLARE_PER_CPUin headers andDEFINE_PER_CPUin exactly one .c file, just like any other extern/definition pair. - Always use kernel-provided per-CPU accessor APIs rather than touching the underlying array yourself.
FAQ
Q1. Are per-CPU variables truly lock-free?
Yes — since each CPU only ever touches its own private slot, there is no shared critical section left to protect.
Q2. How many copies does a per-CPU variable actually create?
Exactly one copy per live CPU core on the running system, reported by num_online_cpus().
Q3. What is the difference between DEFINE_PER_CPU and DECLARE_PER_CPU?
DEFINE_PER_CPU allocates and zero-initializes storage; DECLARE_PER_CPU only declares an extern reference for use in other files.
Q4. Should I use per-CPU variables for large structs?
No — since the data is duplicated per CPU, large structs should stay as shared, lock-protected data instead.
Q5. Is per-CPU similar to user-space Thread-Local Storage?
Conceptually yes — both give each execution context its own private copy so no locking is required.
Q6. Can I read another CPU’s per-CPU slot?
Yes, using per_cpu(variable, cpu_number) for a specific CPU, though this is mainly used for reporting/debugging rather than the hot path.
Q7. Does per-CPU replace spinlocks and mutexes entirely?
No — it is one tool among several (alongside RCU and lock-free data structures) used only where the data can genuinely be split per CPU.
Conclusion
Per-CPU variables are one of the simplest and most effective lock-free techniques in the Linux kernel: by giving each CPU its own private slot, the entire notion of a shared critical section disappears, and with it the need for any lock. In this lecture you statically declared a per-CPU counter with DEFINE_PER_CPU and inspected each CPU’s slot. The next lecture in this free Linux kernel development course covers the accessor APIs needed to safely read and update the currently running CPU’s own slot.

2 Comments