Per-CPU Variables in the Linux Kernel: Lock-Free Programming-Free Linux Kernel Development Course

Per-CPU Variables in the Linux Kernel: Lock-Free Programming
Free Linux Kernel Development Course — Kernel Synchronization Part 2

← Previous Lecture  |  Next Lecture →

Every lock you add to the Linux kernel is, in effect, a checkpoint where CPUs must queue up. On a system with a handful of cores that cost is small; on a modern server with a few hundred cores it can dominate everything. Per-CPU variables linux kernel developers rely on are one of the kernel’s core lock-free techniques for sidestepping this bottleneck entirely, by giving every CPU its own private copy of the data instead of making them fight over a shared one.

What You Will Learn

  • Why locking does not scale well on high core-count systems
  • What lock contention and lock proliferation mean in practice
  • What a per-CPU variable is and how it eliminates the critical section
  • How to statically declare a per-CPU variable with DEFINE_PER_CPU
  • The difference between DEFINE_PER_CPU and DECLARE_PER_CPU

Prerequisites

  • Understanding of critical sections and why shared data needs protection (earlier lectures)
  • Basic familiarity with spinlocks and mutexes
  • Comfortable building and loading a simple kernel module

Why Locking Alone Does Not Scale

A lock is effectively a single-lane checkpoint: whichever CPU wants to touch the protected data has to wait its turn, one at a time, no matter how many cores the machine has. On a two- or four-core embedded board this rarely matters. But scale that same idea up to a server with a few hundred cores, and every thread wanting the same lock funnels through that one narrow point — the more cores you add, the worse the queue gets. This is usually described as “locking does not scale.”

The obvious fix — splitting one big lock into many smaller, finer-grained locks — helps reduce contention between any two threads, but it does not scale forever either. A kernel already manages many thousands of locks internally, and the more locks that exist, the higher the chance that two code paths acquire them in a different order and hit a subtle deadlock.

Locking Bottleneck vs Per-CPU Approach
LOCKING (shared data, many CPUs) CPU0 –\ CPU1 —> [ single lock ] –> [ shared variable ] (CPUs wait in line) CPU2 –/ PER-CPU (private data, many CPUs) CPU0 –> [ CPU0’s own copy ] (no waiting, no lock) CPU1 –> [ CPU1’s own copy ] (no waiting, no lock) CPU2 –> [ CPU2’s own copy ] (no waiting, no lock)

What Exactly Is a Per-CPU Variable?

A per-CPU (PCP) variable is not one variable at all — under the hood it is an array with one private element per live CPU on the system. On a 4-core system, a per-CPU int is really 4 separate integers, one that belongs exclusively to CPU0, one to CPU1, and so on. A thread running on CPU2 only ever touches CPU2’s private slot.

Because no two CPUs ever write to the same slot, the entire critical section disappears — there is nothing left to protect, and therefore no lock is needed. This is conceptually the same idea as a local variable on a thread’s private stack, or POSIX Thread-Local Storage (TLS) in user space: private-per-execution-context data never needs a lock.

Per-CPU Variable as a Private Array (4-Core System)
DEFINE_PER_CPU(int, pkt_count); CPU0 slot: pkt_count = 0 CPU1 slot: pkt_count = 0 CPU2 slot: pkt_count = 0 CPU3 slot: pkt_count = 0 Thread on CPU1 increments -> only CPU1’s slot changes, no other CPU affected

Statically Declaring a Per-CPU Variable

The Linux kernel offers two macros for compile-time (static) per-CPU allocation, both from <linux/percpu.h>:

MacroPurpose
DEFINE_PER_CPU(type, name)Allocates and zero-initializes the per-CPU variable — use this in the .c file that owns it
DECLARE_PER_CPU(type, name)Only declares an extern reference to a per-CPU variable defined elsewhere — use this in a header shared across files

Original Demo: Declaring a Per-CPU Packet Counter

Here is a small, original kernel module (not copied from any book) that statically defines a per-CPU counter and reports how many live CPUs the system currently has — the exact number of private slots that counter owns.

#include <linux/module.h>
#include <linux/percpu.h>
#include <linux/cpumask.h>

/* One private "pkt_count" slot per live CPU - no lock required */
DEFINE_PER_CPU(int, pkt_count);

static int __init ep_percpu_demo_init(void)
{
    int cpu;

    pr_info("ep_percpu_demo: live CPUs = %d\n", num_online_cpus());

    for_each_online_cpu(cpu) {
        pr_info("ep_percpu_demo: initial pkt_count on CPU%d = %d\n",
                cpu, per_cpu(pkt_count, cpu));
    }

    return 0;
}

static void __exit ep_percpu_demo_exit(void)
{
    pr_info("ep_percpu_demo: unloaded\n");
}

module_init(ep_percpu_demo_init);
module_exit(ep_percpu_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala static per-CPU variable demo");

Build and load it on a 4-core kernel 6.x system:

$ sudo insmod ep_percpu_demo.ko
$ dmesg | tail -5
ep_percpu_demo: live CPUs = 4
ep_percpu_demo: initial pkt_count on CPU0 = 0
ep_percpu_demo: initial pkt_count on CPU1 = 0
ep_percpu_demo: initial pkt_count on CPU2 = 0
ep_percpu_demo: initial pkt_count on CPU3 = 0

Note: This lecture only covers static allocation and inspection with per_cpu() for a fixed, known CPU number. The kernel-safe way to read and modify the currently running CPU’s own slot — using helpers like get_cpu_var()/put_cpu_var() and this_cpu_* operations — is covered in the next lecture of this free Linux kernel development course.

Per-CPU Variables vs Locking vs TLS

TechniqueNeeds a lock?Memory costBest for
Shared variable + spinlock/mutexYesLowData genuinely shared/read by all CPUs
Per-CPU variableNoOne copy per live CPUIndependent counters/statistics per CPU
User-space TLS (__thread)NoOne copy per threadPer-thread private data in user space

Common Mistakes

  • Using per-CPU variables for large data structures. Since the item is duplicated once per CPU, large structs waste significant memory on systems with hundreds of cores.
  • Reaching into a per-CPU variable directly by pointer instead of using the kernel-provided accessor macros — always use the official helper API, never raw pointer arithmetic.
  • Forgetting that per-CPU does not mean per-thread. If a thread migrates to a different CPU mid-access, it can end up reading a different slot than intended unless preemption is properly disabled — the accessor macros exist to prevent exactly this bug.

Best Practices

  • Reach for per-CPU variables for small, frequently-updated statistics (packet counts, hit/miss counters, latency samples).
  • Use DECLARE_PER_CPU in headers and DEFINE_PER_CPU in exactly one .c file, just like any other extern/definition pair.
  • Always use kernel-provided per-CPU accessor APIs rather than touching the underlying array yourself.
Key Takeaways
Locking does not scale to many cores Per-CPU = one private copy per core No critical section, no lock needed DEFINE_PER_CPU allocates + zeroes DECLARE_PER_CPU is the extern form

FAQ

Q1. Are per-CPU variables truly lock-free?
Yes — since each CPU only ever touches its own private slot, there is no shared critical section left to protect.

Q2. How many copies does a per-CPU variable actually create?
Exactly one copy per live CPU core on the running system, reported by num_online_cpus().

Q3. What is the difference between DEFINE_PER_CPU and DECLARE_PER_CPU?
DEFINE_PER_CPU allocates and zero-initializes storage; DECLARE_PER_CPU only declares an extern reference for use in other files.

Q4. Should I use per-CPU variables for large structs?
No — since the data is duplicated per CPU, large structs should stay as shared, lock-protected data instead.

Q5. Is per-CPU similar to user-space Thread-Local Storage?
Conceptually yes — both give each execution context its own private copy so no locking is required.

Q6. Can I read another CPU’s per-CPU slot?
Yes, using per_cpu(variable, cpu_number) for a specific CPU, though this is mainly used for reporting/debugging rather than the hot path.

Q7. Does per-CPU replace spinlocks and mutexes entirely?
No — it is one tool among several (alongside RCU and lock-free data structures) used only where the data can genuinely be split per CPU.

Conclusion

Per-CPU variables are one of the simplest and most effective lock-free techniques in the Linux kernel: by giving each CPU its own private slot, the entire notion of a shared critical section disappears, and with it the need for any lock. In this lecture you statically declared a per-CPU counter with DEFINE_PER_CPU and inspected each CPU’s slot. The next lecture in this free Linux kernel development course covers the accessor APIs needed to safely read and update the currently running CPU’s own slot.

← Previous Lecture  |  Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *