Per-CPU Variables in the Linux Kernel: A Practical Guide for Kernel 6.12+-Linux Kernel Development Course Online

← Previous Lecture  |  Next Lecture →

Per-CPU Variables in the Linux Kernel: A Practical Guide for Kernel 6.12+

Modern multi-core processors can execute dozens of kernel threads at the exact same instant, and every one of those threads may want to touch a shared piece of kernel data. The traditional fix — wrap the data in a spinlock — works, but locks are expensive on hot code paths that run millions of times a second, such as network packet counters or scheduler statistics.

The Linux kernel solves this with per-CPU variables: instead of one shared copy of a variable, the kernel keeps one private copy for every CPU core on the system. Each core reads and writes only its own copy, so there is no contention and no need for a lock at all. This free Linux kernel development course lecture walks you through per-CPU variables from first principles, using the current per-CPU APIs found in kernel 6.12 and later.

What You Will Learn
Why per-CPU variables avoid locking overhead Static vs dynamic per-CPU allocation get_cpu_var / put_cpu_var and preemption Modern this_cpu_* operations Per-CPU data under PREEMPT_RT A working per-CPU counter example

Prerequisites

Before working through this lecture, you should be comfortable with:

  • Writing and building a basic loadable kernel module (LKM)
  • Basic C pointers and structures
  • The idea of kernel preemption and why blocking calls (like memory allocation that can sleep) are dangerous inside atomic sections
  • A Linux system running kernel 6.1 or newer for testing (kernel 6.12 LTS is used as the reference version throughout this article)

Why Per-CPU Data Exists

Picture a network driver that increments a “packets received” counter on every incoming packet. On an 8-core system, eight CPUs could be handling interrupts from the same network card simultaneously. If all eight cores update one shared counter, you need a lock or an atomic instruction on every single packet — and that lock becomes a bottleneck precisely when traffic is heaviest.

Per-CPU variables sidestep the problem entirely. Each core gets its own private counter, so there is nothing to contend over. When you need the combined total, you simply add up all the per-core copies — an operation that happens rarely compared to how often the counter itself is updated.

Conceptual Layout of a Per-CPU Counter on a 4-Core System
CPU 0
rx_count = 812
CPU 1
rx_count = 745
CPU 2
rx_count = 903
CPU 3
rx_count = 660

Each core owns an independent copy of the same variable — no shared memory, no lock needed.

This design brings three concrete benefits:

Benefit Why It Matters
No locking overhead Each core only ever touches its own copy, so there is nothing to serialize
No cache line bouncing A shared variable forces the owning cache line to migrate between cores; per-CPU data stays local to each core’s cache
Predictable, low latency Ideal for interrupt handlers and other latency-sensitive paths

Static Allocation: DEFINE_PER_CPU and DECLARE_PER_CPU

The simplest way to create a per-CPU variable is at compile time, using macros from <linux/percpu.h>. DEFINE_PER_CPU() both declares and zero-initializes the variable, while DECLARE_PER_CPU() is used in header files when you only need a forward declaration.

#include <linux/percpu.h>

/* One private copy of 'rx_count' per online CPU */
DEFINE_PER_CPU(unsigned long, rx_count);

At module load time, the kernel silently allocates one instance of rx_count for every possible CPU on the system (not just the ones currently online), so hot-plugging a new core later works without any extra code from you.

Dynamic Allocation for Per-CPU Structures

Static per-CPU variables are fine for a single counter, but a real driver usually needs a whole structure per core — for example, a receive-queue statistics block that is only needed while the driver is loaded. For that you allocate dynamically with alloc_percpu():

struct rxq_stats {
    unsigned long packets;
    unsigned long bytes;
    unsigned long drops;
};

struct rxq_stats __percpu *stats;

stats = alloc_percpu(struct rxq_stats);
if (!stats)
    return -ENOMEM;

If you need to pass GFP allocation flags (say, to allow allocation from an atomic context), use alloc_percpu_gfp() instead:

stats = alloc_percpu_gfp(struct rxq_stats, GFP_ATOMIC);

When your driver is bound to a struct device, prefer the resource-managed variant. It ties the per-CPU memory to the device’s lifetime, so the kernel frees it automatically when the device is removed:

stats = devm_alloc_percpu(dev, struct rxq_stats);

Any memory you allocate manually with alloc_percpu() or alloc_percpu_gfp() must be released explicitly with free_percpu() once you are done with it, typically in your module’s cleanup path:

free_percpu(stats);
Allocation Style Typical Use Case Freed By
DEFINE_PER_CPU Simple counters known at compile time Automatically, at module unload
alloc_percpu() / alloc_percpu_gfp() Structures needed only while the module is loaded free_percpu(), manually
devm_alloc_percpu() Driver data tied to a struct device Automatically, when the device is unbound

Reading and Writing Per-CPU Data Safely

Because “your” copy of a per-CPU variable depends on which core the code happens to be running on, the kernel must guarantee that you are not migrated to a different core in the middle of an access. That guarantee comes from a pair of macros that temporarily disable kernel preemption:

unsigned long val;

val = get_cpu_var(rx_count);
val++;
put_cpu_var(rx_count);

get_cpu_var() calls preempt_disable() internally and returns an lvalue for the current CPU’s copy, so it can be incremented, assigned to, or read directly. put_cpu_var() re-enables preemption with preempt_enable(). The two calls must always be balanced — if you call the getter twice, you must call the setter twice as well.

Critical Rule: Never Sleep Between get_cpu_var and put_cpu_var
get_cpu_var()
→
No sleeping / blocking calls here
→
put_cpu_var()

The section between the two calls is effectively an atomic context. Calling anything that can sleep — a blocking memory allocation, a mutex lock, or a function that internally waits on I/O — inside that window is a kernel bug, and a debug kernel will report it with a “sleeping function called from invalid context” warning the moment it happens.

To read a value that belongs to a specific CPU (not necessarily the one you are running on), use per_cpu() instead. It does not disable preemption because it does not care which core is currently executing:

int cpu;
unsigned long total = 0;

for_each_online_cpu(cpu)
    total += per_cpu(rx_count, cpu);

pr_info("total packets received: %lu\n", total);

For dynamically allocated per-CPU structures, the pointer equivalents are get_cpu_ptr(), put_cpu_ptr(), and per_cpu_ptr():

struct rxq_stats *s;

s = get_cpu_ptr(stats);
s->packets++;
put_cpu_ptr(stats);

/* Read another CPU's structure without disabling preemption */
struct rxq_stats *cpu2_stats = per_cpu_ptr(stats, 2);

The Modern Way: this_cpu_* Operations

Kernels since the 3.x days — and very much still current in 6.12 — provide a family of this_cpu_*() helpers that fold the get/modify/put sequence into a single call. On most architectures these compile down to a single instruction that is inherently safe against migration, so they are both shorter to write and often faster than the get_cpu_var pattern:

this_cpu_inc(rx_count);
this_cpu_add(rx_count, 5);
val = this_cpu_read(rx_count);
this_cpu_write(rx_count, 0);

For dynamically allocated structures, the pointer form works the same way once you have the per-CPU pointer for the current core:

this_cpu_inc(stats->packets);
this_cpu_add(stats->bytes, skb->len);

Current kernel documentation and driver code generally recommend reaching for this_cpu_*() first, and only falling back to get_cpu_var()/put_cpu_var() when you genuinely need the raw pointer for a multi-step update.

Per-CPU Data and PREEMPT_RT

Since the real-time preemption patches merged into mainline Linux, kernel preemption is far more aggressive by default. On a PREEMPT_RT-enabled 6.12 kernel, disabling preemption is treated as a much heavier hammer, and long get_cpu_var-style sections can hurt the very real-time latency guarantees the kernel is trying to provide.

For code that needs to combine per-CPU access with a genuine critical section — more than a single read or increment — the current recommendation is to pair per-CPU data with a local_lock_t, which behaves like a normal spinlock on RT kernels but compiles down to simple preemption disabling on a standard kernel. This keeps your driver correct on both kernel configurations without maintaining two code paths.

A Complete Example: A Per-CPU Packet Counter Module

The following minimal module ties everything together: it statically defines a per-CPU counter, increments it safely, and reports the total across all cores when the module is removed.

#include <linux/module.h>
#include <linux/percpu.h>
#include <linux/kernel.h>

static DEFINE_PER_CPU(unsigned long, sample_hits);

static int __init pcpu_demo_init(void)
{
    int cpu;

    /* Simulate a hit being recorded on the current CPU */
    this_cpu_inc(sample_hits);

    for_each_online_cpu(cpu)
        pr_info("pcpu_demo: cpu %d hits = %lu\n",
                cpu, per_cpu(sample_hits, cpu));

    return 0;
}

static void __exit pcpu_demo_exit(void)
{
    int cpu;
    unsigned long total = 0;

    for_each_online_cpu(cpu)
        total += per_cpu(sample_hits, cpu);

    pr_info("pcpu_demo: total hits across all CPUs = %lu\n", total);
}

module_init(pcpu_demo_init);
module_exit(pcpu_demo_exit);
MODULE_LICENSE("GPL");

Build it with a standard out-of-tree Kbuild Makefile, load it with insmod, and inspect the output with dmesg:

$ sudo insmod pcpu_demo.ko
$ dmesg | tail
$ sudo rmmod pcpu_demo
$ dmesg | tail

Real-World Use Cases

  • Network stack statistics — interface RX/TX counters are per-CPU so packet processing never contends on a lock
  • Scheduler run-queues — every CPU maintains its own run-queue as a per-CPU data structure
  • RCU implementation — read-side tracking in RCU relies heavily on per-CPU counters for scalability
  • Memory allocator caches — SLUB/SLAB keep per-CPU free-object caches to avoid a global allocator lock on every allocation
  • Per-CPU workqueues — many kernel subsystems queue deferred work on a per-CPU basis to keep work local to the core that generated it

Common Mistakes and Troubleshooting

Mistake Symptom Fix
Sleeping between get_cpu_var/put_cpu_var “sleeping function called from invalid context” in dmesg Move any blocking allocation or lock outside the critical section
Unbalanced get/put calls Preemption stays disabled far longer than intended, hurting latency Always match every get_cpu_var/get_cpu_ptr with exactly one put
Forgetting free_percpu() Memory leak reported by kmemleak on module unload Free dynamically allocated per-CPU memory in your exit path, or use devm_alloc_percpu()
Treating per_cpu() as safe for concurrent writers Race conditions when two paths write the same CPU’s copy from different contexts Only the owning CPU should freely write; cross-CPU writes still need explicit synchronization

Best Practices

  • Prefer this_cpu_*() operations over manual get_cpu_var/put_cpu_var whenever a single atomic-style update is all you need
  • Keep the code between get_cpu_var() and put_cpu_var() as short as possible
  • Use devm_alloc_percpu() whenever your per-CPU data is tied to a device’s lifetime
  • Pair per-CPU access with local_lock_t if you need a genuine multi-step critical section that must also behave correctly under PREEMPT_RT
  • Sum per-CPU values with for_each_possible_cpu() during initialization and for_each_online_cpu() for runtime aggregation, since a CPU can be present but currently offline

Performance Considerations

Per-CPU variables trade memory for speed: the kernel allocates one copy per possible CPU, so a structure that seems small can add up on systems with a very high core count. In exchange, you eliminate lock contention and cache-line bouncing entirely on the hot path, which is almost always the right trade for counters, statistics, and per-core caches that are updated far more often than they are read in aggregate.

Security Considerations

Per-CPU data is kernel memory like any other, so the usual rules apply: never expose raw per-CPU pointers to user space, validate any size or index values that come from user-controlled input before using them to compute an offset, and remember that aggregating per-CPU counters for a value exposed through procfs or sysfs can leak timing information about per-core activity if you are not careful about what you publish.

Summary / Key Takeaways

  • Per-CPU variables give each core its own private copy of a variable, removing the need for locks on hot paths
  • Static allocation uses DEFINE_PER_CPU/DECLARE_PER_CPU; dynamic allocation uses alloc_percpu(), alloc_percpu_gfp(), or devm_alloc_percpu()
  • get_cpu_var()/put_cpu_var() create a short atomic section that must never sleep
  • Modern kernels favor the shorter, often single-instruction this_cpu_*() operations
  • PREEMPT_RT changes the cost of disabling preemption, so pair per-CPU access with local_lock_t for anything beyond a trivial update

Conclusion

Per-CPU variables are one of the simplest and most effective scalability tools available inside the Linux kernel. Once you understand the difference between static and dynamic allocation, the atomicity rules around get_cpu_var()/put_cpu_var(), and the modern this_cpu_*() helpers, you have everything you need to write driver and subsystem code that scales cleanly across systems with dozens of cores — without ever reaching for a lock on your hottest code paths. As you continue through this free Linux kernel development course, keep an eye out for per-CPU data structures in the scheduler and networking stack; they are one of the clearest real-world examples of this pattern in action.

FAQ

1. What is a per-CPU variable in the Linux kernel?

A per-CPU variable is a kernel data item that has one independent copy for every CPU on the system, so each core can read and modify its own copy without contending with other cores.

2. When should I use DEFINE_PER_CPU instead of alloc_percpu()?

Use DEFINE_PER_CPU() for simple, always-present data such as a single counter that is known at compile time. Use alloc_percpu() when the data is a structure that should only exist while your module or driver is loaded.

3. Why can’t I sleep between get_cpu_var() and put_cpu_var()?

get_cpu_var() disables kernel preemption to guarantee you stay on the same core. Sleeping requires the scheduler to be able to switch tasks, which is not allowed while preemption is disabled — doing so is a kernel bug.

4. What is the difference between this_cpu_inc() and get_cpu_var() plus an increment?

this_cpu_inc() folds the entire operation into a single, migration-safe step, often a single CPU instruction. get_cpu_var() followed by a manual increment and put_cpu_var() achieves the same result but requires you to manage preemption state yourself.

5. Do per-CPU variables need locks?

Not for the core update itself, since each CPU only touches its own copy. You still need synchronization if a task running on one CPU needs to write another CPU’s copy directly.

6. How do I get the combined value across all CPUs?

Loop over for_each_online_cpu() (or for_each_possible_cpu() during setup) and add up the per_cpu() value from each core.

7. What happens to per-CPU data when a CPU goes offline?

The memory for that CPU’s copy remains allocated; it simply is not updated while the CPU is offline. Aggregation loops should account for this by choosing the correct iterator for the situation.

8. Are per-CPU variables affected by PREEMPT_RT?

Yes. On a PREEMPT_RT kernel, disabling preemption for extended periods is discouraged because it hurts real-time latency. Combine per-CPU data with local_lock_t for anything beyond a trivial single-step update.

9. Can I use per-CPU variables in interrupt handlers?

Yes, and this is one of their most common uses. Just be careful that this_cpu_*() operations are still appropriate for the interrupt context you are running in, and avoid any sleeping calls.

10. Do I need to free statically defined per-CPU variables?

No. Variables created with DEFINE_PER_CPU() are cleaned up automatically when the module is unloaded. Only memory obtained through alloc_percpu() or alloc_percpu_gfp() must be released explicitly with free_percpu().

Continue Your Free Linux Kernel Development Course

Explore more lectures on kernel internals, device drivers, and embedded systems programming.

Browse the Course Index Join EmbeddedPathashala

← Previous Lecture  |  Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *