Per-CPU Variable Access in the Linux Kernel: get_cpu_var, per_cpu() and Pointer APIs-Linux Device Drivers Course

← Previous Lecture  |  Next Lecture →

Per-CPU Variable Access in the Linux Kernel: get_cpu_var, per_cpu() and Pointer APIs

A hands-on, kernel 6.x lecture from the free Linux Kernel Development Course on EmbeddedPathashala

In the previous lecture of this free Linux kernel development course, we introduced per-CPU variables as a lock-free way to avoid the scalability problems of shared-data locking. This lecture completes that picture by focusing entirely on per cpu variable access in the Linux kernel: how a driver actually reads, updates, and iterates over per-CPU data safely, without ever taking a lock. If you are building your embedded Linux or device driver skills through a free embedded systems course, understanding correct per-CPU variable access is one of the most practical performance skills you can pick up, because it shows up constantly in real driver and subsystem code.

Free Embedded Systems Course Free Linux Development Course Free Linux Device Drivers Course Free Linux Kernel Development Course

What You Will Learn

get_cpu_var() / put_cpu_var() per_cpu() macro get_cpu_ptr() / put_cpu_ptr() per_cpu_ptr() this_cpu_* operations alloc_percpu() / free_percpu() Original per-CPU debugfs demo module

Prerequisites

You should already be comfortable with basic Linux kernel module structure (module_init/module_exit), and it helps to have read the earlier lecture on per-CPU variable fundamentals and static allocation with DEFINE_PER_CPU(). Familiarity with spinlocks from the kernel synchronization chapter is useful for comparison but not mandatory.

Why Correct Per-CPU Variable Access Matters

A per-CPU variable is really an array with one private slot for every CPU core, but the kernel does not let you index that array casually. Correct per cpu variable access in the Linux kernel means guaranteeing that, for the duration of your read-modify-write, the kernel thread cannot be moved to a different core and cannot be preempted onto a different context mid-operation. Get this wrong and you silently update the wrong core’s copy, which is a very hard bug to reproduce because it depends on scheduling timing.

The get_cpu_var() and put_cpu_var() Pair

The classic way to safely touch a statically defined per-CPU variable is the get_cpu_var() / put_cpu_var() pair. Internally, get_cpu_var() disables kernel preemption on the current core and returns an lvalue for your variable’s slot on that core; put_cpu_var() re-enables preemption once you are done. Everything between the two calls is, in effect, a short atomic-context critical section.

DEFINE_PER_CPU(int, ep_hit_counter);

static void ep_bump_hit_counter(void)
{
    get_cpu_var(ep_hit_counter)++;
    put_cpu_var(ep_hit_counter);
}

Why the Critical Section Must Never Sleep

Because preemption is disabled between get_cpu_var() and put_cpu_var(), any function that can block — memory allocation with GFP_KERNEL, taking a mutex, copy_to_user(), or similar — is illegal there. Calling a blocking function inside this window triggers a “sleeping function called from invalid context” kernel bug report, usually caught immediately by CONFIG_DEBUG_ATOMIC_SLEEP on a debug kernel. Logging with pr_info() is fine, since it does not block.

Allowed inside get_cpu_var()/put_cpu_var()Not allowed inside get_cpu_var()/put_cpu_var()
Simple arithmetic on the per-CPU valuekmalloc(…, GFP_KERNEL) or vmalloc()
pr_info() / pr_debug() loggingmutex_lock() or any sleeping lock
Reading another already-disabled-preemption valuecopy_to_user() / copy_from_user()

Reading Any Core’s Value with per_cpu()

Sometimes you don’t want the current core’s value — you want to inspect every core’s copy, for example to print a summary. That is what the per_cpu(variable, cpu) macro is for. Unlike get_cpu_var(), it does not disable preemption, so you must only use it when you already know the target CPU is stable (for example, while iterating with for_each_online_cpu() for a read-only report).

static void ep_dump_hit_counters(void)
{
    unsigned int cpu;

    for_each_online_cpu(cpu) {
        int val = per_cpu(ep_hit_counter, cpu);

        pr_info("ep_percpu: cpu %u hit_counter = %d\n", cpu, val);
    }
}

The Modern Shortcut: this_cpu_* Operations

Kernel 6.x still supports get_cpu_var()/put_cpu_var(), but for simple read-modify-write operations the this_cpu_* family (this_cpu_inc(), this_cpu_add(), this_cpu_read(), this_cpu_write()) is now the preferred style of per cpu variable access in the Linux kernel. These helpers fold the preempt-disable, operate, preempt-enable sequence into a single call, and on most architectures they compile down to one instruction using segment-relative addressing, which is faster than the explicit get/put pair.

DEFINE_PER_CPU(int, ep_hit_counter);

static void ep_bump_hit_counter_modern(void)
{
    this_cpu_inc(ep_hit_counter);
}
StyleUse When
get_cpu_var() / put_cpu_var()You need the lvalue for a multi-step or more complex critical section
this_cpu_inc() / this_cpu_add() / this_cpu_read()A single simple operation — the recommended modern default
per_cpu(var, cpu)Read-only inspection of a specific or every core’s value

Dynamically Allocated Per-CPU Data and Pointer Macros

Statically defined per-CPU integers are fine for counters, but real drivers usually need a per-CPU structure allocated at module load time. This is done with alloc_percpu(type), or alloc_percpu_gfp(type, gfp) if you need to control the allocation flags. If you are inside a proper driver’s probe function with a struct device *dev available, the resource-managed devm_alloc_percpu(dev, type) is preferable, since the kernel then frees the per-CPU block automatically on driver detach. Memory from either of the plain routines must be freed manually with free_percpu().

To safely reach the current core’s copy of a dynamically allocated per-CPU pointer, the kernel gives you the pointer equivalents of get_cpu_var()/put_cpu_var(): get_cpu_ptr() and put_cpu_ptr(). To inspect a specific core’s copy without disabling preemption (again, only when you already know the CPU is stable), use per_cpu_ptr(pointer, cpu).

Static per-CPU variableDynamic per-CPU pointer
DEFINE_PER_CPU(type, name)alloc_percpu(type) / devm_alloc_percpu(dev, type)
get_cpu_var(name) / put_cpu_var(name)get_cpu_ptr(ptr) / put_cpu_ptr(ptr)
per_cpu(name, cpu)per_cpu_ptr(ptr, cpu)
Freed automatically at module unloadfree_percpu(ptr) required (unless devm_ managed)
Per-CPU Access Paths at a Glance
CPU 0 private slot    CPU 1 private slot    CPU 2 private slot
    |                         |                         |
    get_cpu_var() / this_cpu_inc() → touches ONLY the running core’s slot, preemption disabled
    per_cpu(var, N) / per_cpu_ptr(ptr, N) → reads a chosen core’s slot, no preemption change

Hands-On: An Original Per-CPU Statistics Debugfs Module

Let’s put this together in a small, original demo module you can build on any kernel 6.x virtual machine. It keeps a per-CPU events/drops structure updated with get_cpu_ptr()/put_cpu_ptr(), and exposes a debugfs file that sums every core’s copy using per_cpu_ptr() and for_each_online_cpu().

#include <linux/module.h>
#include <linux/percpu.h>
#include <linux/debugfs.h>
#include <linux/seq_file.h>
#include <linux/cpumask.h>

struct ep_pcpu_stats {
    u64 events;
    u64 drops;
};

static struct ep_pcpu_stats __percpu *ep_stats;
static struct dentry *ep_dbg_dir;

static void ep_record_event(bool dropped)
{
    struct ep_pcpu_stats *st = get_cpu_ptr(ep_stats);

    st->events++;
    if (dropped)
        st->drops++;

    put_cpu_ptr(ep_stats);
}

static int ep_stats_show(struct seq_file *m, void *v)
{
    unsigned int cpu;
    u64 total_events = 0, total_drops = 0;

    for_each_online_cpu(cpu) {
        struct ep_pcpu_stats *st = per_cpu_ptr(ep_stats, cpu);

        seq_printf(m, "cpu%u: events=%llu drops=%llu\n",
                   cpu, st->events, st->drops);
        total_events += st->events;
        total_drops  += st->drops;
    }

    seq_printf(m, "total: events=%llu drops=%llu\n",
               total_events, total_drops);
    return 0;
}
DEFINE_SHOW_ATTRIBUTE(ep_stats);

static int __init ep_percpu_init(void)
{
    unsigned int cpu;

    ep_stats = alloc_percpu(struct ep_pcpu_stats);
    if (!ep_stats)
        return -ENOMEM;

    for_each_possible_cpu(cpu)
        ep_record_event(false); /* seed a sample event, demo only */

    ep_dbg_dir = debugfs_create_dir("ep_percpu", NULL);
    debugfs_create_file("stats", 0444, ep_dbg_dir, NULL, &ep_stats_fops);

    pr_info("ep_percpu: loaded, see /sys/kernel/debug/ep_percpu/stats\n");
    return 0;
}

static void __exit ep_percpu_exit(void)
{
    debugfs_remove_recursive(ep_dbg_dir);
    free_percpu(ep_stats);
    pr_info("ep_percpu: unloaded\n");
}

module_init(ep_percpu_init);
module_exit(ep_percpu_exit);
MODULE_LICENSE("GPL");

Build and try it out:

make -C /lib/modules/$(uname -r)/build M=$(pwd) modules
sudo insmod ep_percpu.ko
sudo mount -t debugfs none /sys/kernel/debug   # if not already mounted
cat /sys/kernel/debug/ep_percpu/stats
sudo rmmod ep_percpu
dmesg | tail

Notice the module never takes a single lock. Every core writes only its own slot through get_cpu_ptr()/put_cpu_ptr(), and the debugfs reader safely walks every slot with per_cpu_ptr() inside a for_each_online_cpu() loop.

Real-World Use Cases

  • Network drivers keep per-CPU RX/TX packet and drop counters to avoid a global lock on every packet.
  • Block layer and filesystem code use per-CPU counters for statistics that would otherwise bottleneck on high-core-count servers.
  • Tracing and profiling infrastructure (such as the kernel’s own per-CPU ring buffers) rely on this same access pattern for lock-free event recording.

Common Mistakes and Troubleshooting

MistakeConsequenceFix
Calling a blocking function inside get_cpu_var()/put_cpu_var()“sleeping function called from invalid context” bugMove allocation/locking outside the critical section
Using per_cpu(var, cpu) on the current CPU while assuming it’s preemption-safeTask may migrate mid-read on non-debug kernels, giving a stale valueUse get_cpu_var()/this_cpu_* when you need the running core’s own value
Forgetting free_percpu() on a plain alloc_percpu() allocationPer-CPU memory leak on module unloadAlways pair alloc_percpu() with free_percpu(), or use devm_alloc_percpu()
Mismatched get_cpu_var()/put_cpu_var() callsPreemption stays disabled longer than intended, hurting responsivenessAlways call put_cpu_var() exactly as many times as get_cpu_var()

Best Practices

  • Prefer this_cpu_* helpers for simple counters; reserve get_cpu_var()/put_cpu_var() for cases where you genuinely need the lvalue across several statements.
  • Never sleep, allocate with GFP_KERNEL, or take a mutex inside a get_cpu_var()/get_cpu_ptr() critical section.
  • Use devm_alloc_percpu() whenever a struct device pointer is available, so cleanup is automatic.
  • Keep the code between get_cpu_var()/put_cpu_var() as short as possible to minimize the time preemption is disabled.

Performance Considerations

The entire point of correct per cpu variable access in the Linux kernel is performance: no cache-line bouncing between cores and no lock contention. But the preempt_disable() inside get_cpu_var()/this_cpu_* still has a real, if small, cost, and it briefly increases scheduling latency on that core. On latency-sensitive systems keep these critical sections extremely short — a handful of instructions, not loops or function calls.

Security Considerations

Per-CPU data is still kernel memory, so the usual rules apply: never expose a raw per_cpu_ptr() value to user space, and always validate any user-supplied CPU index before passing it to per_cpu() or per_cpu_ptr() — an out-of-range or offline CPU number can lead to an invalid memory access.

Summary / Key Takeaways

  • get_cpu_var()/put_cpu_var() safely access the running core’s static per-CPU slot by disabling preemption around it.
  • this_cpu_* operations are the modern, lower-overhead equivalent for simple read-modify-write operations.
  • per_cpu(var, cpu) and per_cpu_ptr(ptr, cpu) read a specific core’s slot without touching preemption state.
  • get_cpu_ptr()/put_cpu_ptr() are the pointer versions of get_cpu_var()/put_cpu_var(), used with alloc_percpu()/devm_alloc_percpu() allocations.
  • Nothing that can sleep is allowed inside any of these critical sections.

Conclusion

Correct per cpu variable access in the Linux kernel is what turns per-CPU variables from a neat idea into a genuinely lock-free, scalable technique. Once you’re comfortable choosing between get_cpu_var(), this_cpu_*, and the per_cpu()/per_cpu_ptr() read-any-core macros, you have everything you need to add fast, contention-free counters and small state structures to your own drivers. This closes out the per-CPU section of the kernel synchronization chapter in this free Linux kernel development course — next we move on to the remaining synchronization primitives.

Frequently Asked Questions

1. What is the difference between get_cpu_var() and this_cpu_inc()?

get_cpu_var() returns an lvalue you can use across several statements before calling put_cpu_var(); this_cpu_inc() and its relatives perform one simple operation in a single call and are generally preferred when that’s all you need.

2. Can I sleep inside a get_cpu_var()/put_cpu_var() block?

No. Preemption is disabled in that window, so any blocking call is a kernel bug, reported as “sleeping function called from invalid context” on a debug-enabled kernel.

3. When should I use per_cpu() instead of get_cpu_var()?

Use per_cpu(var, cpu) when you want to read a specific core’s value (often while iterating all cores for a summary) and you don’t need preemption disabled, because you’re not depending on staying on that core.

4. What’s the difference between get_cpu_ptr() and per_cpu_ptr()?

get_cpu_ptr() returns the running core’s pointer and disables preemption until put_cpu_ptr(); per_cpu_ptr() returns a specific core’s pointer without touching preemption state, so it’s meant for read-only inspection, not for safe read-modify-write on the current core.

5. Do I still need free_percpu() if I used devm_alloc_percpu()?

No. devm_alloc_percpu() ties the allocation to the driver’s device lifetime, so the kernel frees it automatically on detach. Plain alloc_percpu() always needs a matching free_percpu().

6. Are this_cpu_* operations available on all architectures?

Yes, this_cpu_* is a core kernel API implemented on every supported architecture, though the exact generated instructions vary by CPU architecture for efficiency.

7. Why did my per_cpu_ptr() loop crash with an invalid CPU number?

per_cpu_ptr() does not validate the CPU index for you. Always bound your CPU number using for_each_online_cpu() or for_each_possible_cpu(), or check it against nr_cpu_ids first.

This lecture is part of the free Linux kernel development course and free Linux device drivers course on EmbeddedPathashala.

← Previous Lecture  |  Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *