Linux Kernel Thread CPU Affinity: Building Lock-Free Drivers-Free Linux Device Drivers Course

Linux Kernel Thread CPU Affinity: Lock-Free Per-CPU Programming Guide (Kernel 6.x)
Linux Kernel Thread CPU Affinity: Building Lock-Free Drivers
Part of EmbeddedPathashala’s Free Linux Kernel Development Course — updated for kernel 6.x
Chapter: Kernel Synchronization Part 2
Level: Intermediate
Kernel: 6.x

← Previous Lecture

Next Lecture →

If you have been following this free linux kernel development course, you already know that shared data usually needs a lock. But there is a neat exception: when you understand linux kernel thread cpu affinity, you can pin a kernel thread to one specific CPU core and let it touch per-CPU data with absolutely no locking at all. In this lecture of our free embedded systems course we build an original demo driver that does exactly this, and we explain how the kernel itself leans on the very same trick internally.

linux kernel thread cpu affinity free linux kernel development course free linux device drivers course free embedded systems course kthread_bind

What You Will Learn

  • Why an ordinary kernel module cannot simply call the scheduler’s affinity function
  • How kthread_bind() pins a kernel thread to a chosen CPU the correct, supported way
  • Building an original two-thread, two-core demo that updates per-CPU data with zero locks
  • The internal design pattern the kernel uses for its own current task pointer
  • How to search the kernel source tree for real usage of any internal API
  • Performance, security, and debugging considerations for CPU-pinned code

Prerequisites

  • Comfort writing a basic loadable kernel module (module_init/module_exit)
  • Completion of the earlier lectures on per-CPU variables, get_cpu_var(), and get_cpu_ptr() in this free linux kernel development course
  • A multi-core test VM (2 or more virtual CPUs) running kernel 6.x
  • Basic familiarity with kernel threads (kthread_create(), kthread_run(), kthread_stop())

Why You Can’t Just Call sched_setaffinity() From a Module

The scheduler function that changes a task’s allowed CPU mask lives deep inside the process scheduler, and it is deliberately not exported to loadable modules. That is not an oversight — the kernel maintainers keep scheduler internals private on purpose so that out-of-tree modules cannot destabilise core scheduling behaviour. Some older tutorials work around this by resolving the symbol address at runtime through kallsyms_lookup_name() and calling it as a raw function pointer. On a modern hardened kernel 6.x build this workaround is fragile: kallsyms_lookup_name() itself was made inaccessible to modules starting with kernel 5.7 for exactly this kind of misuse, and CONFIG_KALLSYMS-based symbol hacking is treated as a security smell by most distribution kernels.

The good news: there has always been a clean, exported, fully supported way to pin a kernel thread (not a random user process) to a CPU — kthread_bind(). Because a kernel thread is the kernel’s own creation, the kthread subsystem is allowed to control exactly where it runs.

The Modern, Correct Way: kthread_bind()

Instead of chasing an unexported scheduler symbol, linux kernel thread cpu affinity for kthreads is set with one exported call:

struct task_struct *tsk = kthread_create(worker_fn, arg, "ep_worker/%d", cpu_id);
if (!IS_ERR(tsk)) {
    kthread_bind(tsk, cpu_id);   /* must be called before wake_up_process() */
    wake_up_process(tsk);
}

Two rules matter here:

  • kthread_bind() must be called before the thread is woken up with wake_up_process(); once running, a bound kthread cannot be re-bound with this call.
  • A kthread bound this way is excluded from the load balancer — the scheduler will never migrate it to another core, which is precisely the property that makes lock-free per-CPU access safe.

Original Demo: Two Kthreads, Two Cores, Zero Locks

Let’s build a fresh driver, ep_percpu_affinity_demo, that starts two kernel threads, binds thread A to CPU 0 and thread B to CPU 1, and has each thread hammer away at a per-CPU counter and a small per-CPU statistics structure. Because each thread is nailed to its own core and never leaves it, both threads can update the same named per-CPU variable concurrently without any spinlock or mutex protecting it.

CPU-Pinned Kthreads — No Shared Lock Required
CPU 0
kthread A (bound)
per_cpu(stats,0).hits++
owns this slice only
CPU 1
kthread B (bound)
per_cpu(stats,1).hits++
owns this slice only
No thread ever touches the other core’s slice → no race, no lock needed
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/percpu.h>
#include <linux/delay.h>

struct ep_stats {
    long hits;
    long last_val;
};

static DEFINE_PER_CPU(struct ep_stats, ep_stats_pcpu);
static struct task_struct *ep_worker[2];

static int ep_worker_fn(void *arg)
{
    long id = (long)arg;
    struct ep_stats *st;
    int i;

    for (i = 0; i < 5 && !kthread_should_stop(); i++) {
        st = this_cpu_ptr(&ep_stats_pcpu);
        st->hits++;
        st->last_val = id * 10 + i;
        pr_info("ep_worker/%ld: cpu=%d hits=%ld last_val=%ld\n",
                id, smp_processor_id(), st->hits, st->last_val);
        msleep(50);
    }
    return 0;
}

static int __init ep_percpu_affinity_init(void)
{
    long i;

    for (i = 0; i < 2; i++) {
        ep_worker[i] = kthread_create(ep_worker_fn, (void *)i,
                                       "ep_worker/%ld", i);
        if (IS_ERR(ep_worker[i]))
            return PTR_ERR(ep_worker[i]);

        kthread_bind(ep_worker[i], i);   /* pin BEFORE waking it up */
        wake_up_process(ep_worker[i]);
    }
    pr_info("ep_percpu_affinity_demo: two cpu-pinned kthreads started\n");
    return 0;
}

static void __exit ep_percpu_affinity_exit(void)
{
    int cpu;
    struct ep_stats *st;

    for_each_online_cpu(cpu) {
        st = per_cpu_ptr(&ep_stats_pcpu, cpu);
        pr_info("ep_percpu_affinity_demo: cpu%d hits=%ld last_val=%ld\n",
                cpu, st->hits, st->last_val);
    }
    pr_info("ep_percpu_affinity_demo: unloaded\n");
}

module_init(ep_percpu_affinity_init);
module_exit(ep_percpu_affinity_exit);
MODULE_LICENSE("GPL");

Notice we used this_cpu_ptr() instead of the older get_cpu_var()/put_cpu_var() pair we introduced in the previous lecture. Because each worker is permanently bound to its own core and never preempted onto another one for this data access (a bound kthread simply cannot migrate), a plain this_cpu_ptr() is enough here — there is no need to additionally disable preemption the way you must when a thread is not pinned.

You can confirm each thread’s placement from a shell while the module is loaded:

$ ps -eLo pid,psr,comm | grep ep_worker
  1832    0 ep_worker/0
  1833    1 ep_worker/1

$ dmesg | tail -12
ep_percpu_affinity_demo: two cpu-pinned kthreads started
ep_worker/0: cpu=0 hits=1 last_val=0
ep_worker/1: cpu=1 hits=1 last_val=10
ep_worker/0: cpu=0 hits=2 last_val=1
ep_worker/1: cpu=1 hits=2 last_val=11

The psr column from ps confirms each kthread genuinely stays glued to the CPU we bound it to for its entire life.

How the Kernel Itself Uses This Exact Idea: the current Pointer

This lecture’s design pattern is not just a classroom toy — the kernel relies on it for one of the hottest lookups in the entire system: finding the currently running task via the current macro. Conceptually (without reproducing any kernel source), the idea works like this: the kernel keeps a per-CPU pointer variable that always holds the address of the task presently executing on that core. Every time the scheduler switches to a new task, it simply overwrites that CPU’s private pointer to point at the new task. Reading current then becomes nothing more than “read my own CPU’s private pointer” — no lock, no atomic RMW, no cross-core traffic, because no other core ever touches that particular slice.

You can build a miniature version of the same idiom yourself. Imagine a driver that tracks “the last debugfs client to touch this CPU”:

static DEFINE_PER_CPU(struct task_struct *, ep_last_client);

/* called from an open() or ioctl() handler */
static void ep_record_client(void)
{
    this_cpu_write(ep_last_client, current);
}

static struct task_struct *ep_who_was_here(void)
{
    return this_cpu_read(ep_last_client);
}

That is the whole trick: a per-CPU pointer, written on context-relevant events and read back with no locking, because “which core am I on” already scopes the access. This is exactly why per-CPU design is one of the highest-leverage lock-free techniques in the entire kernel.

Discovering Real Per-CPU API Usage Yourself

A great habit while working through this free linux device drivers course is to check how widely an API is actually used in the mainline tree before you rely on it. Two simple approaches:

  • grep the source tree: grep -rn "alloc_percpu(" drivers/ kernel/ | wc -l gives you an instant usage count.
  • Browse an online cross-referencer: sites such as the Bootlin Elixir cross-reference let you click any function name and see every caller across kernel versions without needing a local cscope index at all.

Either approach quickly shows that per-CPU allocation functions are used across block I/O statistics, GIC interrupt controller code, network routing tables, and slab allocator internals — confirming this is a mainstream, production-proven pattern and not an obscure trick.

CPU Affinity vs Locking: A Comparison

ApproachNeeds a Lock?Scales With Cores?Complexity
Shared global variable + spinlockYesPoor (lock contention grows)Low
Per-CPU variable, thread NOT pinnedYes (disable preemption)GoodMedium
Per-CPU variable + kthread_bind() pinned threadNoExcellentMedium

Real-World Use Cases

  • Per-core network softirq statistics counters that never need cross-core synchronisation
  • Benchmark or latency-measurement kthreads that must run in isolation on a dedicated core
  • Real-time or automotive embedded systems where a specific core is reserved for one deterministic worker (common in embedded systems course case studies)
  • Interrupt-affinity-aware drivers that pair an IRQ pinned to core N with a kthread also pinned to core N to avoid any cache-line bouncing

Common Mistakes and Troubleshooting

  • Calling kthread_bind() after wake_up_process() — the bind is ignored once the thread is already running; always bind first.
  • Assuming this_cpu_ptr() is always lock-free — it is only safe without extra protection when you are certain the calling thread cannot migrate off that core, as with a bound kthread.
  • Binding to a CPU number that may not exist — always validate against num_online_cpus() or iterate with for_each_online_cpu() rather than hardcoding CPU 0/1.
  • Forgetting kthread_stop() cleanup — bound kthreads still need a clean shutdown path in your module’s exit function, exactly like unbound ones.

Best Practices

  • Prefer kthread_bind() over any scheduler-symbol hack — it is exported, stable, and maintained.
  • Document in code comments exactly why a thread is pinned, so future maintainers don’t “fix” it by removing the bind.
  • Always pair per-CPU writes with a matching per-CPU read pattern (per_cpu_ptr/for_each_online_cpu) for diagnostics.
  • Test on a VM with at least as many virtual CPUs as your driver expects to bind to.

Performance Considerations

Pinning removes migration overhead and keeps each core’s cache lines warm with only that core’s data, which is the entire performance win of this pattern. The tradeoff is scheduling flexibility: a pinned kthread cannot be moved even if its core becomes busy with other work, so only pin threads that genuinely benefit from strict locality.

Security Considerations

Never resolve unexported scheduler symbols via kallsyms_lookup_name() in production code; besides being unsupported, this function’s own visibility to modules has been progressively locked down in recent kernels as a hardening measure. Sticking to exported, documented APIs like kthread_bind() keeps your driver compatible with hardened and locked-down kernel configurations.

Summary / Key Takeaways

  • Kernel modules cannot call the scheduler’s internal affinity function directly — and should not work around this with symbol-resolution hacks.
  • kthread_bind() is the correct, exported way to achieve linux kernel thread cpu affinity for kthreads you create yourself.
  • A kthread that is permanently bound to one core can access that core’s per-CPU data with zero locking.
  • The kernel’s own current pointer uses this exact per-CPU idiom internally.

Conclusion

Understanding linux kernel thread cpu affinity closes the loop on everything this free linux kernel development course has covered about per-CPU variables: allocation, access, and now placement. Once a kthread is reliably bound to a single core with kthread_bind(), per-CPU data stops being a synchronisation problem and becomes simple, fast, private storage. This is one of the cleanest lock-free techniques available anywhere in kernel programming, and it is the same idea the kernel trusts for one of its hottest code paths. In the next lecture of this free linux device drivers course we move on to sequence locks (seqlock_t) as the next lock-free building block.

Frequently Asked Questions

Q1. Can I bind a regular user-space process with kthread_bind()?
No. kthread_bind() only works on kernel threads created by kthread_create(); user-space CPU affinity is set with sched_setaffinity() from user space or taskset at the shell.

Q2. What happens if I call kthread_bind() on an already-running thread?
The kernel will refuse to move it correctly; the bind must always happen before the first wake_up_process() call.

Q3. Is this_cpu_ptr() always safe without a lock?
Only when you can guarantee the calling context cannot migrate cores mid-access — a bound kthread satisfies this; an ordinary process context generally does not.

Q4. Why not just use kallsyms_lookup_name() to call sched_setaffinity() directly?
It relies on an unexported, unstable internal symbol, and recent kernels have restricted kallsyms_lookup_name() itself from modules for security reasons. kthread_bind() achieves the same practical result the supported way.

Q5. Does binding a kthread hurt overall system scheduling?
Only pin threads that truly need strict locality; each pinned thread removes one core’s worth of scheduling flexibility from the load balancer.

Q6. How does this relate to the current macro?
current is implemented as a per-CPU pointer that the scheduler updates on every context switch, so each core always reads its own private pointer — the same lock-free idiom this lecture teaches with an original example.

Q7. Can this technique work on a single-CPU (UP) system?
Technically per-CPU variables still exist on UP kernels, but the whole benefit of this lecture comes from having genuinely separate cores; on UP there is only one slice to worry about anyway.

Q8. What is the fastest way to find how many places in the kernel call a given per-CPU API?
Use grep -rn against the kernel source tree, or browse an online kernel source cross-referencer instead of building a local index.

← Previous Lecture

Next Lecture →

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *