Benchmarking Atomic Bitops vs Spinlock-Protected RMW in the Linux Kernel-Linux Device Driver Training

« Previous Lecture | Next Lecture »

Benchmarking Atomic Bitops vs Spinlock-Protected RMW in the Linux Kernel
Free Linux Kernel Programming Course • Kernel Synchronization Part 2 • Kernel 6.x
Level: Intermediate
Kernel: 6.x
Reading Time: 12 min

Free linux kernel development course: in this lecture we measure, on a real kernel module, how much faster a single atomic bitwise operation is compared to a hand-rolled read-modify-write critical section guarded by a spinlock. This is a hands-on continuation of our free linux device drivers course and free embedded systems course series on kernel synchronization.

What you will learn:

  • How to time kernel code accurately with ktime_get_real_ns()
  • Why a single RMW atomic bitwise instruction beats a spinlock-protected critical section
  • How to build your own benchmarking kernel module
  • How to read and interpret dmesg timing output

Prerequisites: familiarity with spinlocks, atomic bitops (set_bit(), clear_bit()), and basic kernel module loading, covered in earlier lectures of this free linux kernel development course.

free kernel programming course free linux device drivers course atomic bitops benchmark linux kernel spinlock ktime_get_real_ns

Why Benchmark Atomics Against Spinlocks?

Earlier in this chapter we learned that a plain tmp = mem; tmp |= FLAG; mem = tmp; sequence is not safe on its own, because it is really three separate CPU operations that can be interleaved with another core’s access. We fixed that two different ways: wrapping the sequence in a spinlock, and replacing it entirely with a single RMW atomic bitwise call such as set_bit(). Knowing both approaches are correct is not enough for real driver work — we also need to know which one is faster, because synchronization primitives sit directly on hot paths in many drivers.

Timing Kernel Code With ktime_get_real_ns()

The kernel exposes several clock sources for timestamping code paths. ktime_get_real_ns() returns a nanosecond-resolution timestamp and remains fully supported on kernel 6.x. It is not meant for production hot-path profiling — for that you would reach for ftrace or perf events — but it is perfectly adequate for a teaching demo where we just want to compare two code paths against each other on the same machine.

u64 t_start, t_end;

t_start = ktime_get_real_ns();
/* code under test */
t_end = ktime_get_real_ns();

pr_info("ep_bench: delta = %llu ns\n", t_end - t_start);

Our Original Benchmark Driver

Below is an original demo module, ep_bitops_bench, built for this lecture. It exposes a single write-only debugfs file. Writing 0 to it runs the atomic path, writing 1 runs the spinlock-protected path, and the result is printed to the kernel log. Keeping the trigger in debugfs (instead of running both paths automatically at init) lets you repeat each measurement multiple times in a single boot, which gives a more reliable picture than a single one-shot sample.

#include <linux/module.h>
#include <linux/debugfs.h>
#include <linux/spinlock.h>
#include <linux/bitops.h>
#include <linux/ktime.h>

#define EP_READY_BIT 0

static unsigned long ep_flags;
static DEFINE_SPINLOCK(ep_lock);
static struct dentry *ep_dir;

static void ep_run_atomic(void)
{
	u64 t1, t2;

	t1 = ktime_get_real_ns();
	set_bit(EP_READY_BIT, &ep_flags);
	t2 = ktime_get_real_ns();

	pr_info("ep_bench: atomic set_bit() took %llu ns\n", t2 - t1);
}

static void ep_run_spinlock(void)
{
	u64 t1, t2;
	unsigned long tmp;

	t1 = ktime_get_real_ns();
	spin_lock(&ep_lock);
	tmp = ep_flags;
	tmp |= BIT(EP_READY_BIT);
	ep_flags = tmp;
	spin_unlock(&ep_lock);
	t2 = ktime_get_real_ns();

	pr_info("ep_bench: spinlock RMW took %llu ns\n", t2 - t1);
}

static ssize_t ep_trigger_write(struct file *f, const char __user *buf,
				 size_t count, loff_t *ppos)
{
	char c;

	if (copy_from_user(&c, buf, 1))
		return -EFAULT;

	if (c == '0')
		ep_run_atomic();
	else
		ep_run_spinlock();

	return count;
}

static const struct file_operations ep_trigger_fops = {
	.write = ep_trigger_write,
};

static int __init ep_bench_init(void)
{
	ep_dir = debugfs_create_dir("ep_bitops_bench", NULL);
	debugfs_create_file("trigger", 0200, ep_dir, NULL, &ep_trigger_fops);
	pr_info("ep_bench: loaded, write 0 (atomic) or 1 (spinlock) to trigger\n");
	return 0;
}

static void __exit ep_bench_exit(void)
{
	debugfs_remove_recursive(ep_dir);
}

module_init(ep_bench_init);
module_exit(ep_bench_exit);
MODULE_LICENSE("GPL");
Two Paths To The Same Result
[ ep_flags bit 0 needs to be set, safely, on an SMP box ] PATH A: set_bit() PATH B: spinlock + manual RMW —————— ————————— 1 CPU instruction (locked) acquire spinlock (may spin) bus/cache-line lock asserted read ep_flags bit modified atomically OR in the flag done write ep_flags back release spinlock Result: fewer instructions, Result: correct, but pays for no separate lock object, lock acquire + release plus no contention window a wider critical section

What The Numbers Usually Show

Run echo 0 > /sys/kernel/debug/ep_bitops_bench/trigger a few times, then echo 1 > ... a few times, and read dmesg. Exact nanosecond figures depend heavily on your CPU generation, whether the kernel is a debug build, and whether CONFIG_MUTEX_SPIN_ON_OWNER-style spinning is in play, so do not chase a specific number — chase the ratio. On most SMP-capable x86_64 and ARM64 systems you should consistently see the atomic path finish in a small double-digit number of nanoseconds, while the spinlock path costs several times more, simply because it has to acquire a separate lock object in addition to doing the same read-modify-write work.

Aspectset_bit() (atomic)spinlock + manual RMW
CPU instructionsOne locked instructionLock acquire + load + modify + store + lock release
Extra memory usedNoneOne spinlock_t per protected word
Contention behaviourHardware-arbitrated, no spinningLosing CPUs spin until the lock is free
Best used forSingle flag/bit updatesMulti-field critical sections

When You Still Need The Spinlock

This benchmark is not an argument for deleting every spinlock in your driver. Atomic bitops only protect a single word. The moment your critical section touches more than one related field — for example updating a status flag and a counter together so the two never disagree — you are back to needing a lock, because there is no atomic instruction that updates two independent memory locations at once. Use atomic bitops when the operation really is “flip one bit,” and reach for a spinlock or mutex the moment the critical section grows beyond that.

Frequently Asked Questions

Q1. Is ktime_get_real_ns() safe to call from interrupt context?
Yes, it is safe from both process and interrupt context, which is one reason it is convenient for quick driver-level benchmarking.

Q2. Why do my numbers change every run?
CPU frequency scaling, cache state, and scheduler noise all affect a single sample. Average several runs instead of trusting one reading.

Q3. Does a debug kernel affect these numbers?
Yes. Kernels built with lock debugging (CONFIG_DEBUG_SPINLOCK, lockdep) add extra bookkeeping to every lock operation, which will make the spinlock path look even slower than on a production kernel.

Q4. Is set_bit() always faster than a spinlock in every kernel version?
The relative ordering has held for as long as atomic bitops have existed in the kernel, because it comes down to a single locked CPU instruction versus a full lock/unlock sequence, not a version-specific optimization.

Q5. Can I use this technique to benchmark my own driver’s locking?
Yes — wrap any critical section with ktime_get_real_ns() calls the same way, though for production profiling prefer ftrace or perf events.

Q6. What is the debugfs path for this module?
/sys/kernel/debug/ep_bitops_bench/trigger, assuming debugfs is mounted at the default location.

Practice Exercises

  1. Extend the module to also time clear_bit() versus a spinlock-protected clear.
  2. Add a loop that runs each path 1000 times and prints the average delta instead of a single sample.
  3. Repeat the benchmark on a CONFIG_PREEMPT_RT kernel and compare the spinlock numbers to a non-RT kernel.

This lecture is part of EmbeddedPathashala’s free linux kernel development course, free linux device drivers course and free embedded systems course.

More Free Lectures Join the Community

« Previous Lecture | Next Lecture »

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *