« Previous Lecture | Next Lecture »
Free linux kernel development course: in this lecture we measure, on a real kernel module, how much faster a single atomic bitwise operation is compared to a hand-rolled read-modify-write critical section guarded by a spinlock. This is a hands-on continuation of our free linux device drivers course and free embedded systems course series on kernel synchronization.
What you will learn:
- How to time kernel code accurately with
ktime_get_real_ns() - Why a single RMW atomic bitwise instruction beats a spinlock-protected critical section
- How to build your own benchmarking kernel module
- How to read and interpret
dmesgtiming output
Prerequisites: familiarity with spinlocks, atomic bitops (set_bit(), clear_bit()), and basic kernel module loading, covered in earlier lectures of this free linux kernel development course.
Why Benchmark Atomics Against Spinlocks?
Earlier in this chapter we learned that a plain tmp = mem; tmp |= FLAG; mem = tmp; sequence is not safe on its own, because it is really three separate CPU operations that can be interleaved with another core’s access. We fixed that two different ways: wrapping the sequence in a spinlock, and replacing it entirely with a single RMW atomic bitwise call such as set_bit(). Knowing both approaches are correct is not enough for real driver work — we also need to know which one is faster, because synchronization primitives sit directly on hot paths in many drivers.
Timing Kernel Code With ktime_get_real_ns()
The kernel exposes several clock sources for timestamping code paths. ktime_get_real_ns() returns a nanosecond-resolution timestamp and remains fully supported on kernel 6.x. It is not meant for production hot-path profiling — for that you would reach for ftrace or perf events — but it is perfectly adequate for a teaching demo where we just want to compare two code paths against each other on the same machine.
u64 t_start, t_end;
t_start = ktime_get_real_ns();
/* code under test */
t_end = ktime_get_real_ns();
pr_info("ep_bench: delta = %llu ns\n", t_end - t_start);
Our Original Benchmark Driver
Below is an original demo module, ep_bitops_bench, built for this lecture. It exposes a single write-only debugfs file. Writing 0 to it runs the atomic path, writing 1 runs the spinlock-protected path, and the result is printed to the kernel log. Keeping the trigger in debugfs (instead of running both paths automatically at init) lets you repeat each measurement multiple times in a single boot, which gives a more reliable picture than a single one-shot sample.
#include <linux/module.h>
#include <linux/debugfs.h>
#include <linux/spinlock.h>
#include <linux/bitops.h>
#include <linux/ktime.h>
#define EP_READY_BIT 0
static unsigned long ep_flags;
static DEFINE_SPINLOCK(ep_lock);
static struct dentry *ep_dir;
static void ep_run_atomic(void)
{
u64 t1, t2;
t1 = ktime_get_real_ns();
set_bit(EP_READY_BIT, &ep_flags);
t2 = ktime_get_real_ns();
pr_info("ep_bench: atomic set_bit() took %llu ns\n", t2 - t1);
}
static void ep_run_spinlock(void)
{
u64 t1, t2;
unsigned long tmp;
t1 = ktime_get_real_ns();
spin_lock(&ep_lock);
tmp = ep_flags;
tmp |= BIT(EP_READY_BIT);
ep_flags = tmp;
spin_unlock(&ep_lock);
t2 = ktime_get_real_ns();
pr_info("ep_bench: spinlock RMW took %llu ns\n", t2 - t1);
}
static ssize_t ep_trigger_write(struct file *f, const char __user *buf,
size_t count, loff_t *ppos)
{
char c;
if (copy_from_user(&c, buf, 1))
return -EFAULT;
if (c == '0')
ep_run_atomic();
else
ep_run_spinlock();
return count;
}
static const struct file_operations ep_trigger_fops = {
.write = ep_trigger_write,
};
static int __init ep_bench_init(void)
{
ep_dir = debugfs_create_dir("ep_bitops_bench", NULL);
debugfs_create_file("trigger", 0200, ep_dir, NULL, &ep_trigger_fops);
pr_info("ep_bench: loaded, write 0 (atomic) or 1 (spinlock) to trigger\n");
return 0;
}
static void __exit ep_bench_exit(void)
{
debugfs_remove_recursive(ep_dir);
}
module_init(ep_bench_init);
module_exit(ep_bench_exit);
MODULE_LICENSE("GPL");
What The Numbers Usually Show
Run echo 0 > /sys/kernel/debug/ep_bitops_bench/trigger a few times, then echo 1 > ... a few times, and read dmesg. Exact nanosecond figures depend heavily on your CPU generation, whether the kernel is a debug build, and whether CONFIG_MUTEX_SPIN_ON_OWNER-style spinning is in play, so do not chase a specific number — chase the ratio. On most SMP-capable x86_64 and ARM64 systems you should consistently see the atomic path finish in a small double-digit number of nanoseconds, while the spinlock path costs several times more, simply because it has to acquire a separate lock object in addition to doing the same read-modify-write work.
| Aspect | set_bit() (atomic) | spinlock + manual RMW |
|---|---|---|
| CPU instructions | One locked instruction | Lock acquire + load + modify + store + lock release |
| Extra memory used | None | One spinlock_t per protected word |
| Contention behaviour | Hardware-arbitrated, no spinning | Losing CPUs spin until the lock is free |
| Best used for | Single flag/bit updates | Multi-field critical sections |
When You Still Need The Spinlock
This benchmark is not an argument for deleting every spinlock in your driver. Atomic bitops only protect a single word. The moment your critical section touches more than one related field — for example updating a status flag and a counter together so the two never disagree — you are back to needing a lock, because there is no atomic instruction that updates two independent memory locations at once. Use atomic bitops when the operation really is “flip one bit,” and reach for a spinlock or mutex the moment the critical section grows beyond that.
Frequently Asked Questions
Q1. Is ktime_get_real_ns() safe to call from interrupt context?
Yes, it is safe from both process and interrupt context, which is one reason it is convenient for quick driver-level benchmarking.
Q2. Why do my numbers change every run?
CPU frequency scaling, cache state, and scheduler noise all affect a single sample. Average several runs instead of trusting one reading.
Q3. Does a debug kernel affect these numbers?
Yes. Kernels built with lock debugging (CONFIG_DEBUG_SPINLOCK, lockdep) add extra bookkeeping to every lock operation, which will make the spinlock path look even slower than on a production kernel.
Q4. Is set_bit() always faster than a spinlock in every kernel version?
The relative ordering has held for as long as atomic bitops have existed in the kernel, because it comes down to a single locked CPU instruction versus a full lock/unlock sequence, not a version-specific optimization.
Q5. Can I use this technique to benchmark my own driver’s locking?
Yes — wrap any critical section with ktime_get_real_ns() calls the same way, though for production profiling prefer ftrace or perf events.
Q6. What is the debugfs path for this module?/sys/kernel/debug/ep_bitops_bench/trigger, assuming debugfs is mounted at the default location.
Practice Exercises
- Extend the module to also time
clear_bit()versus a spinlock-protected clear. - Add a loop that runs each path 1000 times and prints the average delta instead of a single sample.
- Repeat the benchmark on a
CONFIG_PREEMPT_RTkernel and compare the spinlock numbers to a non-RT kernel.
This lecture is part of EmbeddedPathashala’s free linux kernel development course, free linux device drivers course and free embedded systems course.
More Free Lectures Join the Community
2 Comments