← Previous Lecture | Next Lecture →
If you are learning free linux kernel development, one of the most practical topics you will run into is read modify write atomic operations linux kernel style bit manipulation. Any time a driver needs to flip a single flag bit — a “device busy” flag, a “data ready” flag, an error flag — inside a shared status word, a plain read-modify-write on that bit is not safe on multicore systems. The Linux kernel solves this with a dedicated family of RMW bitwise atomic operators: set_bit(), clear_bit(), change_bit(), and their “test_and_” counterparts. This lecture walks through each of them, explains exactly why they are safe, and builds an original demo driver that manages a set of status flags with them on a modern kernel.
What You Will Learn
- Why a single bit inside a shared word still needs atomic protection
- The six core RMW bitwise atomic APIs and what each one returns
- The difference between the “set/clear/change” group and the “test_and_” group
- Why these operators are safe across CPU cores, not just within one core
- When you still need a spinlock even though you are using atomic bitops
- How to build a status-flag driver using these APIs on kernel 6.x
Prerequisites
- Comfortable with basic kernel module structure (init/exit, module_init/module_exit)
- Understanding of critical sections and why plain
i++is not atomic - Familiarity with
atomic_t/refcount_tis helpful but not required
Why Even a Single Bit Needs Protection
Turning a bit on with plain C code usually looks like flags |= (1 << 7);. That single line of source code is not a single CPU instruction. Under the hood, the processor has to read the current value of flags into a register, set the bit inside the register, and write the register back to memory. If two CPU cores execute this sequence on the same word at nearly the same moment, one core’s write can silently overwrite the other core’s update, and one of the two flag changes is lost. This is exactly the kind of read-modify-write race we covered earlier with plain integers — bit flags have the same problem, just at bit granularity instead of whole-word granularity.
The RMW Bitwise Atomic API Family
The kernel groups these operators into two families. The first family — set_bit(), clear_bit(), change_bit() — performs the update and returns nothing. The second family — the test_and_* versions — performs the same update but atomically hands back the bit’s previous value, which is extremely useful when you need to know whether a flag was already set before you set it.
| API | Effect | Returns |
|---|---|---|
void set_bit(unsigned int nr, volatile unsigned long *p) | Atomically sets bit nr of p to 1 | nothing |
void clear_bit(unsigned int nr, volatile unsigned long *p) | Atomically clears bit nr of p to 0 | nothing |
void change_bit(unsigned int nr, volatile unsigned long *p) | Atomically toggles bit nr of p | nothing |
int test_and_set_bit(unsigned int nr, volatile unsigned long *p) | Atomically sets bit nr | previous value of the bit |
int test_and_clear_bit(unsigned int nr, volatile unsigned long *p) | Atomically clears bit nr | previous value of the bit |
int test_and_change_bit(unsigned int nr, volatile unsigned long *p) | Atomically toggles bit nr | previous value of the bit |
There is also a plain, non-atomic test_bit() used only to read a bit’s current value — it does not modify anything, so it does not need RMW protection by itself.
Important: these operators are not just safe with respect to the single CPU core running them — they are safe across every core in the system. That means if several CPUs are calling set_bit() / clear_bit() on the same word at the same time, no update is lost. But atomicity of one bit operation does not automatically make a longer sequence of operations atomic — if your driver logic depends on checking multiple bits and then acting on the combined result, that whole sequence is still a critical section and may need a spinlock around it.
Hands-On: A Status-Flags Driver Using RMW Bitops
Below is an original example driver, ep_bitflags_demo, written for this course. It models a small device status word with three flags — READY, BUSY, and ERROR — and shows how a real driver would use set_bit(), test_and_set_bit(), and clear_bit() from different code paths (open, a simulated worker, and release) without any manual read-modify-write.
#include <linux/module.h>
#include <linux/fs.h>
#include <linux/miscdevice.h>
#include <linux/bitops.h>
#include <linux/printk.h>
#define EP_FLAG_READY 0
#define EP_FLAG_BUSY 1
#define EP_FLAG_ERROR 2
static unsigned long ep_status_flags;
static int ep_bitflags_open(struct inode *inode, struct file *filp)
{
/* Try to atomically claim BUSY; bail out if another opener beat us to it */
if (test_and_set_bit(EP_FLAG_BUSY, &ep_status_flags)) {
pr_info("ep_bitflags_demo: device already busy\n");
return -EBUSY;
}
set_bit(EP_FLAG_READY, &ep_status_flags);
clear_bit(EP_FLAG_ERROR, &ep_status_flags);
pr_info("ep_bitflags_demo: opened, flags = 0x%lx\n", ep_status_flags);
return 0;
}
static int ep_bitflags_release(struct inode *inode, struct file *filp)
{
clear_bit(EP_FLAG_BUSY, &ep_status_flags);
pr_info("ep_bitflags_demo: released, flags = 0x%lx\n", ep_status_flags);
return 0;
}
static const struct file_operations ep_bitflags_fops = {
.owner = THIS_MODULE,
.open = ep_bitflags_open,
.release = ep_bitflags_release,
};
static struct miscdevice ep_bitflags_miscdev = {
.minor = MISC_DYNAMIC_MINOR,
.name = "ep_bitflags_demo",
.fops = &ep_bitflags_fops,
};
static int __init ep_bitflags_init(void)
{
int ret = misc_register(&ep_bitflags_miscdev);
if (ret) {
pr_err("ep_bitflags_demo: misc_register failed: %d\n", ret);
return ret;
}
pr_info("ep_bitflags_demo: loaded\n");
return 0;
}
static void __exit ep_bitflags_exit(void)
{
misc_deregister(&ep_bitflags_miscdev);
pr_info("ep_bitflags_demo: unloaded\n");
}
module_init(ep_bitflags_init);
module_exit(ep_bitflags_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala RMW atomic bitops demo (kernel 6.x)");
Notice that test_and_set_bit() gives us a single atomic “check BUSY and claim it” operation. Without it, we would need to read the flag, check it, and set it in three separate steps — exactly the kind of race window we saw in the diagram above.
Naive Bit Manipulation vs Atomic Bitops vs Spinlock
| Approach | Safe on SMP? | When to use |
|---|---|---|
flags |= BIT(n); | No | Never for shared state — only for a purely local, single-threaded variable |
set_bit() / test_and_set_bit() | Yes, for a single bit operation | Independent flags that are set/cleared/tested one at a time |
| spinlock + plain bit ops | Yes, for the whole critical section | When you must check and act on several bits together as one logical step |
Frequently Asked Questions
Q1. Is set_bit() faster than using a spinlock to protect a single flag?
Yes. The atomic bitops compile down to a single locked CPU instruction on most architectures, so there is no sleep/spin overhead the way a contended lock can have.
Q2. Can I use these operators on a stack variable?
You can, but there is no benefit unless that variable is actually shared between contexts (other CPUs, interrupt handlers, or kernel threads). On purely local data, plain bit operations are fine and faster.
Q3. What is the difference between change_bit() and test_and_change_bit()?
Both toggle the bit atomically. test_and_change_bit() additionally returns what the bit’s value was immediately before the toggle.
Q4. Do I need to include a special header for these functions?
Yes, <linux/bitops.h> provides these declarations on modern kernels.
Q5. Are these functions safe to call from interrupt context?
The bit operations themselves are atomic and interrupt-safe. But as with any shared state, if the same word is also accessed inside a longer non-atomic critical section elsewhere, that surrounding section still needs its own protection (for example a spinlock using the _irqsave variant).
Q6. Can these operators work on device registers, not just RAM?
Yes. Device registers mapped into kernel virtual address space through MMIO behave like ordinary memory locations, so the same set_bit()/clear_bit() family can be applied to them.
Q7. What happens if two CPUs call test_and_set_bit() on the same bit at the exact same instant?
The hardware serializes the two operations. One of them will see the bit as already set and get a return value of 1; the other will see it as free and get 0. There is no scenario where both see 0.
Practice Exercise
Extend the ep_bitflags_demo driver above with a fourth flag, EP_FLAG_SUSPENDED. Use test_and_change_bit() in a new ioctl handler to toggle suspend state and log whether the device was suspended or resumed, based purely on the return value of the call.
Continue the Kernel Synchronization Part 2 series on EmbeddedPathashala.
Next Lecture
3 Comments