RMW Atomic Bitwise Operators in the Linux Kernel-Linux Device Drivers Course

← Previous Lecture  |  Next Lecture →

RMW Atomic Bitwise Operators in the Linux Kernel
set_bit(), clear_bit(), change_bit() and the test_and_* family — updated for kernel 6.x
Free Linux Kernel Course
Kernel 6.x Ready
Hands-On Driver Example

If you are learning free linux kernel development, one of the most practical topics you will run into is read modify write atomic operations linux kernel style bit manipulation. Any time a driver needs to flip a single flag bit — a “device busy” flag, a “data ready” flag, an error flag — inside a shared status word, a plain read-modify-write on that bit is not safe on multicore systems. The Linux kernel solves this with a dedicated family of RMW bitwise atomic operators: set_bit(), clear_bit(), change_bit(), and their “test_and_” counterparts. This lecture walks through each of them, explains exactly why they are safe, and builds an original demo driver that manages a set of status flags with them on a modern kernel.

What You Will Learn

  • Why a single bit inside a shared word still needs atomic protection
  • The six core RMW bitwise atomic APIs and what each one returns
  • The difference between the “set/clear/change” group and the “test_and_” group
  • Why these operators are safe across CPU cores, not just within one core
  • When you still need a spinlock even though you are using atomic bitops
  • How to build a status-flag driver using these APIs on kernel 6.x

Prerequisites

  • Comfortable with basic kernel module structure (init/exit, module_init/module_exit)
  • Understanding of critical sections and why plain i++ is not atomic
  • Familiarity with atomic_t / refcount_t is helpful but not required

Why Even a Single Bit Needs Protection

Turning a bit on with plain C code usually looks like flags |= (1 << 7);. That single line of source code is not a single CPU instruction. Under the hood, the processor has to read the current value of flags into a register, set the bit inside the register, and write the register back to memory. If two CPU cores execute this sequence on the same word at nearly the same moment, one core’s write can silently overwrite the other core’s update, and one of the two flag changes is lost. This is exactly the kind of read-modify-write race we covered earlier with plain integers — bit flags have the same problem, just at bit granularity instead of whole-word granularity.

Lost Update on a Shared Flags Word
CPU 0 CPU 1 —— —— read flags = 0x00 read flags = 0x00 set bit 7 in register set bit 3 in register write flags = 0x80 write flags = 0x08 <– bit 7 update LOST Final value in memory: 0x08 (should have been 0x88)

The RMW Bitwise Atomic API Family

The kernel groups these operators into two families. The first family — set_bit(), clear_bit(), change_bit() — performs the update and returns nothing. The second family — the test_and_* versions — performs the same update but atomically hands back the bit’s previous value, which is extremely useful when you need to know whether a flag was already set before you set it.

API Effect Returns
void set_bit(unsigned int nr, volatile unsigned long *p)Atomically sets bit nr of p to 1nothing
void clear_bit(unsigned int nr, volatile unsigned long *p)Atomically clears bit nr of p to 0nothing
void change_bit(unsigned int nr, volatile unsigned long *p)Atomically toggles bit nr of pnothing
int test_and_set_bit(unsigned int nr, volatile unsigned long *p)Atomically sets bit nrprevious value of the bit
int test_and_clear_bit(unsigned int nr, volatile unsigned long *p)Atomically clears bit nrprevious value of the bit
int test_and_change_bit(unsigned int nr, volatile unsigned long *p)Atomically toggles bit nrprevious value of the bit

There is also a plain, non-atomic test_bit() used only to read a bit’s current value — it does not modify anything, so it does not need RMW protection by itself.

Important: these operators are not just safe with respect to the single CPU core running them — they are safe across every core in the system. That means if several CPUs are calling set_bit() / clear_bit() on the same word at the same time, no update is lost. But atomicity of one bit operation does not automatically make a longer sequence of operations atomic — if your driver logic depends on checking multiple bits and then acting on the combined result, that whole sequence is still a critical section and may need a spinlock around it.

Hands-On: A Status-Flags Driver Using RMW Bitops

Below is an original example driver, ep_bitflags_demo, written for this course. It models a small device status word with three flags — READY, BUSY, and ERROR — and shows how a real driver would use set_bit(), test_and_set_bit(), and clear_bit() from different code paths (open, a simulated worker, and release) without any manual read-modify-write.

#include <linux/module.h>
#include <linux/fs.h>
#include <linux/miscdevice.h>
#include <linux/bitops.h>
#include <linux/printk.h>

#define EP_FLAG_READY   0
#define EP_FLAG_BUSY    1
#define EP_FLAG_ERROR   2

static unsigned long ep_status_flags;

static int ep_bitflags_open(struct inode *inode, struct file *filp)
{
    /* Try to atomically claim BUSY; bail out if another opener beat us to it */
    if (test_and_set_bit(EP_FLAG_BUSY, &ep_status_flags)) {
        pr_info("ep_bitflags_demo: device already busy\n");
        return -EBUSY;
    }

    set_bit(EP_FLAG_READY, &ep_status_flags);
    clear_bit(EP_FLAG_ERROR, &ep_status_flags);

    pr_info("ep_bitflags_demo: opened, flags = 0x%lx\n", ep_status_flags);
    return 0;
}

static int ep_bitflags_release(struct inode *inode, struct file *filp)
{
    clear_bit(EP_FLAG_BUSY, &ep_status_flags);
    pr_info("ep_bitflags_demo: released, flags = 0x%lx\n", ep_status_flags);
    return 0;
}

static const struct file_operations ep_bitflags_fops = {
    .owner   = THIS_MODULE,
    .open    = ep_bitflags_open,
    .release = ep_bitflags_release,
};

static struct miscdevice ep_bitflags_miscdev = {
    .minor = MISC_DYNAMIC_MINOR,
    .name  = "ep_bitflags_demo",
    .fops  = &ep_bitflags_fops,
};

static int __init ep_bitflags_init(void)
{
    int ret = misc_register(&ep_bitflags_miscdev);

    if (ret) {
        pr_err("ep_bitflags_demo: misc_register failed: %d\n", ret);
        return ret;
    }

    pr_info("ep_bitflags_demo: loaded\n");
    return 0;
}

static void __exit ep_bitflags_exit(void)
{
    misc_deregister(&ep_bitflags_miscdev);
    pr_info("ep_bitflags_demo: unloaded\n");
}

module_init(ep_bitflags_init);
module_exit(ep_bitflags_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala RMW atomic bitops demo (kernel 6.x)");

Notice that test_and_set_bit() gives us a single atomic “check BUSY and claim it” operation. Without it, we would need to read the flag, check it, and set it in three separate steps — exactly the kind of race window we saw in the diagram above.

test_and_set_bit() as an Atomic Claim
Process A calls open() Process B calls open() ————————– ————————– test_and_set_bit(BUSY) -> reads old value: 0 -> sets bit to 1 -> returns 0 (was free) => proceeds with open test_and_set_bit(BUSY) -> reads old value: 1 -> sets bit to 1 (no-op) -> returns 1 (was busy) => open() returns -EBUSY The entire read+set happens as one indivisible step, so there is no window where both processes could see BUSY as “free”.

Naive Bit Manipulation vs Atomic Bitops vs Spinlock

Approach Safe on SMP? When to use
flags |= BIT(n);NoNever for shared state — only for a purely local, single-threaded variable
set_bit() / test_and_set_bit()Yes, for a single bit operationIndependent flags that are set/cleared/tested one at a time
spinlock + plain bit opsYes, for the whole critical sectionWhen you must check and act on several bits together as one logical step

Frequently Asked Questions

Q1. Is set_bit() faster than using a spinlock to protect a single flag?
Yes. The atomic bitops compile down to a single locked CPU instruction on most architectures, so there is no sleep/spin overhead the way a contended lock can have.

Q2. Can I use these operators on a stack variable?
You can, but there is no benefit unless that variable is actually shared between contexts (other CPUs, interrupt handlers, or kernel threads). On purely local data, plain bit operations are fine and faster.

Q3. What is the difference between change_bit() and test_and_change_bit()?
Both toggle the bit atomically. test_and_change_bit() additionally returns what the bit’s value was immediately before the toggle.

Q4. Do I need to include a special header for these functions?
Yes, <linux/bitops.h> provides these declarations on modern kernels.

Q5. Are these functions safe to call from interrupt context?
The bit operations themselves are atomic and interrupt-safe. But as with any shared state, if the same word is also accessed inside a longer non-atomic critical section elsewhere, that surrounding section still needs its own protection (for example a spinlock using the _irqsave variant).

Q6. Can these operators work on device registers, not just RAM?
Yes. Device registers mapped into kernel virtual address space through MMIO behave like ordinary memory locations, so the same set_bit()/clear_bit() family can be applied to them.

Q7. What happens if two CPUs call test_and_set_bit() on the same bit at the exact same instant?
The hardware serializes the two operations. One of them will see the bit as already set and get a return value of 1; the other will see it as free and get 0. There is no scenario where both see 0.

Practice Exercise

Extend the ep_bitflags_demo driver above with a fourth flag, EP_FLAG_SUSPENDED. Use test_and_change_bit() in a new ioctl handler to toggle suspend state and log whether the device was suspended or resumed, based purely on the return value of the call.

free linux kernel development course free linux device drivers course free embedded systems course read modify write atomic operations linux kernel

Continue the Kernel Synchronization Part 2 series on EmbeddedPathashala.

Next Lecture

← Previous Lecture  |  Next Lecture →

3 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *