read_lock() and write_lock(): The Complete Reader-Writer Spinlock API in Linux Kernel 6.x-Best Linux Device Driver Training Online

« Previous Lecture | Next Lecture »

read_lock() and write_lock(): The Complete Reader-Writer Spinlock API in Linux Kernel 6.x

Free Linux Kernel Programming Course — Kernel Synchronization Part 2

Kernel 6.x Updated
Original Driver Example
Beginner Friendly

In the previous lecture of this free Linux kernel development course you learned what a reader writer spinlock is and why it exists: to let many CPU cores read shared kernel data at the same time while still giving a writer exclusive access when it needs to change that data. In this lecture we go hands-on with the actual read_lock() / write_lock() API, look at the IRQ-safe and bottom-half-safe variants, and study the single biggest problem with reader-writer spinlocks in real production kernels: writer starvation.

read_lock write_lock rwlock_t api writer starvation free linux kernel development course free linux device drivers course

What You Will Learn

  • The full read_lock() / write_lock() API surface, including IRQ and bottom-half safe variants
  • How to decide, in your own driver, which variant to call
  • Why reader-writer spinlocks can starve a waiting writer, with a step-by-step timeline
  • A complete original kernel module that protects a shared status table with rwlock_t
  • Why modern kernel code is steadily moving away from rwlock_t toward seqlock_t and RCU

Prerequisites

This lecture assumes you have completed the earlier lectures on spinlocks, mutexes, and the introduction to reader-writer spinlocks in this free Linux device drivers course. You should be comfortable building and loading a basic kernel module on a Linux kernel 6.x system.

The Reader-Writer Spinlock API Family

A rwlock_t is declared and initialized the same way as a normal spinlock:

rwlock_t ep_status_lock;
rwlock_init(&ep_status_lock);

/* or, for statically declared locks */
DEFINE_RWLOCK(ep_status_lock);

Once initialized, the kernel gives you four base primitives:

FunctionPurpose
read_lock(&lock)Take the lock for reading; multiple readers can hold it together
read_unlock(&lock)Release a reader’s hold on the lock
write_lock(&lock)Take the lock for writing; blocks until no readers or writers remain
write_unlock(&lock)Release the writer’s exclusive hold

Just like the plain spinlock family, every one of these has interrupt-safe and bottom-half-safe cousins:

ContextReader variantWriter variant
Process context only, no IRQ contentionread_lock() / read_unlock()write_lock() / write_unlock()
Shared with a hardirq handlerread_lock_irqsave() / read_unlock_irqrestore()write_lock_irqsave() / write_unlock_irqrestore()
IRQs already known to be enabled/disabledread_lock_irq() / read_unlock_irq()write_lock_irq() / write_unlock_irq()
Shared with a softirq/taskletread_lock_bh() / read_unlock_bh()write_lock_bh() / write_unlock_bh()

The rule for choosing between these is exactly the same rule you already learned for plain spinlocks earlier in this course: if the data can also be touched from a hardware interrupt handler, you must use the _irqsave (or at minimum _irq) form, or you risk the same self-deadlock scenario we studied with spin_lock(). One subtlety worth remembering here: even the read IRQ-safe lock disables kernel preemption on the CPU that holds it, so a reader critical section should still be kept short.

Reader-Writer Spinlock: Concurrent Readers, Exclusive Writer
CPU0: read_lock() —-[reading]—- read_unlock() CPU1: read_lock() —-[reading]—- read_unlock() CPU2: write_lock() –[BLOCKED until CPU0 & CPU1 unlock]– —-[writing]—- write_unlock()

An Original rwlock_t Demo Driver

Below is an original example (not taken from any book or existing source) that protects a small in-memory status table with a reader-writer spinlock. A background kernel thread acts as the “writer,” periodically updating the table, while reads happen through a debugfs file.

#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/debugfs.h>
#include <linux/spinlock.h>
#include <linux/delay.h>

#define EP_NR_SLOTS 4

struct ep_status_table {
	rwlock_t lock;
	u32 slot[EP_NR_SLOTS];
	u32 update_count;
};

static struct ep_status_table ep_tbl;
static struct task_struct *ep_writer_thr;
static struct dentry *ep_dbgdir;

static int ep_writer_fn(void *data)
{
	int i;

	while (!kthread_should_stop()) {
		write_lock(&ep_tbl.lock);
		for (i = 0; i < EP_NR_SLOTS; i++)
			ep_tbl.slot[i] = ep_tbl.update_count + i;
		ep_tbl.update_count++;
		write_unlock(&ep_tbl.lock);

		msleep(500);
	}
	return 0;
}

static ssize_t ep_status_read(struct file *f, char __user *ubuf,
			       size_t count, loff_t *ppos)
{
	char kbuf[128];
	int len, i;

	read_lock(&ep_tbl.lock);
	len = scnprintf(kbuf, sizeof(kbuf), "update #%u : ",
			 ep_tbl.update_count);
	for (i = 0; i < EP_NR_SLOTS; i++)
		len += scnprintf(kbuf + len, sizeof(kbuf) - len,
				  "%u ", ep_tbl.slot[i]);
	read_unlock(&ep_tbl.lock);

	return simple_read_from_buffer(ubuf, count, ppos, kbuf, len);
}

static const struct file_operations ep_status_fops = {
	.owner = THIS_MODULE,
	.read  = ep_status_read,
};

static int __init ep_rwlock_demo_init(void)
{
	rwlock_init(&ep_tbl.lock);

	ep_dbgdir = debugfs_create_dir("ep_rwlock_demo", NULL);
	debugfs_create_file("status", 0444, ep_dbgdir, NULL, &ep_status_fops);

	ep_writer_thr = kthread_run(ep_writer_fn, NULL, "ep_rwlock_writer");
	if (IS_ERR(ep_writer_thr))
		return PTR_ERR(ep_writer_thr);

	pr_info("ep_rwlock_demo: loaded\n");
	return 0;
}

static void __exit ep_rwlock_demo_exit(void)
{
	kthread_stop(ep_writer_thr);
	debugfs_remove_recursive(ep_dbgdir);
	pr_info("ep_rwlock_demo: unloaded\n");
}

module_init(ep_rwlock_demo_init);
module_exit(ep_rwlock_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original rwlock_t demo");

Any number of user-space processes can cat /sys/kernel/debug/ep_rwlock_demo/status at once and all of them will be allowed to read concurrently, while the background writer thread is guaranteed exclusive access whenever it updates the table.

The Writer Starvation Problem

Reader-writer spinlocks look like a free performance win, but they come with a well-known downside: a waiting writer can starve if readers keep arriving faster than the current set of readers finish.

How a Writer Starves
t0: Reader A holds the lock t1: Reader B arrives, also allowed in (readers stack freely) t2: Writer W arrives, must wait for A and B to finish t3: Reader C arrives — also allowed in, because readers keep being admitted t4: Reader D arrives — same story …: as long as new readers keep showing up, W never gets a turn

This is not a hypothetical concern — it is explicitly called out in the kernel’s own locking documentation, and it is one of the main reasons the kernel community has been steadily moving read-mostly code paths toward lock-free alternatives such as RCU (Read-Copy-Update) instead of rwlock_t. As we noted in the previous lecture, for genuinely read-mostly data a seqlock_t or RCU-based design is usually the better modern choice; plain rwlock_t is best reserved for cases where the read-side critical section is short and reasonably infrequent.

Best Practices

  • Keep both the reader and the writer critical sections as short as possible.
  • Never sleep, allocate memory with GFP_KERNEL, or call blocking APIs while holding a rwlock_t of either kind — it is a spinlock family, not a sleeping lock.
  • If your workload is truly read-heavy and long-lived, evaluate seqlock_t or RCU before reaching for rwlock_t.
  • Always match the lock/unlock IRQ variant — mixing read_lock() with read_unlock_irqrestore() is a bug.

Common Mistakes

MistakeWhy it’s a problemFix
Using plain read_lock() when a hardirq also touches the dataSelf-deadlock if the interrupt fires on the same CPU while holding the lockUse read_lock_irqsave() / write_lock_irqsave()
Long read-side critical sectionsIncreases the chance of writer starvation and cache ping-pongCopy data out under the lock, process it after unlocking
Sleeping inside a read or write critical sectionrwlock_t is a spinning lock — sleeping triggers a scheduling-while-atomic bugUse a reader-writer semaphore (covered in the next lecture) instead

Summary

The read_lock() / write_lock() family gives you concurrent-reader, exclusive-writer protection with the same low-level spinning behaviour as an ordinary spinlock, including IRQ-safe and bottom-half-safe variants. Its biggest weakness is writer starvation under sustained reader load, which is why modern kernel code increasingly favours seqlock_t or RCU for read-mostly data structures.

Frequently Asked Questions

Is rwlock_t still used in the modern Linux kernel?

Yes, but its use is narrowing over time. It still appears in places like filesystem extent trees and networking code, but new read-mostly designs are increasingly written with RCU or seqlock_t instead.

Can a reader-writer spinlock be held across a sleep?

No. Like any spinlock, both the reader and writer sides spin the CPU rather than sleep, so you must never call a blocking function while holding one.

What happens if two writers call write_lock() at the same time?

Only one writer proceeds; the other spins until the first calls write_unlock(), exactly like a normal spinlock.

Why does the IRQ-safe read variant disable preemption too?

Because disabling interrupts on a CPU already implies the scheduler cannot preempt that CPU, so kernel preemption is disabled as a side effect for the duration of the critical section.

When should I prefer a reader-writer semaphore instead?

Whenever your critical section might need to sleep — for example while copying data to or from user space. That is exactly what the next lecture in this free Linux kernel development course covers.

Conclusion

You now know the complete rwlock_t API surface used throughout the Linux kernel, how to pick the correct IRQ/bottom-half variant for your driver, and why writer starvation is the key trade-off to keep in mind. In the next lecture of this free Linux device drivers course, we move to the sleepable sibling of this lock: the reader-writer semaphore.

« Previous Lecture | Next Lecture »

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *