« Previous Lecture | Next Lecture »
Reader-Writer Semaphore (rw_semaphore) in Linux Kernel Drivers: Complete Guide
Free Linux Kernel Programming Course — Kernel Synchronization Part 2
A reader writer semaphore, or rw_semaphore, solves the one problem the reader-writer spinlock cannot: what happens when a reader or a writer needs to sleep while holding the lock — for example while copying data to or from user space. This lecture in our free Linux kernel development course covers the full rw_semaphore API, an original demo driver, and a quick look at two related kernel synchronization primitives you will meet later: completions and sequence locks.
What You Will Learn
- The full
rw_semaphoreAPI:down_read(),up_read(),down_write(),up_write(), and theirtrylock/killablevariants - When to choose a reader-writer semaphore over a reader-writer spinlock
- A complete original kernel module protecting a driver’s configuration data with
rw_semaphore - A short introduction to two related primitives: the completion mechanism and the sequence lock
Prerequisites
This lecture builds directly on the previous lecture’s coverage of read_lock() / write_lock(). You should already be comfortable with mutexes and reader-writer spinlocks from earlier in this free Linux kernel programming course.
Reader-Writer Semaphore vs Reader-Writer Spinlock
Conceptually, a reader-writer semaphore behaves exactly like a reader-writer spinlock: many readers can hold it together, a writer needs it exclusively. The difference is entirely in what happens when the lock is unavailable.
| Property | rwlock_t (spinlock) | rw_semaphore |
|---|---|---|
| Waiting behaviour | Spins the CPU | Sleeps, CPU is released to the scheduler |
| Safe to sleep inside critical section | No | Yes |
| Usable in interrupt context | Yes (with the correct irq variant) | No — process context only |
| Typical use case | Short, fast read of in-memory state | Longer critical section, e.g. involving copy_to_user() |
The rw_semaphore API
A reader-writer semaphore is declared with DECLARE_RWSEM() for a static instance, or initialized at runtime with init_rwsem():
struct rw_semaphore ep_cfg_sem;
init_rwsem(&ep_cfg_sem);
/* or, for a statically declared semaphore */
DECLARE_RWSEM(ep_cfg_sem);
The core operations, all declared in <linux/rwsem.h>:
| Function | Behaviour |
|---|---|
down_read(&sem) | Take the lock for reading; sleeps (uninterruptibly) if a writer currently holds it |
up_read(&sem) | Release a reader’s hold |
down_write(&sem) | Take the lock exclusively for writing; sleeps until no readers or writers remain |
up_write(&sem) | Release the writer’s exclusive hold |
down_read_trylock(&sem) | Non-blocking attempt to take a read lock; returns immediately with success/failure |
down_write_trylock(&sem) | Non-blocking attempt to take the write lock |
down_read_killable(&sem) | Sleeps for a read lock, but can be woken by a fatal signal, returning -EINTR |
down_write_killable(&sem) | Same idea, for the write side |
An Original rw_semaphore Demo Driver
The following original example protects a small driver configuration block that is exposed to user space through a debugfs file. Because the read handler uses simple_read_from_buffer(), which can legitimately take page faults while touching user memory, a sleeping lock is the correct choice here — a plain spinlock would not be safe.
#include <linux/module.h>
#include <linux/debugfs.h>
#include <linux/rwsem.h>
#include <linux/uaccess.h>
struct ep_driver_config {
struct rw_semaphore sem;
char name[32];
u32 threshold;
};
static struct ep_driver_config ep_cfg;
static struct dentry *ep_dbgdir;
static ssize_t ep_cfg_read(struct file *f, char __user *ubuf,
size_t count, loff_t *ppos)
{
char kbuf[64];
int len;
down_read(&ep_cfg.sem);
len = scnprintf(kbuf, sizeof(kbuf), "name=%s threshold=%u\n",
ep_cfg.name, ep_cfg.threshold);
up_read(&ep_cfg.sem);
return simple_read_from_buffer(ubuf, count, ppos, kbuf, len);
}
static ssize_t ep_cfg_write(struct file *f, const char __user *ubuf,
size_t count, loff_t *ppos)
{
u32 new_threshold;
if (kstrtou32_from_user(ubuf, count, 10, &new_threshold))
return -EINVAL;
down_write(&ep_cfg.sem);
ep_cfg.threshold = new_threshold;
up_write(&ep_cfg.sem);
return count;
}
static const struct file_operations ep_cfg_fops = {
.owner = THIS_MODULE,
.read = ep_cfg_read,
.write = ep_cfg_write,
};
static int __init ep_rwsem_demo_init(void)
{
init_rwsem(&ep_cfg.sem);
strscpy(ep_cfg.name, "ep_demo", sizeof(ep_cfg.name));
ep_cfg.threshold = 100;
ep_dbgdir = debugfs_create_dir("ep_rwsem_demo", NULL);
debugfs_create_file("config", 0644, ep_dbgdir, NULL, &ep_cfg_fops);
pr_info("ep_rwsem_demo: loaded\n");
return 0;
}
static void __exit ep_rwsem_demo_exit(void)
{
debugfs_remove_recursive(ep_dbgdir);
pr_info("ep_rwsem_demo: unloaded\n");
}
module_init(ep_rwsem_demo_init);
module_exit(ep_rwsem_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original rw_semaphore demo");
Two Related Primitives You Will Meet Later in This Course
Reader-writer locks are not the only synchronization tools the kernel offers. Two others are worth knowing about by name at this point:
The Completion Mechanism
User-space Pthreads-based programs often use a condition variable to let one thread wait for another to signal that some event or state change has occurred. The Linux kernel’s equivalent is the completion mechanism, built around struct completion and the wait_for_completion() / complete() API pair. We will cover this in detail in a dedicated lecture later in this course.
The Sequence Lock (seqlock_t)
Where rwlock_t and rw_semaphore are optimized for mostly-read workloads, a sequence lock is optimized for the opposite case: data that is written far more often than it is read. Writers never block on readers at all; instead, readers detect and retry if a write happened concurrently, using a sequence counter. This makes it a natural fit for very hot, mostly-write kernel data — the classic example is the kernel’s own jiffies_64 timekeeping counter. We will return to seqlock_t in a future lecture as well.
Best Practices
- Choose
rw_semaphorewhenever the critical section might sleep — user-copy paths, memory allocation withGFP_KERNEL, or calls into other sleeping APIs. - Never use
rw_semaphorefrom interrupt context; it is process-context only. - Prefer the
_killablevariants for locks that a user-space-triggered code path might wait on for a long time, so a fatal signal can interrupt the wait cleanly. - As with any reader-writer primitive, watch out for writer starvation under sustained read pressure.
Common Mistakes
| Mistake | Why it’s a problem | Fix |
|---|---|---|
Calling down_read() from an interrupt handler | rw_semaphore can sleep and is not IRQ-safe at all | Use rwlock_t with an irq-safe variant instead |
| Holding the semaphore across a long-running operation | Blocks writers indefinitely and hurts overall throughput | Copy out what you need, unlock, then process it |
Forgetting to match down_write() with up_write() on every error path | Leaves the semaphore permanently held, deadlocking future writers | Use a single unlock point or goto-based cleanup |
Summary
A reader-writer semaphore gives you the same many-readers/one-writer model as a reader-writer spinlock, but allows sleeping inside the critical section, at the cost of only being usable in process context. Choose based on whether your critical section needs to sleep. Two related primitives — completions and sequence locks — round out the kernel’s synchronization toolbox and will be covered in upcoming lectures.
Frequently Asked Questions
Can I use rw_semaphore in an interrupt handler?
No. rw_semaphore may sleep, and interrupt handlers must never sleep. Use rwlock_t with the appropriate irq-safe variant instead.
What is the difference between down_read() and down_read_trylock()?
down_read() will sleep if a writer currently holds the lock, while down_read_trylock() returns immediately, reporting failure instead of waiting.
Why would I ever want down_write_killable() instead of down_write()?
It allows a fatal signal to interrupt a long wait for the lock, so a process being killed does not have to wait indefinitely — the call simply returns -EINTR.
Is a completion the same thing as a rw_semaphore?
No. A completion is for one-shot event signalling between tasks — “wait until this happens” — while rw_semaphore protects ongoing access to shared data.
When should I reach for seqlock_t instead of rw_semaphore?
When your data is written very frequently and read rarely, and the write side cannot tolerate being blocked by readers at all.
Conclusion
With both the reader-writer spinlock and reader-writer semaphore now covered, you have the two most common tools for many-readers/one-writer synchronization in Linux kernel drivers. The next lectures in this free Linux kernel development course move on to CPU cache effects — an equally important, often overlooked source of performance problems in concurrent kernel code.

2 Comments