Reader-Writer Semaphore (rw_semaphore) in Linux Kernel Drivers: Complete Guide-Linux Device Driver Training Online

« Previous Lecture | Next Lecture »

Reader-Writer Semaphore (rw_semaphore) in Linux Kernel Drivers: Complete Guide

Free Linux Kernel Programming Course — Kernel Synchronization Part 2

Kernel 6.x Updated
Original Driver Example
Beginner Friendly

A reader writer semaphore, or rw_semaphore, solves the one problem the reader-writer spinlock cannot: what happens when a reader or a writer needs to sleep while holding the lock — for example while copying data to or from user space. This lecture in our free Linux kernel development course covers the full rw_semaphore API, an original demo driver, and a quick look at two related kernel synchronization primitives you will meet later: completions and sequence locks.

reader writer semaphore linux kernel rw_semaphore api down_read up_read free linux device drivers course

What You Will Learn

  • The full rw_semaphore API: down_read(), up_read(), down_write(), up_write(), and their trylock/killable variants
  • When to choose a reader-writer semaphore over a reader-writer spinlock
  • A complete original kernel module protecting a driver’s configuration data with rw_semaphore
  • A short introduction to two related primitives: the completion mechanism and the sequence lock

Prerequisites

This lecture builds directly on the previous lecture’s coverage of read_lock() / write_lock(). You should already be comfortable with mutexes and reader-writer spinlocks from earlier in this free Linux kernel programming course.

Reader-Writer Semaphore vs Reader-Writer Spinlock

Conceptually, a reader-writer semaphore behaves exactly like a reader-writer spinlock: many readers can hold it together, a writer needs it exclusively. The difference is entirely in what happens when the lock is unavailable.

Propertyrwlock_t (spinlock)rw_semaphore
Waiting behaviourSpins the CPUSleeps, CPU is released to the scheduler
Safe to sleep inside critical sectionNoYes
Usable in interrupt contextYes (with the correct irq variant)No — process context only
Typical use caseShort, fast read of in-memory stateLonger critical section, e.g. involving copy_to_user()

The rw_semaphore API

A reader-writer semaphore is declared with DECLARE_RWSEM() for a static instance, or initialized at runtime with init_rwsem():

struct rw_semaphore ep_cfg_sem;
init_rwsem(&ep_cfg_sem);

/* or, for a statically declared semaphore */
DECLARE_RWSEM(ep_cfg_sem);

The core operations, all declared in <linux/rwsem.h>:

FunctionBehaviour
down_read(&sem)Take the lock for reading; sleeps (uninterruptibly) if a writer currently holds it
up_read(&sem)Release a reader’s hold
down_write(&sem)Take the lock exclusively for writing; sleeps until no readers or writers remain
up_write(&sem)Release the writer’s exclusive hold
down_read_trylock(&sem)Non-blocking attempt to take a read lock; returns immediately with success/failure
down_write_trylock(&sem)Non-blocking attempt to take the write lock
down_read_killable(&sem)Sleeps for a read lock, but can be woken by a fatal signal, returning -EINTR
down_write_killable(&sem)Same idea, for the write side
rw_semaphore: Readers and Writer Sleep Instead of Spin
Task A: down_read() —[reading, may sleep inside]— up_read() Task B: down_read() —[reading, may sleep inside]— up_read() Task C: down_write() –[goes to sleep, taken off CPU, woken later]– —-[writing]—- up_write()

An Original rw_semaphore Demo Driver

The following original example protects a small driver configuration block that is exposed to user space through a debugfs file. Because the read handler uses simple_read_from_buffer(), which can legitimately take page faults while touching user memory, a sleeping lock is the correct choice here — a plain spinlock would not be safe.

#include <linux/module.h>
#include <linux/debugfs.h>
#include <linux/rwsem.h>
#include <linux/uaccess.h>

struct ep_driver_config {
	struct rw_semaphore sem;
	char name[32];
	u32 threshold;
};

static struct ep_driver_config ep_cfg;
static struct dentry *ep_dbgdir;

static ssize_t ep_cfg_read(struct file *f, char __user *ubuf,
			    size_t count, loff_t *ppos)
{
	char kbuf[64];
	int len;

	down_read(&ep_cfg.sem);
	len = scnprintf(kbuf, sizeof(kbuf), "name=%s threshold=%u\n",
			 ep_cfg.name, ep_cfg.threshold);
	up_read(&ep_cfg.sem);

	return simple_read_from_buffer(ubuf, count, ppos, kbuf, len);
}

static ssize_t ep_cfg_write(struct file *f, const char __user *ubuf,
			     size_t count, loff_t *ppos)
{
	u32 new_threshold;

	if (kstrtou32_from_user(ubuf, count, 10, &new_threshold))
		return -EINVAL;

	down_write(&ep_cfg.sem);
	ep_cfg.threshold = new_threshold;
	up_write(&ep_cfg.sem);

	return count;
}

static const struct file_operations ep_cfg_fops = {
	.owner = THIS_MODULE,
	.read  = ep_cfg_read,
	.write = ep_cfg_write,
};

static int __init ep_rwsem_demo_init(void)
{
	init_rwsem(&ep_cfg.sem);
	strscpy(ep_cfg.name, "ep_demo", sizeof(ep_cfg.name));
	ep_cfg.threshold = 100;

	ep_dbgdir = debugfs_create_dir("ep_rwsem_demo", NULL);
	debugfs_create_file("config", 0644, ep_dbgdir, NULL, &ep_cfg_fops);

	pr_info("ep_rwsem_demo: loaded\n");
	return 0;
}

static void __exit ep_rwsem_demo_exit(void)
{
	debugfs_remove_recursive(ep_dbgdir);
	pr_info("ep_rwsem_demo: unloaded\n");
}

module_init(ep_rwsem_demo_init);
module_exit(ep_rwsem_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original rw_semaphore demo");

Two Related Primitives You Will Meet Later in This Course

Reader-writer locks are not the only synchronization tools the kernel offers. Two others are worth knowing about by name at this point:

The Completion Mechanism

User-space Pthreads-based programs often use a condition variable to let one thread wait for another to signal that some event or state change has occurred. The Linux kernel’s equivalent is the completion mechanism, built around struct completion and the wait_for_completion() / complete() API pair. We will cover this in detail in a dedicated lecture later in this course.

The Sequence Lock (seqlock_t)

Where rwlock_t and rw_semaphore are optimized for mostly-read workloads, a sequence lock is optimized for the opposite case: data that is written far more often than it is read. Writers never block on readers at all; instead, readers detect and retry if a write happened concurrently, using a sequence counter. This makes it a natural fit for very hot, mostly-write kernel data — the classic example is the kernel’s own jiffies_64 timekeeping counter. We will return to seqlock_t in a future lecture as well.

Best Practices

  • Choose rw_semaphore whenever the critical section might sleep — user-copy paths, memory allocation with GFP_KERNEL, or calls into other sleeping APIs.
  • Never use rw_semaphore from interrupt context; it is process-context only.
  • Prefer the _killable variants for locks that a user-space-triggered code path might wait on for a long time, so a fatal signal can interrupt the wait cleanly.
  • As with any reader-writer primitive, watch out for writer starvation under sustained read pressure.

Common Mistakes

MistakeWhy it’s a problemFix
Calling down_read() from an interrupt handlerrw_semaphore can sleep and is not IRQ-safe at allUse rwlock_t with an irq-safe variant instead
Holding the semaphore across a long-running operationBlocks writers indefinitely and hurts overall throughputCopy out what you need, unlock, then process it
Forgetting to match down_write() with up_write() on every error pathLeaves the semaphore permanently held, deadlocking future writersUse a single unlock point or goto-based cleanup

Summary

A reader-writer semaphore gives you the same many-readers/one-writer model as a reader-writer spinlock, but allows sleeping inside the critical section, at the cost of only being usable in process context. Choose based on whether your critical section needs to sleep. Two related primitives — completions and sequence locks — round out the kernel’s synchronization toolbox and will be covered in upcoming lectures.

Frequently Asked Questions

Can I use rw_semaphore in an interrupt handler?

No. rw_semaphore may sleep, and interrupt handlers must never sleep. Use rwlock_t with the appropriate irq-safe variant instead.

What is the difference between down_read() and down_read_trylock()?

down_read() will sleep if a writer currently holds the lock, while down_read_trylock() returns immediately, reporting failure instead of waiting.

Why would I ever want down_write_killable() instead of down_write()?

It allows a fatal signal to interrupt a long wait for the lock, so a process being killed does not have to wait indefinitely — the call simply returns -EINTR.

Is a completion the same thing as a rw_semaphore?

No. A completion is for one-shot event signalling between tasks — “wait until this happens” — while rw_semaphore protects ongoing access to shared data.

When should I reach for seqlock_t instead of rw_semaphore?

When your data is written very frequently and read rarely, and the write side cannot tolerate being blocked by readers at all.

Conclusion

With both the reader-writer spinlock and reader-writer semaphore now covered, you have the two most common tools for many-readers/one-writer synchronization in Linux kernel drivers. The next lectures in this free Linux kernel development course move on to CPU cache effects — an equally important, often overlooked source of performance problems in concurrent kernel code.

« Previous Lecture | Next Lecture »

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *