Lockdep: Catching Kernel Deadlocks Before They Happen-Free Linux Device Driver Training in Hyderabad

Lockdep: Catching Kernel Deadlocks Before They Happen
A free, practical guide to the Linux kernel lock validator, updated for kernel 6.12 and later
Free Linux Kernel Development Course
Free Linux Device Drivers Course
Beginner Friendly

If you write Linux kernel modules or device drivers that use spinlocks, mutexes, or any other locking primitive, sooner or later you will make a locking mistake. The kernel does not forgive locking mistakes gracefully — a bad lock ordering can freeze an entire CPU, hang a driver forever, or bring down a production server with no warning at all. This is exactly the class of bug that lockdep, the kernel’s runtime lock dependency validator, was built to catch before it turns into a real deadlock.

In this lecture of our free Linux kernel development course, we will build a clear, from-scratch understanding of how lockdep works, how to enable it on a modern kernel, and how to read its warning reports when something goes wrong in your driver code.

What You Will Learn
What a kernel deadlock actually is How lockdep builds a lock dependency graph Enabling lockdep on kernel 6.12+ Self-deadlock (AA deadlock) Circular dependency (AB-BA) deadlock Reading a lockdep warning report Lock ordering best practices

Prerequisites

  • Comfort writing and inserting a basic Linux kernel module (insmod / rmmod)
  • Basic understanding of spinlocks and mutexes in the kernel
  • A Linux VM you can safely crash and reboot (never test deadlocks on a production box)
  • Kernel 6.12 or newer is assumed for configuration option names in this lecture

What Is a Deadlock, in Plain Terms?

A deadlock happens when two or more threads of execution are each waiting for a resource that another one of them is holding, and none of them can ever proceed. In kernel space, that “resource” is usually a lock — a spinlock, mutex, semaphore, or rwlock. Unlike a userspace deadlock, a kernel deadlock on a spinlock does not just freeze one process; it can spin a CPU core at 100% forever, effectively taking that core out of service and, depending on which lock is involved, hanging the whole machine.

The frustrating part of deadlocks is that they are usually timing dependent. Your driver might work perfectly for months and then lock up the very first time two code paths race against each other in just the wrong order. This is precisely why static code review alone is not reliable for catching them — you need a tool that watches actual lock acquisition order at runtime.

What Is Lockdep, the Kernel Lock Validator?

Lockdep is a debugging subsystem built directly into the Linux kernel. Instead of waiting for a deadlock to actually occur, lockdep tracks the order in which every lock in the system is acquired and released, and it builds a running dependency graph out of that history. The moment it observes a locking sequence that could theoretically lead to a deadlock — even if the deadlock does not actually occur on that particular run — it prints a detailed warning to the kernel log and identifies exactly which locks and which code paths are responsible.

This proactive design is what makes lockdep so valuable during development. A deadlock might only manifest once in ten thousand boots on real hardware, but lockdep can flag the unsafe lock ordering the very first time your driver code takes that path, long before it becomes a 3 a.m. production incident.

How Lockdep Thinks About Locks
Lock Class Every lock is grouped into a “class” based on where it is initialized in the source code, not by its runtime address. This lets lockdep track a lock type used by thousands of instances (e.g. one mutex per inode) as a single class.
Dependency Edge Whenever lock B is taken while lock A is already held, lockdep records a directed edge A → B in its graph.
Cycle Detection If the graph ever contains a cycle — for example A → B and also B → A — a deadlock is theoretically possible, and lockdep reports it immediately.

Enabling Lockdep on Kernel 6.12 and Later

Most distribution kernels ship with lockdep disabled by default because it adds measurable runtime overhead. For kernel development and driver debugging, you should build a debug kernel with it turned on. The relevant options, unchanged in structure through the 6.12 series, live under kernel hacking in menuconfig:

CONFIG_DEBUG_KERNEL=y
CONFIG_LOCKDEP=y
CONFIG_LOCKDEP_SUPPORT=y
CONFIG_PROVE_LOCKING=y
CONFIG_DEBUG_LOCK_ALLOC=y
CONFIG_DEBUG_ATOMIC_SLEEP=y

The important one is CONFIG_PROVE_LOCKING. This is the option that actually enables the full “proof” engine that walks the dependency graph on every lock acquisition looking for cycles. On a self-built kernel, enable it with:

make menuconfig
# Kernel hacking ---> Lock Debugging (spinlocks, mutexes, etc...)
#   [*] Lock debugging: prove locking correctness

make -j$(nproc)
sudo make modules_install install

Once you boot into a lockdep-enabled kernel, no extra step is needed in your own module code. Every spinlock, mutex, and rwlock you use is automatically tracked, and lockdep starts building its dependency graph from the moment the kernel boots.

Self-Deadlock: The AA Deadlock

The simplest possible deadlock is a self-deadlock, sometimes called an AA deadlock, where a single thread tries to acquire a lock it is already holding. This typically happens by accident when a helper function acquires the same lock as its caller, and nobody notices the recursion during code review.

Here is a minimal, original example of the mistake inside a kernel module worker function:

static DEFINE_MUTEX(drv_state_lock);

static void drv_update_status(struct my_device *dev)
{
    mutex_lock(&drv_state_lock);
    dev->status = DEV_STATUS_BUSY;
    /* forgot to unlock before calling the helper below */
    drv_refresh_counters(dev);
    mutex_unlock(&drv_state_lock);
}

static void drv_refresh_counters(struct my_device *dev)
{
    mutex_lock(&drv_state_lock);   /* BUG: same lock, same thread! */
    dev->counter++;
    mutex_unlock(&drv_state_lock);
}

On a non-lockdep kernel, calling drv_refresh_counters() from inside drv_update_status() would spin the CPU forever waiting for a mutex that will never be released, because the only thread that can release it is the one currently spinning. With lockdep enabled, the instant mutex_lock() is called a second time on the same class by the same task, the kernel log immediately shows a report similar to this:

=============================================
[ INFO: possible recursive locking detected ]
---------------------------------------------
drv_worker/842 is trying to acquire lock:
 drv_state_lock, at: drv_refresh_counters+0x1c/0x40

but task is already holding lock:
 drv_state_lock, at: drv_update_status+0x22/0x60

other info that might help us debug this:
 Possible unsafe locking scenario:

       CPU0
       ----
  lock(drv_state_lock);
  lock(drv_state_lock);

 *** DEADLOCK ***

Notice that lockdep names the exact lock, the exact function, and the exact offset where the second acquisition happened. Fixing this class of bug almost always means restructuring the code so the lock is released before calling a helper, or introducing an “already locked” internal variant of the helper that assumes the lock is held.

Circular Dependency Deadlock: The AB-BA Pattern

A more subtle and far more common real-world deadlock involves two different locks acquired by two different threads in opposite order. This is called an AB-BA deadlock, and it is one of the most frequent bugs found in multi-threaded and multi-core driver code.

Suppose a driver uses two spinlocks, one protecting a hardware register cache and one protecting a software queue. If one code path always takes the register lock first and then the queue lock, but another code path takes the queue lock first and then the register lock, you have set a trap that will eventually spring on a busy multi-core system.

AB-BA Deadlock Sequence Across Two CPU Cores
Time CPU 0 (Thread A) CPU 1 (Thread B)
t0 takes reg_lock takes queue_lock
t1 tries queue_lock → blocked tries reg_lock → blocked
t2 Both CPUs spin forever — neither lock is ever released

Here is an original, simplified example of the two code paths that create this trap:

static DEFINE_SPINLOCK(reg_lock);
static DEFINE_SPINLOCK(queue_lock);

/* Path A: interrupt handler updates a register, then queues work */
static void drv_irq_handler(void)
{
    spin_lock(&reg_lock);
    spin_lock(&queue_lock);
    /* update register, push to queue */
    spin_unlock(&queue_lock);
    spin_unlock(&reg_lock);
}

/* Path B: a workqueue handler drains the queue, then touches a register */
static void drv_queue_worker(struct work_struct *w)
{
    spin_lock(&queue_lock);
    spin_lock(&reg_lock);   /* BUG: opposite order from drv_irq_handler() */
    /* drain queue, update register */
    spin_unlock(&reg_lock);
    spin_unlock(&queue_lock);
}

The two functions look completely correct in isolation. Each one locks, does its work, and unlocks cleanly. The bug only exists in the relationship between them, which is exactly the kind of bug that is nearly impossible to spot with a plain code review, especially once a driver grows past a few thousand lines. Lockdep, however, has been recording the dependency edges reg_lock → queue_lock (from Path A) and queue_lock → reg_lock (from Path B), and the moment both edges exist it reports a circular dependency, whether or not the two paths ever actually raced against each other yet on your test machine:

======================================================
WARNING: possible circular locking dependency detected
------------------------------------------------------
drv_worker/901 is trying to acquire lock:
 reg_lock, at: drv_queue_worker+0x30/0x90

but task is already holding lock:
 queue_lock, at: drv_queue_worker+0x18/0x90

which lock already depends on the new lock.

Possible unsafe locking scenario:

       CPU0                CPU1
       ----                ----
  lock(reg_lock);
                            lock(queue_lock);
                            lock(reg_lock);
  lock(queue_lock);

 *** DEADLOCK ***

This is the single most valuable thing lockdep does: it can flag a circular dependency the very first time both code paths have ever run on the system, even under light testing, long before a real deadlock would show up under production load.

How to Read a Lockdep Warning Report

Report Field What It Tells You
“is trying to acquire lock” The lock and exact code location where the unsafe attempt happened.
“but task is already holding lock” The lock currently held that creates the conflict.
“Possible unsafe locking scenario” A two-column CPU timeline showing exactly how the two acquisition orders conflict.
Stack trace (below the summary) The full call chain leading to each lock acquisition, useful for finding the offending code path fast.

One important detail: a lockdep circular dependency warning does not always mean a deadlock has actually occurred on that boot. It means the kernel has proven that a deadlock is possible given the lock orders it has observed. Treat every such warning as a real bug to fix, even if your test run did not actually freeze.

Common Mistakes and Troubleshooting Tips

  • Ignoring a lockdep splat because “it still boots fine.” A proven possible deadlock will eventually trigger under the right timing; treat every warning as a must-fix bug.
  • Testing only on a single-core VM. Many AB-BA deadlocks need two CPUs racing at the same time to actually freeze; lockdep catches them either way, but manual reproduction needs multi-core hardware.
  • Forgetting to rebuild with CONFIG_PROVE_LOCKING. Without this option, the kernel logs almost nothing useful about lock ordering.
  • Assuming irqsave variants are always required. Lockdep also flags irq-context lock misuse separately; read the report’s “irq” annotations carefully rather than adding _irqsave everywhere by habit.
  • Not distinguishing AA from AB-BA reports. The fix for a self-deadlock is almost always a code restructuring; the fix for AB-BA is almost always a consistent, documented lock ordering rule.

Best Practices for Lock Ordering

  • Document a strict, global lock ordering rule for every pair of locks your driver uses, and put it in a code comment right next to the lock definitions.
  • Prefer acquiring the “outer” lock in a fixed hierarchy every single time, in every code path, including error handling paths.
  • Keep critical sections as short as possible; the less time a lock is held, the smaller the window for contention or misuse.
  • Where two independent locks genuinely need to be taken together, consider whether a single combined lock is actually simpler and safer.
  • Always test new locking code on a lockdep-enabled debug kernel before it ever reaches a staging or production build.

Performance Considerations

Lockdep is a development and debugging tool, not something you want enabled on a production kernel. Because it validates every single lock acquisition against a growing dependency graph, it adds real, measurable overhead, particularly on workloads that take and release locks at very high frequency. The standard practice is to run lockdep-enabled debug kernels throughout development and CI testing, then ship production kernels with CONFIG_PROVE_LOCKING disabled once the code has been proven clean.

Security Considerations

An undetected deadlock is not just an inconvenience — it is a denial-of-service condition. A kernel module that can be driven into a self-deadlock or an AB-BA deadlock by a malicious or malformed input can be used to hang an entire system remotely, which is a serious concern for network-facing drivers and filesystems. Treat every lockdep warning found during fuzzing or security testing with the same seriousness as a memory-safety bug, since the practical impact — a fully unresponsive system — can be just as severe.

Real-World Use Cases

Lockdep runs continuously on kernel developers’ test machines, in automated continuous integration systems like the kernel test robot, and on distribution build farms. It has caught countless real bugs in mainline drivers, filesystems, and networking code over the years, often in code paths that only two subsystems combined would ever trigger together, which is exactly the kind of interaction a human reviewer is unlikely to spot by reading code alone.

Summary and Key Takeaways

  • Lockdep tracks every lock acquisition in the kernel and builds a live dependency graph.
  • It detects both self-deadlocks (AA) and circular dependency deadlocks (AB-BA) before they actually freeze the system.
  • Enable it with CONFIG_PROVE_LOCKING=y on your development kernel.
  • A lockdep warning means a deadlock is provably possible — always fix it, even if nothing froze on that particular run.
  • A documented, consistent lock ordering rule is the single best defense against AB-BA deadlocks.

Conclusion

Deadlocks are among the hardest kernel bugs to reproduce and the easiest to introduce by accident, especially once a driver grows multiple locks and multiple entry points. Lockdep turns an unpredictable, timing-dependent failure into a clear, actionable report the moment an unsafe lock order is even theoretically possible. Making a lockdep-enabled debug kernel part of your everyday development workflow is one of the highest-value habits you can build as an embedded Linux or device driver engineer.

Frequently Asked Questions

Q1. Does lockdep only detect deadlocks that have already happened?
No. Lockdep’s entire value is that it proves a deadlock is possible from the observed lock ordering, often long before the exact unlucky timing needed to trigger the freeze actually occurs.

Q2. Can I use lockdep on an ARM or embedded board, or only x86?
Lockdep is architecture independent and works on ARM, ARM64, RISC-V, and every other architecture the mainline kernel supports.

Q3. Will lockdep slow down my production system?
Yes, noticeably, which is why it should be enabled only on development and CI kernels, not shipped in production builds.

Q4. What is the difference between an AA deadlock and an AB-BA deadlock?
An AA deadlock is a single thread trying to reacquire a lock it already holds. An AB-BA deadlock involves two threads acquiring two different locks in opposite order.

Q5. Does lockdep work with mutexes as well as spinlocks?
Yes. Lockdep tracks spinlocks, mutexes, rwlocks, semaphores, and several other kernel locking primitives.

Q6. My kernel log shows a lockdep warning but the system did not freeze. Should I still fix it?
Yes, always. A proven unsafe lock order will eventually trigger under different timing, load, or hardware.

Q7. Is lockdep the same thing as KASAN or KMEMLEAK?
No. KASAN detects memory-safety bugs and KMEMLEAK detects memory leaks; lockdep is specifically focused on lock ordering and deadlock detection.

Q8. How do I turn lockdep off if I only want it temporarily?
Rebuild the kernel with CONFIG_PROVE_LOCKING disabled, or boot a separate non-debug kernel image for performance testing.

Continue Your Free Linux Kernel Development Journey

This lecture is part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.

Browse the Full Course Join the Community

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *