← Previous Lecture | Next Lecture →
In this free linux kernel development course lecture we cover lockdep annotations in depth — specifically the lockdep_assert_held() macro — and we watch what actually happens on the console when an AB-BA deadlock is not just detected on paper but genuinely freezes two CPUs. You will see how the kernel’s soft lockup watchdog backs up lockdep, and how to write your own runtime lock assertions in a kernel 6.x driver.
Two Warnings, One Root Cause
A genuine AB-BA deadlock isn’t announced by a single event — the kernel gives you two independent signals, and understanding both is the point of this lockdep annotations lesson. The first signal is lockdep’s own circular-dependency warning, printed the instant the second thread tries to complete the opposite-order lock chain. The second signal, arriving many seconds later, is the soft lockup watchdog, which exists precisely for the case where a CPU never comes back to the scheduler at all.
It helps to be precise about timing here: lockdep’s warning is not a rescue. It is a printk. By the time the “possible circular locking dependency detected” message reaches your console, the second thread has already blocked on the contended lock. lockdep observed the cycle forming and reported it, but it does not unwind the call stack or release anything for you. The thread is still stuck.
Reproducing an AB-BA Deadlock Safely
Warning: the module below is designed to hang two CPUs on purpose. Only load it inside a disposable virtual machine that you can force-reset, never on hardware you care about.
The demo module below pins one kernel thread to CPU 0 and one to CPU 1 using kthread_bind(), then lets a module parameter decide whether the second thread follows the documented lock order or violates it:
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/spinlock.h>
#include <linux/delay.h>
static DEFINE_SPINLOCK(ep_lockA);
static DEFINE_SPINLOCK(ep_lockB);
static struct task_struct *ep_t0, *ep_t1;
static int violate_order;
module_param(violate_order, int, 0644);
MODULE_PARM_DESC(violate_order, "1 = deliberately break the A-then-B lock order");
static int ep_worker0(void *arg)
{
pr_info("ep_ab_ba: thread0 taking lockA then lockB\n");
spin_lock(&ep_lockA);
mdelay(50);
spin_lock(&ep_lockB);
pr_info("ep_ab_ba: thread0 got both locks\n");
spin_unlock(&ep_lockB);
spin_unlock(&ep_lockA);
return 0;
}
static int ep_worker1(void *arg)
{
if (violate_order) {
pr_info("ep_ab_ba: thread1 taking lockB then lockA (WRONG order)\n");
spin_lock(&ep_lockB);
mdelay(50);
spin_lock(&ep_lockA);
} else {
pr_info("ep_ab_ba: thread1 taking lockA then lockB (correct order)\n");
spin_lock(&ep_lockA);
mdelay(50);
spin_lock(&ep_lockB);
}
pr_info("ep_ab_ba: thread1 got both locks\n");
spin_unlock(&ep_lockB);
spin_unlock(&ep_lockA);
return 0;
}
static int __init ep_ab_ba_init(void)
{
ep_t0 = kthread_create(ep_worker0, NULL, "ep_ab_ba0");
ep_t1 = kthread_create(ep_worker1, NULL, "ep_ab_ba1");
if (IS_ERR(ep_t0) || IS_ERR(ep_t1))
return -ENOMEM;
kthread_bind(ep_t0, 0);
kthread_bind(ep_t1, 1);
wake_up_process(ep_t0);
wake_up_process(ep_t1);
return 0;
}
static void __exit ep_ab_ba_exit(void)
{
pr_info("ep_ab_ba: module removed\n");
}
module_init(ep_ab_ba_init);
module_exit(ep_ab_ba_exit);
MODULE_LICENSE("GPL");
With violate_order=0 both threads acquire the locks in the same order, so at worst one thread waits briefly for the other — no cycle is possible. With violate_order=1, thread 0 is holding lockA and waiting for lockB at the same moment thread 1 is holding lockB and waiting for lockA. Neither can ever make progress. That mutual, permanent block is the deadlock itself.
What the Console Actually Shows
On a lock-debugging kernel 6.x build (CONFIG_PROVE_LOCKING=y), loading the module with violate_order=1 produces, in order: a lockdep “possible circular locking dependency detected” block naming both locks and both call sites, followed some seconds later by watchdog: BUG: soft lockup - CPU#0 stuck for 22s! and the matching message for CPU 1. The soft lockup message is printed at KERN_EMERG, which is why it can still reach your console even though two CPUs are effectively frozen — it’s the kernel’s last resort for telling you something has gone badly wrong.
lockdep only ever prints its circular-dependency warning the first time a given pair of lock classes is used in a conflicting order. If you reload the module a second time, you will not see the warning repeat, even though the underlying bug is unchanged. This is a deliberate design choice to avoid flooding the log, and it is worth remembering when a bug seems to have “disappeared” after a reboot.
/proc/lockdep_chains: The Recorded Evidence
Every distinct locking sequence lockdep has observed is kept as a “chain.” After the AB-BA run above, reading /proc/lockdep_chains shows both orderings recorded against the same pair of lock classes — proof, independent of any single stack trace, that the two threads really did take the locks in opposite sequence:
$ sudo cat /proc/lockdep_chains
irq_context: 0
[ep_lockA]
[ep_lockB]
irq_context: 0
[ep_lockB]
[ep_lockA]
Seeing both chains for the same two locks is itself a red flag worth searching your codebase for, independent of whether a deadlock has actually occurred yet.
lockdep_assert_held(): Proactive Lock Assertions
Circular-dependency detection is automatic — it needs no help from you. But there is a second, complementary lockdep feature that you do have to opt into: the lockdep_assert_held() family of annotations. Where circular-dependency detection catches wrong lock ordering, assertions catch code that assumes a lock is held when it might not be — a common bug in helper functions called from multiple paths, some of which forgot to lock first.
The macro itself is a thin, debug-only wrapper:
// include/linux/lockdep.h (kernel 6.x)
#define lockdep_assert_held(l) do { \
WARN_ON(debug_locks && !lockdep_is_held(l)); \
} while (0)
When lock debugging is compiled in, the assertion checks — at zero cost in a release build, since it compiles away entirely when CONFIG_LOCKDEP is off — whether the current task genuinely holds the specified lock right now. If it does not, WARN_ON() fires with a full stack trace, telling you exactly which call path forgot to lock.
Here is an original example: a helper that updates shared driver state and documents its locking contract with an assertion instead of a comment:
#include <linux/module.h>
#include <linux/spinlock.h>
static DEFINE_SPINLOCK(ep_state_lock);
static int ep_counter;
/* Caller MUST already hold ep_state_lock. */
static void ep_update_counter(int delta)
{
lockdep_assert_held(&ep_state_lock);
ep_counter += delta;
}
static void ep_safe_caller(void)
{
spin_lock(&ep_state_lock);
ep_update_counter(1); /* fine: lock is held */
spin_unlock(&ep_state_lock);
}
static void ep_buggy_caller(void)
{
ep_update_counter(1); /* BUG: lockdep_assert_held() fires */
}
The assertion turns an unwritten contract (“call this only while holding the lock”) into something the kernel verifies for you every time the function runs on a debug build. This is exactly the kind of annotation you’ll find throughout core kernel code and mainline drivers.
Comparison: Detection Mechanisms at a Glance
| Mechanism | Triggers when | What it needs | What it tells you |
|---|---|---|---|
| lockdep circular dependency check | A new lock-order edge completes a cycle | CONFIG_PROVE_LOCKING, automatic | Which two locks, which two call sites |
| lockdep_assert_held() | Annotated function runs without the required lock | Manually added by the developer | Exact call site missing the lock |
| Soft lockup watchdog | A CPU makes no scheduling progress for ~20s (kernel.watchdog_thresh) | CONFIG_SOFTLOCKUP_DETECTOR, always on by default | Which CPU and which task is stuck, as a backstop |
Common Mistakes and Troubleshooting Tips
- Assuming lockdep prevents deadlocks: it only reports them; your code still has to avoid the bad ordering.
- Expecting a repeat warning on reload: lockdep reports a given chain violation only once per boot.
- Testing on a non-debug kernel: without CONFIG_PROVE_LOCKING, neither the circular-dependency check nor your assertions produce useful output — always validate new locking code on a debug kernel first.
- Forgetting the watchdog threshold:
kernel.watchdog_threshcontrols how many seconds of unresponsiveness trigger a soft lockup report; tune it in test environments if you need faster feedback.
Best Practices
- Document lock ordering in a single comment near the lock declarations, not scattered across call sites.
- Add
lockdep_assert_held()to any static helper that is only ever safe to call under a specific lock. - Run new locking code under a debug kernel with CONFIG_PROVE_LOCKING before it ever reaches a shared test machine.
- Treat a second entry appearing for the same lock pair in
/proc/lockdep_chainsas a bug report, not a curiosity.
Performance and Security Considerations
lockdep tracking and assertions add measurable overhead, which is why production kernels typically ship with CONFIG_PROVE_LOCKING disabled — the checks compile away completely, leaving zero runtime cost. From a security-hardening angle, catching lock misuse during development matters because an unprotected shared-state bug is often the same class of flaw that turns into a use-after-free or a privilege-escalation primitive if it ships in production.
Summary / Key Takeaways
- lockdep’s circular-dependency warning and the actual thread hang occur almost simultaneously — the warning is diagnostic, not preventive.
- The soft lockup watchdog is an independent, time-based backstop for CPUs that never return to the scheduler.
/proc/lockdep_chainsgives you hard evidence of every lock order the kernel has observed.lockdep_assert_held()lets you turn an unwritten locking contract into a checked assertion.- lockdep reports a given violation only once per boot, so don’t rely on a repeat warning during testing.
Conclusion
Lock ordering bugs are some of the hardest kernel issues to reproduce on demand, which is exactly why lockdep exists: it turns a bug that might take days of stress-testing to hit into a warning on the very first conflicting acquisition. Pairing that automatic detection with your own lockdep_assert_held() annotations gives you two independent nets under the same rope — one that catches wrong ordering, and one that catches a lock that was simply never taken at all. Both are effectively free on a release build and invaluable on a debug one, which is why disciplined kernel and driver developers treat them as a normal part of writing new locking code, not an occasional debugging tool.
Frequently Asked Questions
What does lockdep_assert_held() actually check at runtime?
It checks whether the currently running task holds the specific lock instance you pass in, using lockdep’s internal held-lock tracking, and fires WARN_ON() if it does not.
Does lockdep_assert_held() have any cost on a production kernel?
No. When CONFIG_LOCKDEP is not enabled, the macro compiles down to nothing, so shipping the annotation costs nothing in a release build.
Why did my system hang even though lockdep printed a warning?
lockdep only reports the cycle it detects; it does not release any lock or unwind any call stack for you. The thread that triggered the warning is still blocked immediately afterward.
What is the difference between a soft lockup and a hard lockup?
A soft lockup means a CPU has not returned to the scheduler for the configured threshold, but interrupts are still being serviced. A hard lockup means even interrupts on that CPU have stopped, which the NMI watchdog detects separately.
How do I change how long the watchdog waits before reporting a soft lockup?
Set the kernel.watchdog_thresh sysctl to the desired number of seconds; the softlockup report fires at roughly twice that threshold.
Can lockdep_assert_held() be used with mutexes as well as spinlocks?
Yes. The same annotation family works across spinlocks, mutexes, rw_semaphores and rwlocks, since they all register with the same lockdep tracking infrastructure.
Why does lockdep only report a given chain violation once?
It is a deliberate design decision to avoid flooding the kernel log with the same warning on every subsequent conflicting acquisition once the cycle is already known.
Is it safe to run the AB-BA demo module on my main development machine?
No. With violate_order=1 it is designed to genuinely hang two CPUs. Only run it inside a disposable virtual machine that you can force-reset.
More free linux device drivers course and free embedded systems course lectures are on the way.

2 Comments