← Previous Lecture | Next Lecture →
If you have spent any real time debugging kernel locking with lockdep, you already know it is not a perfect tool. Understanding common linux kernel lockdep issues — like lock class exhaustion after repeated module reloads, or the confusing “lock debugging disabled” warning — will save you hours of chasing a ghost bug that lockdep itself created. In this lecture we cover these lockdep quirks, then move on to a genuinely useful diagnostic feature many driver developers never touch: kernel lock statistics, which tell you exactly which locks in your system are being fought over and for how long.
This lecture is part of our free, community-driven Linux kernel programming course. All examples are written from scratch for kernel 6.x and tested conceptually against current mainline behavior — no outdated APIs, no dead command paths.
You should already be comfortable with basic mutex and spinlock usage in kernel modules, and ideally have gone through our earlier lockdep introduction lecture in this series. A kernel 6.x build with CONFIG_PROVE_LOCKING and debugfs mounted is recommended so you can follow the hands-on steps yourself.
Understanding Linux Kernel Lockdep Issues and Lock Statistics
Why Lockdep Itself Can Misbehave
Lockdep is a static analysis engine built into the kernel, and like any analysis engine it has resource limits. The most common of the linux kernel lockdep issues developers run into during driver development is running out of lock classes. Every time you insmod a kernel module, the kernel allocates a fresh set of lock classes for every lock defined inside that module. Unloading the module with rmmod does not free those classes — the kernel keeps them around so it can reuse the tracking data if the same module is loaded again later. Reload the same module a few dozen times during a debugging session, and you can quietly exhaust lockdep’s internal lock class table.
The practical fix is simple: avoid tight insmod/rmmod cycles once you notice lockdep behaving oddly, and reboot the test machine between long debugging sessions instead of repeatedly reloading. This is especially relevant for driver code with large arrays of embedded locks, such as a per-descriptor spinlock inside a ring buffer structure — a single array-heavy driver can consume a large chunk of the class table on one load.
rmmod driver.ko -> classes NOT freed, marked reusable
insmod driver.ko -> reuses same N classes (fine)
… repeat 50+ times while iterating on a bug …
insmod driver.ko -> lockdep: BUG: MAX_LOCKDEP_KEYS exceeded
The “Lock Debugging Disabled” Warning
A second common source of confusion is this message appearing in dmesg:
WARNING: lock debugging disabled!! - possibly due to a lockdep warning
This is not a new bug — it is lockdep telling you it already gave up earlier. Internally the kernel keeps a flag, debug_locks, which gets forced to zero the moment lockdep detects something it considers unrecoverable, even on a fully debug-enabled kernel. Once that flag flips, all further lock debugging output is suppressed, which means a completely unrelated warning later in your session may really be a symptom of an earlier lockdep failure you scrolled past. The correct response is to scroll back through dmesg for the first lockdep warning, fix that, and reboot before trusting any subsequent lock debugging output.
KCSAN as a Complementary Data-Race Detector
Lockdep proves things about lock ordering, but it does not directly catch plain data races on unprotected memory. The Kernel Concurrency Sanitizer (KCSAN) fills that gap using compile-time instrumentation, and by kernel 6.x it has matured into a standard, well-supported option rather than an experimental patch. If you are chasing intermittent memory corruption that lockdep says nothing about, building a kernel with KCSAN enabled is usually the next step, since it can catch races that never manifest as a classic self-deadlock or AB-BA lock ordering violation.
Enabling Kernel Lock Statistics
Once your locking is passing lockdep cleanly, the next question is usually a performance one: which lock is actually slowing my driver down under load? This is where kernel lock statistics come in. Lockdep already inserts hooks into every lock acquire and release path, and the lock statistics subsystem reuses those same hooks to record timing data instead of just correctness data.
To use this feature your kernel must be built with CONFIG_LOCK_STAT=y. Without it, /proc/lock_stat simply will not exist, which is the case on most stock distribution kernels — you generally need a custom debug build for this, the same kind of build most of this course’s synchronization lectures assume.
| Action | Command |
|---|---|
| Clear existing statistics | echo 0 | sudo tee /proc/lock_stat |
| Enable statistics collection | echo 1 | sudo tee /proc/sys/kernel/lock_stat |
| Disable statistics collection | echo 0 | sudo tee /proc/sys/kernel/lock_stat |
| View all recorded stats | sudo cat /proc/lock_stat |
Reading /proc/lock_stat Fields
The output columns look intimidating the first time you see them, but each one maps to a specific moment in a lock’s life cycle. Understanding these fields is the real payoff of this section on linux kernel lock statistics.
| Field | Meaning |
|---|---|
| class name | The symbolic lock class, i.e. which lock in your source this row represents |
| acquisitions | Total number of times this lock was successfully taken |
| contentions | Number of times a task had to wait because the lock was already held |
| con-bounces | Contentions that additionally bounced the cacheline between CPUs |
| waittime-min/max/total/avg | Time spent waiting for a contended lock, in microseconds |
| holdtime-min/max/total/avg | Time the lock was actually held once acquired, in microseconds |
A lock with high contentions and a large holdtime-avg is your prime suspect for a scalability bottleneck — it is both popular and slow, which is the worst combination for a multi-core system. A lock with high con-bounces specifically points at cacheline ping-pong between CPU cores, a topic we covered in more depth in this course’s false-sharing lecture.
Hands-On: An Original Lock Statistics Demo
Below is an original demo driver, ep_lockstat_demo, written for this course. It deliberately creates contention on a single spinlock by running two kernel threads that both hammer a shared counter, so you have something worth measuring in /proc/lock_stat.
#include <linux/module.h>
#include <linux/kthread.h>
#include <linux/spinlock.h>
#include <linux/delay.h>
static DEFINE_SPINLOCK(ep_counter_lock);
static unsigned long ep_shared_counter;
static struct task_struct *ep_worker_a, *ep_worker_b;
static int ep_worker_fn(void *data)
{
while (!kthread_should_stop()) {
spin_lock(&ep_counter_lock);
ep_shared_counter++;
udelay(50); /* hold the lock briefly to force contention */
spin_unlock(&ep_counter_lock);
usleep_range(100, 200);
}
return 0;
}
static int __init ep_lockstat_demo_init(void)
{
ep_worker_a = kthread_run(ep_worker_fn, NULL, "ep_lockstat_a");
ep_worker_b = kthread_run(ep_worker_fn, NULL, "ep_lockstat_b");
pr_info("ep_lockstat_demo: two workers contending on ep_counter_lock\n");
return 0;
}
static void __exit ep_lockstat_demo_exit(void)
{
kthread_stop(ep_worker_a);
kthread_stop(ep_worker_b);
pr_info("ep_lockstat_demo: stopped\n");
}
module_init(ep_lockstat_demo_init);
module_exit(ep_lockstat_demo_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala lock statistics demo");
With CONFIG_LOCK_STAT enabled, load the module, let it run for a few seconds, then filter the output for just this lock:
sudo insmod ep_lockstat_demo.ko
sleep 5
sudo grep -A2 "ep_counter_lock" /proc/lock_stat
sudo rmmod ep_lockstat_demo
Because both worker threads hold the spinlock for roughly 50 microseconds on every iteration, you should see non-zero contentions and a measurable holdtime-avg close to that value — a concrete, self-generated example instead of an abstract one.
Beyond Lockdep: Other Concurrency Debugging Tools
Kernel-space lockdep and lock statistics are powerful, but they are not the only tools worth knowing. Userspace equivalents include Helgrind from the Valgrind suite and ThreadSanitizer (TSan), both of which instrument multithreaded applications for data races the same way KCSAN does for the kernel. Lockdep itself even has a userspace port for library developers who want the same deadlock-ordering proofs outside the kernel. On the tracing side, the modern eBPF-based deadlock-bpfcc tool from the BCC toolkit can attach to a running process or thread and flag potential lock order inversions without needing a special debug kernel at all — a handy option when you cannot rebuild the kernel on a production-like system.
Best Practices for Lock Debugging
- Keep a dedicated debug kernel build with
CONFIG_PROVE_LOCKINGandCONFIG_LOCK_STATfor development, separate from anything resembling production. - Reboot between heavy insmod/rmmod debugging sessions to avoid silent lockdep lock class exhaustion.
- Treat the first lockdep warning in dmesg as authoritative; ignore or disbelieve anything that appears after
debug_lockshas already flipped to zero. - Clear
/proc/lock_statbefore each measurement run so old numbers do not skew your analysis. - Correlate high
con-bounceswith your system’s actual CPU topology before assuming it is a software bug.
Common Mistakes and Troubleshooting
| Symptom | Likely Cause | Fix |
|---|---|---|
| Lockdep suddenly stops reporting anything | debug_locks was disabled by an earlier warning | Reboot and re-run the test from a clean state |
/proc/lock_stat does not exist | Kernel not built with CONFIG_LOCK_STAT | Rebuild your debug kernel with the option enabled |
| Lock class overflow warning during development | Repeated insmod/rmmod of the same module | Reboot instead of endlessly reloading |
| All lock stats show zero contention | You cleared stats but never re-enabled collection | Re-run the enable command before generating load |
Summary and Key Takeaways
- Lockdep allocates lock classes per module load and never frees them on unload, which can exhaust its internal table during heavy debugging sessions.
- The “lock debugging disabled” warning means lockdep already gave up earlier — trust the first warning in dmesg, not the latest one.
- KCSAN is the modern complement to lockdep, catching plain data races that lock-ordering analysis alone cannot see.
CONFIG_LOCK_STATexposes real per-lock timing data through/proc/lock_stat, letting you find genuinely contended locks instead of guessing.- con-bounces, waittime and holdtime together tell you whether a lock is a correctness problem, a performance problem, or both.
Conclusion
Lockdep and kernel lock statistics are two sides of the same debugging hook: one proves your locking is ordered correctly, the other tells you whether it is fast enough. Knowing the common linux kernel lockdep issues ahead of time — lock class exhaustion, the disabled-debugging warning — means you will spend less time debugging the debugger, and more time using /proc/lock_stat for what it is genuinely good at: finding the one contended lock that is quietly capping your driver’s throughput on a multi-core system.
Why does lockdep run out of lock classes during driver development?
Every module load allocates a fresh set of lock classes, and unloading the module does not free them — they are only marked reusable. Repeated insmod/rmmod cycles during a long debugging session can exhaust lockdep’s internal class table.
What does “lock debugging disabled” mean in dmesg?
It means the kernel’s internal debug_locks flag was already forced to zero by an earlier lockdep warning. Any lock debugging output after that point is suppressed, so you need to find the original warning further back in the log.
Do I need a special kernel build to see /proc/lock_stat?
Yes. You need CONFIG_LOCK_STAT=y at build time. Most stock distribution kernels ship without it, so a custom debug kernel is required.
What is the difference between contentions and con-bounces?
Contentions count how many times a task had to wait for an already-held lock. Con-bounces count the subset of those contentions that also caused the lock’s cacheline to migrate between CPU cores, which is more expensive.
Is KCSAN a replacement for lockdep?
No, they check different things. Lockdep proves lock acquisition ordering is deadlock-free. KCSAN detects plain unsynchronized memory accesses (data races) that may not involve any lock at all.
Can I get lock statistics without rebuilding the kernel?
Not through /proc/lock_stat specifically, since that requires CONFIG_LOCK_STAT at build time. For a production-like system you cannot rebuild, eBPF-based tracing tools are a more practical alternative.
What is a healthy holdtime-avg for a spinlock?
There is no universal number, but spinlocks are meant for very short critical sections — anything creeping into the hundreds of microseconds range is worth investigating, since spinning CPUs are wasting cycles the whole time.
More lectures on kernel synchronization, atomics, and lock debugging are on the way.
Previous Lecture Next Lecture
2 Comments