How to Read Lockdep Splats and Debug Kernel Deadlocks-Free Linux Device Drivers Course

How to Read Lockdep Splats and Debug Kernel Deadlocks
Lock debugging tutorial part 2: interpreting warnings, lock statistics, and KASAN/KCSAN, for kernel 6.12 and later
Level: Intermediate
Kernel: 6.12+
Reading Time: 15 min

In the previous lecture of this free Linux kernel development course, we configured a debug kernel with lockdep, lock statistics, and the core lock debugging Kconfig options. Turning those options on is only half the job. The real skill in kernel lock debugging is being able to read the warning lockdep prints, called a “splat,” and translate it into the actual bug in your driver or subsystem code.

This lecture, part of our free Linux device drivers course, breaks down the anatomy of a lockdep splat, walks through the most common deadlock patterns it detects, and shows how to combine lockdep with lock statistics, KASAN, and KCSAN for a complete debugging workflow.

Topics covered in this free embedded systems course lecture
Lockdep Splat Lock Statistics KASAN KCSAN Deadlock Patterns

What You Will Learn

  • How a lockdep warning is structured, section by section
  • The four classic deadlock and locking-misuse patterns lockdep detects
  • How to use CONFIG_LOCK_STAT to find real lock contention, not just correctness bugs
  • Where KASAN and KCSAN fit alongside lockdep in a modern debugging workflow
  • A repeatable step-by-step process for turning a splat into a fix

Prerequisites

This lecture assumes you have already built and booted a debug kernel with CONFIG_PROVE_LOCKING, CONFIG_LOCK_STAT, and the basic spinlock, mutex, and rwsem debug options enabled, as covered in the previous lecture of this free Linux kernel development course.

Anatomy of a Lockdep Warning

When lockdep detects a possible problem, it prints a multi-part block to the kernel log. Understanding each part quickly is what separates a five-minute fix from an hour of confused scrolling through dmesg.

Structure of a Lockdep Splat
Section What It Tells You
Header line Names the class of problem, such as a possible circular locking dependency, along with the kernel version and taint flags
Task context Process name and PID that triggered the detection, plus the exact lock it was trying to acquire
Existing dependency chain Shows the earlier code path that established the conflicting lock order, with its own stack trace
Held locks list Every lock the current task is holding right now, in acquisition order
Stack trace The call path that led to the problematic lock acquisition, which is usually where you start reading your own driver code

A practical habit: always start reading from the bottom of the stack trace section, since that is closest to your own driver’s entry point, and work upward to see how deep the call chain went before hitting the conflicting lock.

Common Deadlock Patterns Lockdep Detects

Four Locking Patterns to Recognize
AA deadlock
A task tries to acquire a lock it already holds
ABBA deadlock
Path 1 takes A then B, path 2 takes B then A
Interrupt-unsafe locking
A lock is taken in process context without disabling interrupts, but the same lock is also taken from an interrupt handler
Recursive read lock misuse
A reader-writer lock is used in a way that could starve a pending writer indefinitely

The interrupt-unsafe pattern deserves special attention for embedded and driver developers, since it is one of the most common bugs lockdep catches in real hardware drivers. If your driver’s interrupt handler and its normal I/O path both touch the same spinlock, the I/O path must disable local interrupts while holding that lock, or lockdep will flag it as a possible deadlock even if you never actually hit it during testing.

Using Lock Statistics to Find Contention, Not Just Bugs

CONFIG_LOCK_STAT answers a different question than lockdep: instead of “is this lock ordering safe,” it answers “how much time are tasks losing waiting for this lock.” This is invaluable when a driver is functionally correct but slow under load.

# Reset lock statistics before a test run
echo 0 > /proc/lock_stat

# Run your workload, then inspect the results
cat /proc/lock_stat | head -n 20

The output ranks locks by contention count and wait time, letting you spot the single hot lock that is serializing an otherwise parallel workload. This is often the next step after lockdep confirms your locking is correct but a profiler shows unexpected latency.

KASAN and KCSAN: Modern Companions to Lockdep

Lockdep proves locking order is sound, but it cannot tell you about memory corruption or genuine data races on unprotected memory. Two additional tools round out a modern kernel debugging setup:

  • KASAN (Kernel Address Sanitizer) uses compile-time instrumentation to catch out-of-bounds accesses and use-after-free bugs, which frequently corrupt lock structures themselves and produce confusing lockdep output as a side effect.
  • KCSAN (Kernel Concurrency Sanitizer) specifically targets data races on shared memory that is accessed without any lock at all, a category of bug lockdep cannot see since there is no lock acquisition to track in the first place.

Running lockdep, KASAN, and KCSAN together in a CI test kernel gives you overlapping coverage: lockdep for lock ordering, KASAN for memory safety, and KCSAN for unlocked shared-state races.

A Practical Lock Debugging Workflow

  1. Boot your lock-debug kernel and reproduce the workload that triggers the suspected issue
  2. Capture the full lockdep splat from dmesg, including both stack traces
  3. Identify which of the four deadlock patterns above matches the warning
  4. Trace the held-locks list back to the exact line in your driver where the conflicting lock is taken
  5. Fix the ordering, typically by taking locks in a consistent global order or by using a trylock with backoff where strict ordering is not possible
  6. Re-run the same workload and confirm the warning no longer appears
  7. If the driver is now correct but slow, switch to CONFIG_LOCK_STAT to check for contention

Troubleshooting Tips

  • Splat appears once and lockdep goes quiet afterward. This usually means lockdep hit an internal limit or turned itself off after detecting an inconsistency; reboot and capture the log again with more buffer space.
  • Warning references a lock class you do not recognize. Lock classes are named after the source line where the lock is initialized, not the variable name, so search your tree for that init call rather than the printed name.
  • Same warning appears in an unrelated subsystem. Third-party or older out-of-tree modules are common sources of lock ordering violations that only surface once your own driver is loaded alongside them.
  • Lock statistics show high contention but lockdep is silent. That is expected; contention is a performance issue, not a correctness issue, and the two tools answer different questions.

Real-World Use Cases

Kernel lock debugging is especially valuable in three situations common to embedded and driver engineering: bringing up a new interrupt-driven device driver where IRQ handlers and ioctl paths share state, porting a driver to a PREEMPT_RT kernel where locking semantics change, and investigating an intermittent system hang reported from the field that could not be reproduced under normal testing.

Best Practices

  • Always take locks in the same global order across every code path in a subsystem
  • Prefer a single well-documented locking hierarchy over ad-hoc locking added file by file
  • Run your driver test suite under a lock-debug kernel as a required CI step, not an optional one
  • Combine lockdep with KASAN and KCSAN in automated testing rather than relying on any single tool

Summary and Key Takeaways

  • A lockdep splat has a predictable structure: header, task context, prior dependency chain, held locks, and stack trace
  • AA, ABBA, interrupt-unsafe, and recursive read-lock patterns cover the majority of real-world lockdep warnings
  • CONFIG_LOCK_STAT complements lockdep by surfacing contention and performance issues rather than correctness bugs
  • KASAN and KCSAN close the gaps lockdep leaves open around memory safety and unlocked data races
  • A repeatable workflow, reproduce, capture, classify, trace, fix, and re-verify, turns lock debugging from guesswork into a routine engineering task

Conclusion

Lock debugging is one of the highest-leverage skills you can build as part of this free Linux kernel development course. A single afternoon spent learning to read a lockdep splat can save days of chasing an intermittent production hang later. Combine lockdep with lock statistics, KASAN, and KCSAN, and make lock-debug kernel testing a standard part of your driver development and free Linux device drivers course workflow rather than a last resort.

Frequently Asked Questions

Q1. Why did lockdep stop printing warnings after the first one?
Lockdep disables further checking once it detects certain internal inconsistencies, to avoid flooding the log with unreliable output. Fix the first warning and reboot to continue checking.

Q2. Is an ABBA warning always a guaranteed deadlock?
Not always in practice, since the exact timing needed may never occur, but it represents a real possibility and should be treated as a bug to fix, not ignored.

Q3. Can lockdep catch bugs in code that never actually deadlocked during testing?
Yes, that is its main strength. It proves an ordering is unsafe based on the lock classes it observed, regardless of whether the unsafe interleaving actually occurred.

Q4. What is the difference between lockdep and KCSAN?
Lockdep checks the ordering of lock acquisitions. KCSAN checks for data races on memory that may not be protected by any lock at all, which is a separate category of bug.

Q5. Do I need CONFIG_LOCK_STAT if I already have CONFIG_PROVE_LOCKING enabled?
They serve different purposes: PROVE_LOCKING finds correctness bugs, LOCK_STAT finds performance bottlenecks caused by lock contention. Enable both for a complete picture.

Q6. Should I run KASAN and lockdep together?
Yes, this is recommended for CI and development kernels, since memory corruption can sometimes trigger confusing lockdep warnings that are actually rooted in a memory safety bug.

Q7. How do I know which line in my driver to fix from a lockdep warning?
Read the stack trace section of the splat starting from the frame closest to your driver’s own function names, since that is where the problematic lock acquisition originates.

Q8. Does interrupt-unsafe locking only matter for interrupt handlers?
It matters for any code path that shares a lock with an interrupt handler, including timers and other asynchronous contexts, not only the primary IRQ handler function.

Continue Your Free Linux Kernel Development Course

Explore more lectures in this free embedded systems course and free Linux device drivers course series.

Next Lecture Back to Course Index

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *