« Previous Lecture | Next Lecture »
If you have ever pasted a wall of kernel log text into a search engine because a driver hung after loading, you already know that a lockdep report looks intimidating the first time you see it. It is not. Once you know what each line means, a lockdep splat becomes one of the fastest ways to find a locking bug in a Linux kernel module — often faster than adding print statements. In this lecture we take a self-deadlock warning apart piece by piece, on a modern kernel 6.x system, and show exactly how to turn that warning into a one-line code fix.
You should already be comfortable with basic spinlock usage (spin_lock() / spin_unlock()) and have lockdep enabled on your build (CONFIG_PROVE_LOCKING=y), which we covered in the previous lecture. If you have not built a lockdep-enabled kernel config yet, go back one lecture before continuing here.
The Scenario That Produces the Report
Lockdep prints a self-deadlock warning whenever a single thread of execution tries to acquire a lock it is already holding — without going through an unlock first. This most commonly happens indirectly: your code takes a lock, and then somewhere deeper in the call chain it calls a helper function that, unknown to you, tries to take that exact same lock again. On kernel builds without lockdep this simply hangs forever with no explanation. On a lockdep-enabled kernel, you get a full diagnostic report the moment the second acquire is attempted — before the deadlock actually happens.
Anatomy of the Warning — Line by Line
+ kernel version and taint flags
▶ “<task>/<pid> is trying to acquire lock:”
+ lock address, lock class name, calling function+offset
▶ “but task is already holding lock:”
+ SAME lock address and class name, different calling function
▶ “other info that might help us debug this:”
+ a small CPU/lock diagram showing the same lock taken twice
▶ “1 lock held by <task>/<pid>:” + backtrace
Read it top to bottom as a story: lockdep first tells you what the code was about to do (“is trying to acquire”), then tells you what the code had already done earlier in the same call path (“already holding”). Everything else in the report exists to help you connect those two moments.
The Lock Identity Hash — How You Know It’s the Same Lock
Each of the two lock lines begins with a long hexadecimal number. That number is a hash lockdep computes to uniquely identify that specific lock instance in memory. If the hash on the “trying to acquire” line matches the hash on the “already holding” line, you have proof — not a guess — that both acquire attempts refer to the exact same lock object. This is the single most useful fact in the whole report: it immediately rules out “maybe these are two different locks that just look similar,” which is normally the hardest thing to confirm by manual code reading alone.
Understanding the Lock State Notation
Next to the lock class name you will see a short bracketed code made up of symbols such as plus signs and dots. This is lockdep’s compact record of the interrupt state the lock has previously been acquired under — for example, whether it has ever been taken with hardware interrupts enabled, with interrupts disabled, or from softirq context. Lockdep tracks this because a lock that is safe when only ever taken from process context can become unsafe the moment it is also taken from an interrupt handler, opening the door to a completely different class of deadlock that we cover later in this chapter. For a self-deadlock report specifically, this notation is secondary — the two matching hashes and the “already holding” line are what matter first.
Finding the Real Bug in Your Own Code
Once you know two acquire attempts share a lock, the fix is almost always the same shape: somewhere in your call chain, a helper function takes a lock that the caller already holds. Here is a minimal, original example of the pattern — a device driver context structure protected by a spinlock, and a status-formatting helper that mistakenly re-takes it:
struct ep_dev_ctx {
spinlock_t lock;
int status_code;
char status_text[32];
};
/* BUG: this helper takes the same lock its only caller
* already holds, producing a textbook self-deadlock. */
static void ep_format_status(struct ep_dev_ctx *ctx, char *out, size_t len)
{
unsigned long flags;
spin_lock_irqsave(&ctx->lock, flags);
snprintf(out, len, "code=%d text=%s",
ctx->status_code, ctx->status_text);
spin_unlock_irqrestore(&ctx->lock, flags);
}
static void ep_dump_status(struct ep_dev_ctx *ctx)
{
unsigned long flags;
char buf[64];
spin_lock_irqsave(&ctx->lock, flags); /* lock taken here */
ctx->status_code = 1;
ep_format_status(ctx, buf, sizeof(buf)); /* tries to take it again! */
spin_unlock_irqrestore(&ctx->lock, flags);
}
The Fix: Unlock, Call, Relock
The correct fix is not to remove locking from either function — both genuinely need to protect their access to ctx. The fix is to release the lock immediately before calling the helper, and take it again right after, so the two critical sections never overlap:
static void ep_dump_status_fixed(struct ep_dev_ctx *ctx)
{
unsigned long flags;
char buf[64];
spin_lock_irqsave(&ctx->lock, flags);
ctx->status_code = 1;
spin_unlock_irqrestore(&ctx->lock, flags); /* release first */
ep_format_status(ctx, buf, sizeof(buf)); /* helper locks internally */
spin_lock_irqsave(&ctx->lock, flags); /* reacquire */
/* continue with any remaining protected work */
spin_unlock_irqrestore(&ctx->lock, flags);
}
This is exactly the pattern lockdep is nudging you toward whenever it reports recursive locking: shrink each critical section so that a lock is never held across a call into code that might take it again. On kernel 6.x you can immediately confirm the fix by reloading the module and checking that the “possible recursive locking detected” warning no longer appears in dmesg.
lock(ctx->lock) lock(ctx->lock)
call helper() …work…
lock(ctx->lock) ❌ unlock(ctx->lock)
DEADLOCK call helper() ✅
unlock(ctx->lock) lock(ctx->lock)
unlock(ctx->lock)
Looking Ahead: When Two Locks Are Involved
A single-lock self-deadlock is the simplest case lockdep detects. In the next lecture we move to a more common real-world scenario: two separate locks acquired by two different threads in opposite order, known as an AB-BA deadlock. Lockdep catches this too, and the report format you just learned to read is the exact same format it uses there — just with a second lock’s identity hash added to the picture.
Does lockdep detect the deadlock before or after it actually hangs the system?
Before. Lockdep tracks lock acquisition order continuously and prints the warning at the moment a dangerous pattern is attempted, so in most self-deadlock cases you get the report and the system may still recover or at worst hang only that one thread, rather than the whole machine silently freezing with no clue why.
Why does the same lock show two different calling functions in the report?
Because the lock was acquired twice from two different places in your call chain — once by the outer function and once by the inner helper it called. Lockdep prints the function+offset for both acquisitions so you can locate each one in your source.
Can I ignore a lockdep warning if the module seems to still work?
No. A self-deadlock warning describes a real bug that will freeze that code path under the right timing or workload, even if your specific test run happened not to trigger the full hang. Treat every lockdep warning as a bug report, not a suggestion.
Is the unlock-then-relock fix always safe?
It is safe as long as you re-check any assumptions about shared state after reacquiring the lock, since another thread could have modified that state during the brief window the lock was released. For the simple status-formatting case shown here that is not a concern, but for more complex state machines you should re-validate conditions after the relock.
Does this apply to mutexes as well as spinlocks?
Yes. Lockdep tracks mutexes, spinlocks, rwlocks, and semaphores through the same underlying lock-class engine, so a recursive-locking warning can appear for any of them, and the unlock-then-call-then-relock fix pattern applies equally.
What kernel config do I need to see these reports?
You need CONFIG_PROVE_LOCKING=y (which also pulls in the required lock-debugging infrastructure) enabled in your running kernel, as covered in the previous lecture on setting up a debug kernel build.
Continue the Free Linux Kernel Programming Course
More hands-on lectures on kernel synchronization, device drivers, and debugging — all free.
Next Lecture Course Index
2 Comments