Debugging Atomic Context Bugs: Lockdep and Modern Kernel Tracing
Free linux kernel development course — catching “sleeping while atomic” bugs on kernel 6.12+
Every experienced kernel developer eventually ships a bug where code that must never sleep ends up sleeping anyway. This lecture, part of our free linux kernel development course, shows you how to catch that class of bug before it reaches production, using the debugging tools available on a modern kernel rather than the older manual tracing techniques found in dated references. If you’re building driver skills for a free embedded systems course project, this is one of the highest-value debugging skills you can learn.
Prerequisites
This lecture builds directly on the previous module of our free Linux device drivers course covering spinlocks and interrupt context. You should already know the difference between a sleeping lock (mutex) and a non-sleeping lock (spinlock), and ideally have access to a kernel build environment on 6.6 LTS or newer.
What Counts as Atomic Context
“Atomic context” is any point in kernel execution where the code is not allowed to give up the CPU voluntarily. This includes:
- Hardware interrupt handlers (hardirq context)
- Softirq and tasklet handlers
- Any region between
spin_lock()andspin_unlock() - Any region where preemption has been explicitly disabled with
preempt_disable()
Calling a function that can block — a mutex acquisition, a memory allocation with GFP_KERNEL, msleep(), or a blocking I/O wait — from inside any of these regions is a bug, commonly called “sleeping in atomic context” or a scheduling-while-atomic bug.
Why This Bug Is So Easy to Miss in a free Linux Kernel Development Course Project
The dangerous part of this bug class is that it frequently does nothing observable most of the time. A rarely-taken error branch inside a locked region might call kfree() then unexpectedly touch a path that allocates memory, or a driver might call a helper function that only sleeps under specific hardware timing conditions. On a production kernel built without extra safety checks, the corruption or stall this causes can appear intermittent, occur far from the real cause, and be nearly impossible to reproduce on demand — which is exactly why a debug-enabled kernel is non-negotiable during development.
CONFIG_DEBUG_ATOMIC_SLEEP and Lockdep
Modern kernels bundle the atomic-sleep checker into the lock-dependency validator, Lockdep, rather than treating it as a standalone feature. Enabling the relevant debug options gives you an immediate, loud kernel warning — a full stack trace printed to the kernel log — the instant a sleeping call happens inside atomic context, rather than a silent corruption you’d have to hunt for later.
preempt_count++
blocking function
preempt_count > 0
function called”
The kernel tracks a preempt_count for every task. Any code path that can sleep checks this counter first via might_sleep(); a non-zero count means “you are inside a locked or otherwise atomic region,” and the warning fires immediately with a full backtrace.
Building and Booting a Debug Kernel
To get these checks, you need a kernel configured for debugging. On a distribution kernel this is usually off by default for performance reasons.
$ make menuconfig
# Enable under Kernel hacking -> Lock Debugging:
# CONFIG_PROVE_LOCKING=y
# CONFIG_DEBUG_ATOMIC_SLEEP=y
# CONFIG_DEBUG_SPINLOCK=y
$ make -j$(nproc)
$ sudo make modules_install install
$ sudo reboot
After rebooting into the debug kernel, confirm the option is active:
$ grep DEBUG_ATOMIC_SLEEP /boot/config-$(uname -r)
CONFIG_DEBUG_ATOMIC_SLEEP=y
might_sleep(): Your Own Early-Warning Marker
You don’t have to wait for a real bug to test whether your locking is safe. Kernel APIs that can block call might_sleep() internally, and you can also call it yourself in helper functions to document and enforce a “this must never be called from atomic context” contract:
#include <linux/kernel.h>
/* This helper must only ever be called from process context */
static void mydrv_slow_helper(struct mydrv_data *data)
{
might_sleep(); /* triggers a warning if called under a spinlock */
mutex_lock(&data->config_mutex);
/* ... work that may legitimately block ... */
mutex_unlock(&data->config_mutex);
}
Tracing a Live Bug with Ftrace and Bpftrace
When a warning fires, the printed backtrace usually tells you exactly which function slept and where. For the rarer case where the warning is too generic or the call happens deep inside a shared helper, function-graph tracing narrows it down further:
$ cd /sys/kernel/tracing
$ echo function_graph > current_tracer
$ echo mydrv_isr > set_graph_function
$ echo 1 > tracing_on
# reproduce the issue
$ cat trace | less
On a modern system, bpftrace is often faster to reach for than raw ftrace files, since you can target the exact function and print only what you need:
$ sudo bpftrace -e '
kprobe:mydrv_isr { @start[tid] = nsecs; }
kretprobe:mydrv_isr /@start[tid]/ {
printf("isr duration: %d ns\n", nsecs - @start[tid]);
delete(@start[tid]);
}'
This gives you interrupt-handler timing without recompiling anything, which is invaluable when you suspect a handler is taking long enough to interact badly with locking elsewhere in the driver.
Debug Options Comparison
| Tool / Option | Catches | Overhead |
|---|---|---|
| CONFIG_DEBUG_ATOMIC_SLEEP | Sleeping calls inside atomic context | Low |
| CONFIG_PROVE_LOCKING (Lockdep) | Lock ordering violations, potential deadlocks | Moderate to high |
| ftrace function_graph | Exact call sequence and timing | Moderate (target specific functions) |
| bpftrace | Ad-hoc, targeted probes without rebuilding | Low, scoped to your probes |
Real-World Use Cases
- Catching an accidental
kmalloc(..., GFP_KERNEL)call inside a spinlock-protected error path. - Finding a mutex acquisition buried inside a helper function that a driver calls from its interrupt handler.
- Diagnosing intermittent system stalls that only reproduce under specific timing on production-like hardware.
Common Mistakes and Troubleshooting
Best Practices
- Always develop and test on a kernel with CONFIG_DEBUG_ATOMIC_SLEEP and CONFIG_PROVE_LOCKING enabled.
- Add
might_sleep()to your own helper functions that assume process context. - Audit every function called from inside a locked region, including nested helper calls.
- Use targeted tracing (bpftrace/ftrace) rather than broad system-wide tracing to keep overhead manageable.
Performance Considerations
Lock-debugging options carry real runtime overhead and should never ship in a production kernel. Keep a dedicated debug-kernel build for development and testing, and switch to a release configuration for anything performance-sensitive or shipped to end users.
Security Considerations
An unnoticed sleeping-in-atomic bug can turn into a system-wide hang, which is a denial-of-service vector on any system where an attacker can trigger the vulnerable code path. Treat every Lockdep or atomic-sleep warning as a must-fix issue, not a cosmetic one.
Summary and Key Takeaways
- Atomic context covers interrupt handlers, softirqs, and any spinlock-protected region.
- CONFIG_DEBUG_ATOMIC_SLEEP and Lockdep turn silent corruption into a loud, immediate warning.
might_sleep()lets you enforce your own atomic-context contracts.- bpftrace and ftrace’s function_graph tracer are the modern go-to tools for pinpointing exact call timing.
Conclusion
Atomic-context bugs are quiet by nature, which is precisely why they need loud tools. Building the habit of developing on a debug-enabled kernel, and knowing how to reach for Lockdep, might_sleep(), and bpftrace the moment something looks off, will save you hours of guesswork on real hardware. That habit, more than any single API, is what separates a driver author who merely writes code that boots from one who writes code that survives contact with production — a distinction this free Linux kernel development course aims to build in every module.
Frequently Asked Questions
Q1. What exactly triggers a “sleeping while atomic” warning?
Any call to a function that can voluntarily give up the CPU — a mutex_lock(), a GFP_KERNEL allocation, msleep(), and similar — while preempt_count is non-zero, which happens inside spinlock-protected regions or interrupt handlers.
Q2. Is CONFIG_DEBUG_ATOMIC_SLEEP safe to enable on a production system?
It is not recommended. It adds runtime checking overhead and is intended for development and testing kernels, not shipping production builds.
Q3. What’s the difference between Lockdep and CONFIG_DEBUG_ATOMIC_SLEEP?
Lockdep (CONFIG_PROVE_LOCKING) mainly tracks lock acquisition order to catch potential deadlocks, while CONFIG_DEBUG_ATOMIC_SLEEP specifically flags a blocking call happening inside atomic context. Modern kernels use them together for the fullest coverage.
Q4. Can I use bpftrace on a stock distro kernel without a custom build?
Yes, bpftrace only needs BPF and kprobe support, which is enabled by default on virtually all modern distribution kernels, unlike Lockdep-based checks which need a debug config.
Q5. Why does my driver work fine on my dev machine but hang on the target board?
This is a classic symptom of an atomic-context bug that only reproduces under specific timing. Reproduce it on a debug-enabled kernel with CONFIG_DEBUG_ATOMIC_SLEEP to get an immediate stack trace instead of guessing.
Q6. Does the RCU or PREEMPT_RT model change how these bugs manifest?
The underlying rule — no sleeping in atomic context — stays the same, but on PREEMPT_RT some spinlocks behave as sleeping locks internally, which changes which specific call sites can legally block. The debug options above remain the right way to verify this on any kernel configuration.

2 Comments