What is Debugging Atomic Context Bugs: Lockdep and Modern Kernel Tracing- Best Linux Device Drivers Training In Hyderabad

Debugging Atomic Context Bugs: Lockdep and Modern Kernel Tracing

Free linux kernel development course — catching “sleeping while atomic” bugs on kernel 6.12+

Every experienced kernel developer eventually ships a bug where code that must never sleep ends up sleeping anyway. This lecture, part of our free linux kernel development course, shows you how to catch that class of bug before it reaches production, using the debugging tools available on a modern kernel rather than the older manual tracing techniques found in dated references. If you’re building driver skills for a free embedded systems course project, this is one of the highest-value debugging skills you can learn.

What You Will Learn
What “atomic context” really means Why CONFIG_DEBUG_ATOMIC_SLEEP matters Lockdep’s sleeping-while-atomic detector might_sleep() and its diagnostic role Modern tracing with ftrace and bpftrace

Prerequisites

This lecture builds directly on the previous module of our free Linux device drivers course covering spinlocks and interrupt context. You should already know the difference between a sleeping lock (mutex) and a non-sleeping lock (spinlock), and ideally have access to a kernel build environment on 6.6 LTS or newer.

What Counts as Atomic Context

“Atomic context” is any point in kernel execution where the code is not allowed to give up the CPU voluntarily. This includes:

  • Hardware interrupt handlers (hardirq context)
  • Softirq and tasklet handlers
  • Any region between spin_lock() and spin_unlock()
  • Any region where preemption has been explicitly disabled with preempt_disable()

Calling a function that can block — a mutex acquisition, a memory allocation with GFP_KERNEL, msleep(), or a blocking I/O wait — from inside any of these regions is a bug, commonly called “sleeping in atomic context” or a scheduling-while-atomic bug.

Why This Bug Is So Easy to Miss in a free Linux Kernel Development Course Project

The dangerous part of this bug class is that it frequently does nothing observable most of the time. A rarely-taken error branch inside a locked region might call kfree() then unexpectedly touch a path that allocates memory, or a driver might call a helper function that only sleeps under specific hardware timing conditions. On a production kernel built without extra safety checks, the corruption or stall this causes can appear intermittent, occur far from the real cause, and be nearly impossible to reproduce on demand — which is exactly why a debug-enabled kernel is non-negotiable during development.

CONFIG_DEBUG_ATOMIC_SLEEP and Lockdep

Modern kernels bundle the atomic-sleep checker into the lock-dependency validator, Lockdep, rather than treating it as a standalone feature. Enabling the relevant debug options gives you an immediate, loud kernel warning — a full stack trace printed to the kernel log — the instant a sleeping call happens inside atomic context, rather than a silent corruption you’d have to hunt for later.

How the Sleeping-While-Atomic Check Fires
spin_lock() taken
preempt_count++
→
code calls a
blocking function
→
scheduler checks
preempt_count > 0
→
“BUG: sleeping
function called”

The kernel tracks a preempt_count for every task. Any code path that can sleep checks this counter first via might_sleep(); a non-zero count means “you are inside a locked or otherwise atomic region,” and the warning fires immediately with a full backtrace.

Building and Booting a Debug Kernel

To get these checks, you need a kernel configured for debugging. On a distribution kernel this is usually off by default for performance reasons.

$ make menuconfig
# Enable under Kernel hacking -> Lock Debugging:
#   CONFIG_PROVE_LOCKING=y
#   CONFIG_DEBUG_ATOMIC_SLEEP=y
#   CONFIG_DEBUG_SPINLOCK=y

$ make -j$(nproc)
$ sudo make modules_install install
$ sudo reboot

After rebooting into the debug kernel, confirm the option is active:

$ grep DEBUG_ATOMIC_SLEEP /boot/config-$(uname -r)
CONFIG_DEBUG_ATOMIC_SLEEP=y

might_sleep(): Your Own Early-Warning Marker

You don’t have to wait for a real bug to test whether your locking is safe. Kernel APIs that can block call might_sleep() internally, and you can also call it yourself in helper functions to document and enforce a “this must never be called from atomic context” contract:

#include <linux/kernel.h>

/* This helper must only ever be called from process context */
static void mydrv_slow_helper(struct mydrv_data *data)
{
    might_sleep();  /* triggers a warning if called under a spinlock */

    mutex_lock(&data->config_mutex);
    /* ... work that may legitimately block ... */
    mutex_unlock(&data->config_mutex);
}

Tracing a Live Bug with Ftrace and Bpftrace

When a warning fires, the printed backtrace usually tells you exactly which function slept and where. For the rarer case where the warning is too generic or the call happens deep inside a shared helper, function-graph tracing narrows it down further:

$ cd /sys/kernel/tracing
$ echo function_graph > current_tracer
$ echo mydrv_isr > set_graph_function
$ echo 1 > tracing_on
# reproduce the issue
$ cat trace | less

On a modern system, bpftrace is often faster to reach for than raw ftrace files, since you can target the exact function and print only what you need:

$ sudo bpftrace -e '
kprobe:mydrv_isr { @start[tid] = nsecs; }
kretprobe:mydrv_isr /@start[tid]/ {
    printf("isr duration: %d ns\n", nsecs - @start[tid]);
    delete(@start[tid]);
}'

This gives you interrupt-handler timing without recompiling anything, which is invaluable when you suspect a handler is taking long enough to interact badly with locking elsewhere in the driver.

Debug Options Comparison

Tool / Option Catches Overhead
CONFIG_DEBUG_ATOMIC_SLEEP Sleeping calls inside atomic context Low
CONFIG_PROVE_LOCKING (Lockdep) Lock ordering violations, potential deadlocks Moderate to high
ftrace function_graph Exact call sequence and timing Moderate (target specific functions)
bpftrace Ad-hoc, targeted probes without rebuilding Low, scoped to your probes

Real-World Use Cases

  • Catching an accidental kmalloc(..., GFP_KERNEL) call inside a spinlock-protected error path.
  • Finding a mutex acquisition buried inside a helper function that a driver calls from its interrupt handler.
  • Diagnosing intermittent system stalls that only reproduce under specific timing on production-like hardware.

Common Mistakes and Troubleshooting

Mistake: Developing exclusively on a distro kernel with debug options disabled. Bugs will pass local testing and surface later in the field.
Mistake: Assuming a rarely-hit error branch is “probably fine” without auditing it for blocking calls under a lock.
Tip: Keep a permanent debug-kernel VM in your development workflow and run every new driver code path through it at least once before merging.

Best Practices

  1. Always develop and test on a kernel with CONFIG_DEBUG_ATOMIC_SLEEP and CONFIG_PROVE_LOCKING enabled.
  2. Add might_sleep() to your own helper functions that assume process context.
  3. Audit every function called from inside a locked region, including nested helper calls.
  4. Use targeted tracing (bpftrace/ftrace) rather than broad system-wide tracing to keep overhead manageable.

Performance Considerations

Lock-debugging options carry real runtime overhead and should never ship in a production kernel. Keep a dedicated debug-kernel build for development and testing, and switch to a release configuration for anything performance-sensitive or shipped to end users.

Security Considerations

An unnoticed sleeping-in-atomic bug can turn into a system-wide hang, which is a denial-of-service vector on any system where an attacker can trigger the vulnerable code path. Treat every Lockdep or atomic-sleep warning as a must-fix issue, not a cosmetic one.

Summary and Key Takeaways

  • Atomic context covers interrupt handlers, softirqs, and any spinlock-protected region.
  • CONFIG_DEBUG_ATOMIC_SLEEP and Lockdep turn silent corruption into a loud, immediate warning.
  • might_sleep() lets you enforce your own atomic-context contracts.
  • bpftrace and ftrace’s function_graph tracer are the modern go-to tools for pinpointing exact call timing.

Conclusion

Atomic-context bugs are quiet by nature, which is precisely why they need loud tools. Building the habit of developing on a debug-enabled kernel, and knowing how to reach for Lockdep, might_sleep(), and bpftrace the moment something looks off, will save you hours of guesswork on real hardware. That habit, more than any single API, is what separates a driver author who merely writes code that boots from one who writes code that survives contact with production — a distinction this free Linux kernel development course aims to build in every module.

Frequently Asked Questions

Q1. What exactly triggers a “sleeping while atomic” warning?
Any call to a function that can voluntarily give up the CPU — a mutex_lock(), a GFP_KERNEL allocation, msleep(), and similar — while preempt_count is non-zero, which happens inside spinlock-protected regions or interrupt handlers.

Q2. Is CONFIG_DEBUG_ATOMIC_SLEEP safe to enable on a production system?
It is not recommended. It adds runtime checking overhead and is intended for development and testing kernels, not shipping production builds.

Q3. What’s the difference between Lockdep and CONFIG_DEBUG_ATOMIC_SLEEP?
Lockdep (CONFIG_PROVE_LOCKING) mainly tracks lock acquisition order to catch potential deadlocks, while CONFIG_DEBUG_ATOMIC_SLEEP specifically flags a blocking call happening inside atomic context. Modern kernels use them together for the fullest coverage.

Q4. Can I use bpftrace on a stock distro kernel without a custom build?
Yes, bpftrace only needs BPF and kprobe support, which is enabled by default on virtually all modern distribution kernels, unlike Lockdep-based checks which need a debug config.

Q5. Why does my driver work fine on my dev machine but hang on the target board?
This is a classic symptom of an atomic-context bug that only reproduces under specific timing. Reproduce it on a debug-enabled kernel with CONFIG_DEBUG_ATOMIC_SLEEP to get an immediate stack trace instead of guessing.

Q6. Does the RCU or PREEMPT_RT model change how these bugs manifest?
The underlying rule — no sleeping in atomic context — stays the same, but on PREEMPT_RT some spinlocks behave as sleeping locks internally, which changes which specific call sites can legally block. The debug options above remain the right way to verify this on any kernel configuration.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *