← Previous Lecture Next Lecture →
If you have ever hit a scheduling while atomic call trace in your kernel log and felt lost staring at fifty lines of hexadecimal offsets, this lecture is for you. In the previous lecture of this free Linux kernel development course we built a driver that deliberately slept inside a spinlock. Now we go one step further: we actually trigger that bug, capture the kernel’s diagnostic dump, and learn how to read it line by line – on a modern 6.x kernel. Just as important, we will repeat the exact same experiment on a kernel that does not have the debug checks enabled, and see why that difference should worry every driver author.
What You Will Learn
- Decoding the “BUG: scheduling while atomic” header line
- Reading a kernel call trace from the bottom up
- Why some stack frames are prefixed with a “?”
- What CONFIG_DEBUG_ATOMIC_SLEEP actually checks
- Why the identical bug can be completely silent on a production kernel
- A safe, original demo driver you can build and trace yourself
Prerequisites
This lecture assumes you have already gone through the earlier lectures in this chapter, in particular the ones covering spinlocks, mutexes, and the difference between process context and atomic context. You should be comfortable building and loading a kernel module, and you should always run experiments like this one inside a disposable virtual machine, never on real hardware, since a genuine scheduling-while-atomic condition can wedge the whole system.
A One-Line Recap
A spinlock critical section must never call anything that can block, because a spinning CPU cannot be told to go away and come back later – it is busy-waiting with preemption disabled. Calling a sleeping API such as schedule_timeout(), mutex_lock(), or even a blocking memory allocation from inside that window is exactly the mistake the kernel’s debug instrumentation is built to catch.
Anatomy of the Bug Header Line
When the debug checks fire, the very first line printed to the kernel log looks something like this:
BUG: scheduling while atomic: worker_app/4210/0x00000002
Every field in that line carries information, and none of it is decorative:
The task name and PID tell you who was running when the violation happened – usually the process or kernel thread whose code path took the lock. The last field, the preempt count, tells you how deep you were nested. A value of 0x00000002 generally means one spinlock was held (each spinlock acquisition adds 1 to the preempt count internally, and IRQ-disabling variants add more). If you ever see this number climb into double digits, it is a strong hint that locks are being taken recursively somewhere they should not be.
Reading the Call Trace: Always Bottom-Up
Right after the header line, the kernel prints the call trace – the sequence of function calls that were active on the kernel stack at the exact moment the bug fired. The single most important habit to build is this: read a kernel call trace from the bottom toward the top. The bottom-most frame is the oldest call, closest to where the system call entered the kernel; the top-most frame is the innermost function, the one actually executing when the trace was captured.
Reading bottom-up, the story reconstructs itself naturally: a user-space write() call entered the kernel through the standard system call gate, was routed through the generic VFS layer, reached your driver’s write handler, and from inside that handler the code eventually called something that put the current task to sleep. The topmost frames are your smoking gun – they show exactly which sleeping API was called and from which function.
Why Some Frames Start With a “?”
Scroll through a real trace and you will notice many lines prefixed with a question mark. Those are frames the unwinder considered questionable – leftover addresses sitting in the same region of stack memory from an earlier, unrelated function call. This is simply how stack memory works: the kernel does not zero out a stack page every time a function returns, because that would be wasteful. As functions are called and return, the same physical stack memory gets reused over and over, and old byte patterns can still resemble valid return addresses even though they no longer are. Modern frame-pointer and ORC unwinders used in kernel 6.x are quite good at filtering these out, but a few can still slip through. The practical rule is simple: focus on the lines without a question mark – those are the frames the unwinder is confident about, and they are almost always enough to find the culprit function.
What Actually Detects the Bug: CONFIG_DEBUG_ATOMIC_SLEEP
None of this diagnostic output appears by magic – it is produced by a kernel debug feature called CONFIG_DEBUG_ATOMIC_SLEEP. When this option is enabled in your kernel configuration, every function that is allowed to sleep begins with an internal check (functionally, a call to might_sleep()) that inspects the current preempt count. If that count is non-zero – meaning some lock is held or preemption has been explicitly disabled – the kernel immediately prints the bug header and call trace you just learned to read, then normally lets execution continue in a degraded, best-effort way so you can still capture the log.
This check is a build-time and boot-time decision, not something a running production kernel does automatically. On kernel 6.x you will find this option under the Kernel Hacking menu of make menuconfig, and it is bundled with debug-focused kernel builds. A stock distribution kernel almost never ships with it turned on, because the extra check on every scheduling point costs a small amount of performance.
The Uncomfortable Part: The Bug Does Not Disappear Without Debug Checks
Here is the part every driver author needs to internalize. If you take the exact same buggy driver and load it on a generic distribution kernel that was not built with CONFIG_DEBUG_ATOMIC_SLEEP, the bug does not go away – it simply stops being reported. The write still sleeps inside the locked region. The kernel log stays completely clean. From the outside, everything looks fine.
| Behaviour | Debug Kernel (CONFIG_DEBUG_ATOMIC_SLEEP=y) | Generic Distro Kernel |
|---|---|---|
| Kernel log output | Full “BUG: scheduling while atomic” header + call trace | Nothing at all |
| Immediate visible effect | Warning only, task usually still completes | May appear to “just work” |
| Real risk under load | Caught early, in development | Latent risk of stalls, missed deadlines, or a hang under heavier scheduling pressure |
| When you find out | During testing, on a VM | Possibly in production, under load, much harder to reproduce |
This is exactly why the earlier lecture insisted on testing this kind of experiment on a VM running a debug kernel. The absence of a warning is not proof of correctness; it is only proof that nobody happened to be watching closely enough at build time.
A Small Original Demo: Logging Preempt Depth Yourself
Rather than reproducing anyone else’s driver, here is a small, original helper you can drop into your own kernel 6.x module to make the preempt count visible on demand, so you can sanity-check a critical section before you ever trust it in production:
#include <linux/kernel.h>
#include <linux/preempt.h>
#include <linux/spinlock.h>
static DEFINE_SPINLOCK(ep_demo_lock);
static void ep_report_atomic_depth(const char *tag)
{
/* preempt_count() reflects locks held + irq/softirq disables */
pr_info("ep_demo: [%s] preempt_count=%d (in_atomic=%d)\n",
tag, preempt_count(), in_atomic());
}
static void ep_safe_critical_section(void)
{
unsigned long flags;
ep_report_atomic_depth("before lock");
spin_lock_irqsave(&ep_demo_lock, flags);
ep_report_atomic_depth("inside lock"); /* preempt_count() > 0 here */
/* Do only fast, non-blocking work in here: no sleeping APIs allowed */
spin_unlock_irqrestore(&ep_demo_lock, flags);
ep_report_atomic_depth("after unlock"); /* back to preempt_count() == 0 */
}
Loading a module that calls ep_safe_critical_section() and watching dmesg will show the preempt count rise the moment the spinlock is taken and fall back to zero the moment it is released. If you ever see it non-zero somewhere it should not be, that is your early warning sign, well before any bug report is ever printed.
Best Practices Going Forward
- Always develop and test kernel modules on a debug kernel with CONFIG_DEBUG_ATOMIC_SLEEP enabled
- Never assume a clean dmesg on a distro kernel means your locking is correct
- Read every call trace bottom-up, and trust the frames without a leading “?”
- Keep spinlock critical sections short and free of any sleeping API
- Use lockdep alongside CONFIG_DEBUG_ATOMIC_SLEEP for deeper locking validation
Frequently Asked Questions
Q1. What does “scheduling while atomic” actually mean?
It means the kernel tried to put the current task to sleep while it was inside a section where sleeping is forbidden, such as while holding a spinlock or with preemption explicitly disabled.
Q2. Is preempt_count() the same as the number of locks held?
Roughly, yes for spinlocks – each acquisition increments it – but IRQ-disabling calls and softirq/hardirq context also contribute to the same counter.
Q3. Why did I not see this bug on my desktop distro kernel?
Because CONFIG_DEBUG_ATOMIC_SLEEP is a debug option that most distribution kernels do not enable by default, so the check simply never runs.
Q4. Should I enable CONFIG_DEBUG_ATOMIC_SLEEP on a production system?
Generally no – use it on your development and testing kernels; production kernels favour the small performance gain of skipping the extra checks.
Q5. Why are some call trace lines marked with a “?”
They are frames the stack unwinder is not fully confident about, usually stale leftover bytes from a previous, unrelated function call that reused the same stack memory.
Q6. What is the safest way to test this kind of bug?
Always inside a disposable virtual machine, since a genuine scheduling-while-atomic condition can potentially hang the entire system.
Q7. Does this apply to spinlocks only?
No – the same check fires for any atomic context, including hardirq handlers, softirq handlers, and any code region where preemption has been explicitly disabled.
Continue this free Linux kernel development course and go deeper into kernel synchronization primitives.
Next Lecture → ← Previous Lecture
2 Comments