Catching the Bug That Doesn’t Show Up: Ftrace & trace-cmd Debugging-Free Linux Device Driver Course Online

← Previous Lecture    Next Lecture →

Catching the Bug That Doesn’t Show Up: Ftrace & trace-cmd Debugging
Free Linux Kernel Development Course — Kernel Synchronization Series (Kernel 6.x)
Level: Intermediate
Reading Time: 14 min
Kernel: 6.x

Welcome back to our free linux kernel development course. One of the scariest lessons in kernel programming is this: a driver bug can be present, dangerous, and completely silent on your machine — and then crash a customer’s device in production. In this lecture we use ftrace trace-cmd kernel debugging to expose a bug that a normal kernel build will never warn you about, and we explain exactly why the warning disappears.

ftrace trace-cmd function_graph tracer CONFIG_DEBUG_ATOMIC_SLEEP scheduling while atomic free linux kernel development course

Prerequisites

  • Comfortable building and loading kernel modules on a 6.x kernel
  • Understanding of atomic context vs. process context (covered in earlier lectures)
  • A test VM where you don’t mind installing a debug kernel

What You Will Learn

  • Why a genuinely buggy driver can run for months without any visible error
  • How the Ftrace function_graph tracer, driven through trace-cmd, lets you watch kernel function calls in order
  • How to spot a forbidden schedule() call happening inside a locked, non-blocking section
  • What CONFIG_DEBUG_ATOMIC_SLEEP does, and why most distribution kernels ship with it switched off
  • How the modern rtla latency-analysis tool complements Ftrace on kernel 6.x

The Trap: “It Works On My Machine”

Imagine a character-device driver whose write handler grabs a spinlock and then, somewhere deep in a helper function, ends up calling something that can sleep — say, a delay routine that internally calls schedule_timeout(). Sleeping while holding a spinlock is a serious bug: a spinlock protects a critical section by spinning, not sleeping, and the CPU that owns it must never give up the CPU voluntarily until it unlocks.

Here’s the unsettling part. On an ordinary distribution kernel, that broken driver can be loaded, exercised repeatedly from a user-space test application, and never print a single warning. Everything looks fine. The only way to know something is wrong is to watch the kernel’s own function calls as they happen — which is exactly what Ftrace is built for.

A Quick Tour of Ftrace and trace-cmd

Ftrace is the tracing engine built into the Linux kernel itself; it can record every function entered and exited in kernel space. Driving Ftrace directly through /sys/kernel/tracing is powerful but fiddly, so most engineers reach for trace-cmd, a user-space front-end that wraps Ftrace in a handful of simple sub-commands.

The general workflow, using the function_graph tracer, looks like this:

#  1. Start recording while your test app runs against the driver
$ sudo trace-cmd record -p function_graph -F ./my_test_app /dev/my_driver

#  2. Convert the binary trace into a readable report
$ sudo trace-cmd report -I -S -l > report.txt

#  3. Search the report for the write handler and everything under it
$ grep -A 40 "write_my_driver" report.txt | less

The -F flag tells trace-cmd to only trace while the given command is running, which keeps the report focused instead of drowning you in unrelated system activity.

Reading a function_graph Report (Conceptual Flow)
write(2) syscall
  → vfs_write()
    → security_file_permission()
  → your_driver_write()
    → spin_lock(&lock) <– critical section begins
    → your_delay_helper()
      → schedule_timeout() <– suspicious!
        → schedule() <– confirmed bug: sleeping under a spinlock
    → spin_unlock(&lock)

Once you can see schedule() being entered between a spin_lock() and its matching spin_unlock(), the bug is proven beyond doubt — regardless of whether the kernel printed anything about it.

Why the Bug Stayed Silent: CONFIG_DEBUG_ATOMIC_SLEEP

The kernel does have a built-in check for exactly this mistake, but it’s a debug feature that has to be compiled in. You can check whether your running kernel has it with:

$ grep DEBUG_ATOMIC_SLEEP /boot/config-$(uname -r)

On most general-purpose distribution kernels — Ubuntu, Fedora, Debian and friends — this line will read is not set. Distribution maintainers disable heavyweight debug checks like this one because they add runtime overhead on every single lock/unlock path, which production systems can’t afford. The consequence for you as a driver author is simple: your day-to-day development kernel will hide this class of bug from you.

This is why serious kernel development is done on a debug kernel — a kernel built with options such as CONFIG_DEBUG_ATOMIC_SLEEP, CONFIG_PROVE_LOCKING (lockdep), and CONFIG_DEBUG_SPINLOCK switched on. With those enabled, the exact same buggy driver would immediately print a “scheduling while atomic” warning with a full call trace the very first time it ran, instead of staying quiet for months.

A Modern Companion: rtla on Kernel 6.x

Kernel 6.x adds another useful tool to this picture: rtla (Real-Time Linux Analysis), a user-space suite built on top of Ftrace’s newer tracers. Where trace-cmd is a general-purpose recorder, rtla is purpose-built for finding scheduling latency and unexpected delays — the same family of problem we just diagnosed. A quick timerlat sample looks like this:

$ sudo rtla timerlat top -d 10s

This doesn’t replace the function-level detail of trace-cmd report, but it’s a fast way to notice that something is causing latency spikes before you go hunting for the exact function call responsible.

Practical Takeaways

  • Never trust “no error printed” as proof your locking is correct.
  • Keep a debug kernel image around specifically for testing new driver code.
  • Reach for trace-cmd record -p function_graph whenever a bug is suspected but nothing is being logged.
  • On kernel 6.x, pair Ftrace with rtla for a quicker first signal of scheduling trouble.
Coming Up Next

Now that we can prove a driver is misbehaving, we turn to a new source of the same kind of bug: hardware interrupt handlers racing with your driver’s normal read/write code. That’s where a plain spinlock, on its own, starts to fall short.

Interrupt-safe locking spin_lock_irqsave Race conditions

FAQ

Q1. Is trace-cmd the same as ftrace?
No. Ftrace is the kernel’s built-in tracing engine; trace-cmd is a user-space tool that drives Ftrace for you so you don’t have to write to tracing files by hand.

Q2. Do I need to rebuild my kernel to use trace-cmd?
No, basic function tracing works on a normal kernel. You only need a special debug kernel build to get automatic warnings like “scheduling while atomic” printed for you.

Q3. What does the function_graph tracer actually show?
It shows every kernel function entered and exited, indented by call depth, with timestamps and duration — effectively a live call stack over time.

Q4. Why is CONFIG_DEBUG_ATOMIC_SLEEP disabled by default?
It adds a runtime check on every code path that might sleep, which costs performance. Distributions optimize for speed on production kernels, not debug visibility.

Q5. Is rtla available on older kernels?
rtla depends on newer Ftrace tracers (osnoise/timerlat) that matured in the 6.x series; older kernels may lack the required tracers.

Q6. Can this same technique find other silent bugs?
Yes. Any bug that “should” trigger a kernel warning but only does so under specific debug configs can be exposed the same way: trace the call path and look for the offending function directly.

Q7. What’s the safest way to test a debug kernel?
Always test on a disposable virtual machine or a spare board, never on a production system, since debug kernels are slower and more verbose by design.

{ “@context”: “https://schema.org”, “@type”: “FAQPage”, “mainEntity”: [ {“@type”: “Question”, “name”: “Is trace-cmd the same as ftrace?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “No. Ftrace is the kernel’s built-in tracing engine; trace-cmd is a user-space tool that drives Ftrace for you so you don’t have to write to tracing files by hand.”}}, {“@type”: “Question”, “name”: “Do I need to rebuild my kernel to use trace-cmd?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “No, basic function tracing works on a normal kernel. You only need a special debug kernel build to get automatic warnings like scheduling while atomic printed for you.”}}, {“@type”: “Question”, “name”: “What does the function_graph tracer actually show?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “It shows every kernel function entered and exited, indented by call depth, with timestamps and duration, effectively a live call stack over time.”}}, {“@type”: “Question”, “name”: “Why is CONFIG_DEBUG_ATOMIC_SLEEP disabled by default?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “It adds a runtime check on every code path that might sleep, which costs performance. Distributions optimize for speed on production kernels, not debug visibility.”}}, {“@type”: “Question”, “name”: “Is rtla available on older kernels?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “rtla depends on newer Ftrace tracers (osnoise and timerlat) that matured in the 6.x series; older kernels may lack the required tracers.”}}, {“@type”: “Question”, “name”: “Can this same technique find other silent bugs?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “Yes. Any bug that should trigger a kernel warning but only does so under specific debug configs can be exposed the same way, by tracing the call path and looking for the offending function directly.”}}, {“@type”: “Question”, “name”: “What’s the safest way to test a debug kernel?”, “acceptedAnswer”: {“@type”: “Answer”, “text”: “Always test on a disposable virtual machine or a spare board, never on a production system, since debug kernels are slower and more verbose by design.”}} ] }
Keep Learning Linux Kernel Programming for Free

Part of EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course.

Browse the Full Course

← Previous Lecture    Next Lecture →

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *