trace-cmd and KernelShark Tutorial: Watch the Linux Scheduler Work in Real Time
Part of our free Linux kernel development course — record, report, and visualize scheduler activity on modern 6.x kernels
Intermediate
6.x (EEVDF era)
~18 minutes
In the previous lecture of this free Linux kernel development course we learned to decode the Ftrace latency format by hand. Writing to tracefs files works, but it gets tedious fast. trace-cmd is the official command-line front end for Ftrace: one tool that records, filters, saves, and reports kernel traces without you ever touching /sys/kernel/tracing manually. And when a text report is not enough, KernelShark renders the same data as an interactive, zoomable timeline.
In this lecture of our free Linux kernel programming course, we record real scheduler activity with trace-cmd, then read the report to watch something remarkable: the core scheduler walking its list of scheduling classes in strict priority order — stop, deadline, real-time, and finally the fair class, which on kernels 6.6 and later is implemented by EEVDF rather than the old CFS. This is the same scheduler theory from our free embedded systems course lectures, now observed live on your own machine.
What You Will Learn
trace-cmd record / report
function_graph plugin
Scheduler class walk in traces
EEVDF fair class
sched events tracing
KernelShark 2.x GUI
Latency analysis workflow
Prerequisites
- Completion of the previous lecture on the Ftrace latency format (you must be able to decode the status characters).
- Familiarity with modular scheduling classes from earlier lectures in this free Linux kernel development course.
- A Linux system with kernel 5.10+ (6.x recommended), root access, and a package manager.
- For KernelShark: a desktop environment or the ability to forward X/Wayland from your target board to a host PC.
Installing trace-cmd and KernelShark
Both tools are packaged by every major distribution. Current releases are trace-cmd 3.x and KernelShark 2.x, which added a rewritten Qt interface and plugin support:
# Debian / Ubuntu
sudo apt update
sudo apt install trace-cmd kernelshark
# Fedora
sudo dnf install trace-cmd kernelshark
# Arch
sudo pacman -S trace-cmd kernelshark
# Verify versions
trace-cmd --version
kernelshark --version # or check Help > About in the GUI
For embedded targets built with Yocto or Buildroot, enable the trace-cmd recipe on the target and install KernelShark on your host machine only — you record on the board, copy the data file over, and analyze on the desktop.
Recording a Scheduler Trace with trace-cmd
trace-cmd wraps the whole tracefs workflow into a single command. Let us capture the kernel functions executed while a short-lived process runs. The -p function_graph option selects the function graph plugin, and -F restricts tracing to just the command we launch, keeping the capture small and readable:
# Record kernel function calls made while 'ls -lR /etc' executes
sudo trace-cmd record -p function_graph -F ls -lR /etc > /dev/null
# The capture is saved as trace.dat in the current directory
ls -lh trace.dat
# Generate the human-readable report
trace-cmd report | less
Every line of the report carries the fields we decoded in the last lecture: the task name and PID, the five-character latency status field, a timestamp in seconds and microseconds, the event name (funcgraph_entry or funcgraph_exit for this plugin), the optional duration with its delay marker, a separator bar, and finally the kernel function name with C-style braces showing entry and exit nesting.
| Line # | Task-PID | Status Field | Timestamp | Event | Duration | Function |
| record index | who ran | CPU + IRQ + N + ctx + depth | sec : µs | entry / exit | time + marker | call tree with { } |
A useful trade-off to know: if you drop the -p function_graph option and record with the plain function tracer instead, you lose the beautiful indented call-graph — but you gain the runtime argument values of many traced functions printed on each line, which can be exactly what you need when debugging a driver in your free Linux device drivers course projects.
Watching the Scheduler Walk Its Class List
Here is where theory meets reality. In our scheduler lectures, we learned that Linux organizes scheduling policies into modular scheduling classes consulted in strict priority order. When a context switch is needed, the core scheduler asks each class in turn: do you have a task that wants to run? The first class that says yes wins the CPU.
| 1. Stop per-CPU stopper |
→ | 2. Deadline SCHED_DEADLINE |
→ | 3. Real-Time FIFO / RR |
→ | 4. Fair (EEVDF) SCHED_NORMAL / BATCH |
→ | 5. Idle idle task |
Kernel 6.12+ can also insert the sched_ext (BPF) class into this walk when a custom BPF scheduler is loaded. On most desktops, the fair class answers “yes” almost every time.
In a function_graph capture of a context switch, you can literally see this negotiation in the report: the pick-next helpers of the higher-priority classes are invoked one after another — the stop class, then the deadline class, then the real-time class — each returning without a task on an ordinary desktop. Search your own capture for them:
# Find the class-walk functions inside your capture
trace-cmd report | grep -E 'pick_next_task' | head -15
Then comes a lovely subtlety. On a typical system, the winning class is almost always the fair class — so why might its own pick function be missing from the walk in your trace? Because the kernel developers optimized for the common case: the core scheduler first checks whether only fair-class (and idle) tasks are runnable, and if so, it short-circuits directly into the fair-class fast path instead of politely walking the whole list. On 6.6+ kernels that fast path lands in EEVDF code, which selects the runnable task with the earliest eligible virtual deadline. Optimizations like this are visible only through tracing — no amount of reading source code in isolation makes them as obvious as one good capture.
Tracing Scheduler Events Instead of Functions
Function tracing shows how the scheduler works internally. For day-to-day latency work, tracing scheduler events is lighter and often more useful. This original example records wakeups and context switches system-wide for three seconds:
# Record scheduler tracepoints for 3 seconds, system-wide
sudo trace-cmd record -e sched_switch -e sched_wakeup sleep 3
# Report, showing which tasks ran on which CPUs and when
trace-cmd report | head -40
# Measure wakeup-to-run latency per task (built-in profile)
sudo trace-cmd record --profile -e sched sleep 3
trace-cmd report --profile | less
Visual Analysis with KernelShark
KernelShark is the GUI companion to trace-cmd. It opens the same trace.dat file and draws a per-CPU, per-task timeline you can zoom, filter, and click through — far friendlier than scrolling a hundred thousand text lines when you are hunting for one bad context switch:
# Open your capture in the GUI
kernelshark ./trace.dat
A practical KernelShark workflow for scheduling latency looks like this:
- Graph area (top): one row per CPU; colored bars show which task occupied each core over time. Gaps and rapid color changes reveal idle time and thrashing.
- Zoom: drag a region to zoom in; keep zooming until individual
sched_switchevents become visible as markers. - Filters: restrict the view to one task or one event type to isolate, say, the wakeup path of your real-time thread.
- List area (bottom): the raw event lines, synchronized with the graph — clicking a bar jumps to the matching text record, where you can read the latency status field with the skills from the previous lecture.
KernelShark 2.x also supports plugins and can load multiple data streams side by side, which is handy for comparing a “good” run against a “bad” run of the same workload. That said, for very advanced multi-source analysis, tools like Trace Compass remain more feature-rich; KernelShark’s strength is being fast, simple, and perfectly matched to trace-cmd data.
Real-World Use Cases
- Embedded audio/BLE stacks: trace
sched_wakeuptosched_switchlatency for your audio or Bluetooth thread to prove (or disprove) that scheduling jitter causes glitches. - Driver bring-up: record with the plain function tracer around your driver’s IRQ path to see argument values and confirm the top half stays short.
- PREEMPT_RT validation: combine the latency status field (watch for long
dand non-zero preempt depth) with duration markers to find preemption-off regions on a real-time kernel. - CPU affinity debugging: KernelShark’s per-CPU rows instantly expose a pinned thread migrating to the wrong core.
Common Mistakes and Troubleshooting
- Running trace-cmd record without root: recording needs privileged access to tracefs; use
sudo. Reading an existingtrace.datwithtrace-cmd reportworks as a normal user. - Huge trace.dat files: recording function_graph system-wide fills disks fast. Use
-Fto follow one command,-P <pid>for one process, or event tracing (-e sched) instead of full function tracing. - Analyzing on a different machine and missing symbols:
trace.datembeds everything needed for report/KernelShark, so copying it from an embedded target to a host is fine — but keep the file from the same boot as any side notes you take. - Expecting CFS function names on a 6.6+ kernel: older articles reference CFS-era internals; on EEVDF kernels, some fair-class helper names and behaviors differ. Trust your trace over an old book.
- KernelShark showing an empty graph: usually a filter left enabled from a previous session — clear all filters from the Filter menu and reload.
Performance Considerations
Tracing is observation with a cost. Function_graph tracing adds measurable overhead to every kernel function call, which can itself shift timings on a slow embedded core — a classic observer effect. Event tracing (sched tracepoints) is dramatically cheaper and is the right default for latency measurement. Bound your buffers (trace-cmd record -b <kb-per-cpu>), record for seconds rather than minutes, and when numbers matter, disable CPU frequency scaling for the duration of the experiment so durations are comparable across runs.
Best Practices
- Start with sched events, escalate to function_graph only when you need to see inside the scheduler.
- Always capture a short “known-good” baseline trace so you have something to compare against in KernelShark.
- Name your data files descriptively (
trace_ble_audio_glitch_6.12.dat) — a week later,trace.dattells you nothing. - Read the manual pages:
man trace-cmd-record,man trace-cmd-report, and friends each end with worked examples that answer most questions.
Key Takeaways
- trace-cmd is the front end that makes Ftrace practical:
recordcaptures totrace.dat,reportrenders it as text. - A function_graph capture of a context switch shows the scheduler consulting its classes in priority order — stop, deadline, real-time — before the fair class (EEVDF on 6.6+) takes the CPU.
- The fair-class fast path is a real, visible optimization: when only fair tasks are runnable, the full class walk is skipped.
- Scheduler event tracing (
-e sched) is the low-overhead default for latency measurement; function tracing is for understanding internals. - KernelShark turns the same trace.dat into an interactive per-CPU timeline — record on the target, analyze on the host.
Conclusion
You now own a complete scheduler-observation workflow on modern Linux: decode the latency status field by eye, capture precisely-scoped traces with trace-cmd, recognize the scheduling-class walk (and its fast-path shortcut) in the report, and zoom through the timeline in KernelShark when text is not enough. These tools cost nothing, ship with every distribution, and work on everything from a workstation to a tiny ARM board — which is exactly why they anchor this stage of our free Linux kernel development course.
From here, the same workflow scales into every corner of this free embedded systems course: profiling driver IRQ paths in the free Linux device drivers course modules, validating PREEMPT_RT latencies, and dissecting wakeup chains in real products. Keep your practice traces — we will refer back to them in upcoming lectures.
Frequently Asked Questions (FAQ)
1. What is the difference between Ftrace and trace-cmd?
Ftrace is the tracing engine inside the kernel; trace-cmd is a user-space tool that drives it. Everything trace-cmd does could be done by writing tracefs files manually — trace-cmd just makes it fast, scriptable, and saves results in a portable trace.dat format.
2. Which scheduling classes does the core scheduler consult, and in what order?
Highest to lowest: stop, deadline (SCHED_DEADLINE), real-time (SCHED_FIFO/SCHED_RR), fair (SCHED_NORMAL/SCHED_BATCH, implemented by EEVDF on kernels 6.6+), and finally idle. Kernels 6.12+ can additionally place a BPF-driven sched_ext class into the hierarchy.
3. Why doesn’t the fair class’s pick function appear in the class walk in my trace?
Because of a fast-path optimization: when the kernel detects that only fair-class tasks are runnable (the overwhelmingly common case), it skips the full walk and calls the fair-class selection code directly.
4. Is CFS gone from the Linux kernel?
Since kernel 6.6, the fair scheduling class is implemented by EEVDF (Earliest Eligible Virtual Deadline First). The class and its policies (SCHED_NORMAL, SCHED_BATCH) remain; the algorithm choosing which fair task runs next changed.
5. Can I run trace-cmd on an embedded board and KernelShark on my PC?
Yes — that is the recommended workflow. Record trace.dat on the target with trace-cmd, copy it to your host over scp or a card, and open it in KernelShark. The data file is self-contained.
6. How much overhead does tracing add?
Idle Ftrace costs essentially nothing. Scheduler event tracing adds very small overhead per event. Full function or function_graph tracing is the heavy option and can noticeably slow a small embedded core — scope it with filters and short sessions.
7. KernelShark vs Trace Compass — which should I learn?
Learn KernelShark first: it is lightweight, opens trace-cmd files natively, and covers most scheduling analysis. Trace Compass is more powerful for large multi-source investigations but has a steeper learning curve.
8. Where can I learn more after this lecture?
The kernel’s own tracing documentation (under Documentation/trace in the source tree and on docs.kernel.org) and the trace-cmd manual pages (man trace-cmd-record, man trace-cmd-report) are the authoritative references and include many worked examples.
Keep Learning — Completely Free
EmbeddedPathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded systems course are open to everyone. Practice these traces on your own board and join us in the next lecture.

2 Comments