Linux CPU Idle And CPUFreq
Part of EmbeddedPathashala’s free Linux kernel development course — a deep dive into the two dynamic power-management interfaces every kernel and driver engineer must understand.
CPU idle and CPUFreq are the two pillars of dynamic power management inside the Linux kernel. Every time a processor core has nothing to run, the CPU idle framework decides how deeply it should sleep; every time a core is busy, CPUFreq (the kernel’s DVFS layer) decides how fast it should run. Together, CPU idle and CPUFreq are why a phone SoC or an embedded ARM board can idle at a few milliwatts and burst to several watts under load without a single line of application code knowing the difference. This lecture is part of our free Linux kernel development course and continues directly from the previous lecture on power management fundamentals (suspend/resume and system sleep states). Here we go one layer deeper: the always-active, always-running mechanisms that shave power off the CPU even while the system stays fully awake.
What You Will Learn
Prerequisites
Why Dynamic CPU Power Management Matters
Suspend-to-RAM and hibernation, which we covered in the previous lecture, save power when the whole system is not in use. But a running system spends the overwhelming majority of its time neither fully busy nor fully asleep — it is bursty. A UART interrupt fires, a task runs for 200 microseconds, and the CPU goes back to waiting. Multiply that pattern across dozens of interrupts per second and it becomes clear that the biggest energy savings on a live system come from what happens between those bursts, not from full system sleep. That is precisely the job split between CPU idle and CPUFreq: CPU idle governs what a core does when it has no work, and CPUFreq governs how fast a core executes when it does have work. Understanding CPU idle and CPUFreq is fundamental to writing embedded Linux drivers that do not silently defeat the kernel’s power-saving logic — an interrupt handler that never completes, or a driver that busy-polls a register, can keep a core out of deep idle states indefinitely.
CPU Idle: Putting Cores To Sleep Between Bursts Of Work
The CPUIdle subsystem is invoked from the idle loop — the code the scheduler runs on a CPU when there are no runnable tasks left except the special idle task. Every idle processor architecture exposes one or more low-power modes, generically called C-states in ACPI terminology (C0 is “running”, C1 and deeper are idle states of increasing depth). The kernel abstracts these hardware-specific states behind a platform-supplied CPUIdle driver, which presents them to the rest of the kernel as a flat, ordered array of idle state descriptors. The decision of which state to actually request is delegated to a separate module called the governor, so the CPU idle framework as a whole is a clean split between “what states exist” (driver) and “which one to pick right now” (governor).
C-States, Exit Latency, And Target Residency
Each idle state the driver exposes carries two numbers that the governor treats as the ground truth for its decision: exit latency and target residency. Exit latency is the worst-case time, in microseconds, between a wakeup event and the CPU actually executing its first instruction again — deeper states power down more logic (caches, voltage rails, sometimes the whole cluster) and therefore take longer to bring back up. Target residency is the minimum amount of time the core must remain in that state for the energy spent entering and leaving it to be worth it at all; entering and exiting a deep C-state is not free, since ramping voltage rails and re-synchronizing PLLs itself costs energy and time. A governor that requests a deep C-state for an idle period shorter than its target residency actually wastes more power than staying in a shallower state or simply polling. This is why the governor’s real job is prediction: it has to estimate how long the CPU is about to be idle before it commits to a state.
Idle Governors: menu, teo, ladder, And haltpoll
Current mainline kernels ship four CPUIdle governors, and which one is active by default depends on whether the running kernel is “tickless” (built with CONFIG_NO_HZ_IDLE, which lets the periodic scheduler tick be stopped on an idle CPU):
- menu — the long-standing default for tickless systems. It keeps a rolling history of the last several observed idle durations, applies a pattern-recognition pass to find a “typical interval,” and combines that with the time until the next known timer event to pick the deepest state whose target residency still fits.
- TEO (Timer Events Oriented) — a newer governor built around the observation that on most systems timer interrupts dominate wakeup patterns far more than any other source. Instead of menu’s statistical history, TEO tracks per-idle-state “hit” and “intercept” counters and leans on the time until the next timer as the primary signal. Many current distributions now default to TEO on tickless systems because it reacts faster to changing workloads with less bookkeeping overhead than menu.
- ladder — the default governor when the tick cannot be stopped (non-tickless / periodic-tick kernels). It steps one state deeper or shallower per idle-loop iteration based on how the previous sleep compared to that state’s residency — simple, low overhead, but far less adaptive than menu or TEO.
- haltpoll — intended for virtual machines. It briefly polls before actually halting the vCPU, avoiding the cost of a VM-exit for very short idle periods, and falls back to a real idle state (via ladder or menu semantics) if the idle period turns out to be longer than expected.
Both the CPU idle driver and the active governor are visible from user space, and the governor can be swapped at runtime — a detail we exercise hands-on in the next lecture.
The CPU Idle sysfs Layout
Two separate sysfs trees expose CPU idle to user space. Global driver and governor identity live directly under /sys/devices/system/cpu/cpuidle/, while the per-CPU, per-state statistics live under each CPU’s own cpuidle directory.
/sys/devices/system/cpu/cpuidle/current_driver # active CPUIdle driver, e.g. intel_idle
/sys/devices/system/cpu/cpuidle/current_governor_ro # active governor (read-only view)
/sys/devices/system/cpu/cpuidle/current_governor # writable on kernels that allow switching
/sys/devices/system/cpu/cpuidle/available_governors # governors compiled/loaded into this kernel
/sys/devices/system/cpu/cpu0/cpuidle/state0/ # shallowest state (often POLL)
/sys/devices/system/cpu/cpu0/cpuidle/state1/
/sys/devices/system/cpu/cpu0/cpuidle/state2/ # deepest state on this CPU
Every stateN directory exposes the same set of attributes: name and desc (identity), latency (exit latency, microseconds), residency (target residency, microseconds), power (rarely populated accurately), usage (times entered), time (cumulative time spent there), and a writable disable attribute that lets an administrator forbid the governor from ever selecting that state on that particular CPU — useful when a deep C-state introduces unacceptable interrupt latency on a real-time workload.
CPU Idle Decision Path
CPU Hotplug And Its Relationship To Idle Power Saving
CPU hotplug is the kernel facility that lets logical CPUs be taken fully offline and brought back online at runtime, independent of the idle framework. An offline CPU does not merely enter a deep C-state — it is removed from the scheduler’s list of runnable CPUs entirely, its interrupts are migrated away, and its cache/TLB state is torn down. This is a coarser, more expensive operation than idling, but it is the right tool when a whole core (not just its current burst of work) is not needed for an extended period, such as thermal mitigation, or reducing static leakage power on big.LITTLE-style systems by parking an entire cluster. Hotplug and CPU idle are complementary: idle governors handle sub-millisecond decisions many times per second, while hotplug handles coarse-grained decisions (often made by userspace power daemons or thermal governors) that persist for seconds or longer. Hotplugging a CPU is controlled through a simple per-CPU sysfs attribute:
/sys/devices/system/cpu/cpuX/online # 1 = online, 0 = offline; writable
Note that CPU0 is special on most architectures and historically could not be offlined at all; many modern kernels now allow it via CONFIG_BOOTPARAM_HOTPLUG_CPU0, but treat it as the exception rather than the rule.
CPUFreq And DVFS: Scaling Voltage And Frequency
Where CPU idle deals with “no work,” CPUFreq — the kernel’s Dynamic Voltage and Frequency Scaling (DVFS) subsystem — deals with “how much work, how fast.” Most modern processors can run at several discrete clock frequency and voltage combinations, and as a rule, higher frequency and voltage retire more instructions per second at the cost of a roughly cubic increase in dynamic power draw. CPUFreq is built from three cooperating layers: the core (common infrastructure and the sysfs interface), scaling governors (the algorithms that decide the target frequency), and scaling drivers (the platform code that actually reprograms the clock/voltage hardware). CPUs whose frequency control hardware is shared — for example, all cores in a cluster driven by the same PLL — are grouped by the core into a single cpufreq_policy object, so a “policy” rather than an individual CPU is really the unit CPUFreq operates on.
Operating Performance Points (OPP)
An Operating Performance Point, or OPP, is simply a validated (frequency, voltage) pair that the silicon vendor guarantees is stable and within thermal/power budget. On device-tree based embedded platforms the OPP table is typically described directly in the device tree and consumed through the kernel’s generic OPP framework (drivers/opp), which both CPUFreq and devfreq scaling drivers query to discover which frequencies are legal and what voltage each one requires from the regulator. Thinking in terms of OPPs rather than raw frequencies matters because voltage does not scale linearly with frequency, and skipping the OPP abstraction (hardcoding frequency-to-voltage tables in a driver) is one of the most common sources of instability when porting CPUFreq to a new SoC.
CPUFreq Governors Compared
CPUFreq ships several generic governors, and picking the right one is the single biggest lever an embedded engineer has over a platform’s performance/power tradeoff without touching silicon. A key shift since the 4.19-era material this lecture updates: schedutil — not ondemand — is now the default governor on most current distributions and is the governor upstream actively develops, because it reads CPU utilization directly from the scheduler instead of re-measuring load with a separate timer, and it is a prerequisite for Energy Aware Scheduling (EAS) on asymmetric (big.LITTLE) ARM platforms.
| Governor | Strategy | Optimizes For |
|---|---|---|
| performance | Always requests the maximum allowed frequency | Lowest latency, worst power |
| powersave | Always requests the minimum allowed frequency | Lowest power, worst latency |
| userspace | Frequency set explicitly by a user-space writer via scaling_setspeed | Custom / test harnesses, manual tuning |
| ondemand | Samples load on a workqueue timer; jumps to max above a threshold | Legacy responsive scaling, higher context-switch overhead |
| conservative | Same load sampling as ondemand, but steps frequency gradually | Smoother scaling for battery-sensitive systems |
| schedutil | Reads scheduler PELT utilization directly, no separate sampling timer | Current default: fast, scheduler-integrated, EAS-ready |
Governor tunables are exposed as extra files under a per-policy or global <governor-name> directory — for instance ondemand’s up_threshold and sampling_rate, or schedutil’s single rate_limit_us, which bounds how often the governor is allowed to recompute frequency, keeping its scheduler-context overhead bounded.
The CPUFreq sysfs Interface
Every CPUFreq policy is exposed under /sys/devices/system/cpu/cpufreq/policyX/, with a symlink of the same name reachable from each affected CPU’s own directory:
/sys/devices/system/cpu/cpuX/cpufreq/ -> symlink into policyX
/sys/devices/system/cpu/cpufreq/policyX/scaling_governor # active governor, writable
/sys/devices/system/cpu/cpufreq/policyX/scaling_available_governors
/sys/devices/system/cpu/cpufreq/policyX/scaling_available_frequencies
/sys/devices/system/cpu/cpufreq/policyX/scaling_cur_freq # last frequency requested
/sys/devices/system/cpu/cpufreq/policyX/cpuinfo_cur_freq # frequency read back from hardware
/sys/devices/system/cpu/cpufreq/policyX/scaling_min_freq
/sys/devices/system/cpu/cpufreq/policyX/scaling_max_freq
/sys/devices/system/cpu/cpufreq/policyX/affected_cpus # CPUs sharing this policy
The distinction between scaling_cur_freq and cpuinfo_cur_freq trips up a lot of newcomers: the former is what the driver last asked for, the latter (where supported) is what the hardware actually reports back, and on systems with aggressive hardware-level boosting the two can legitimately disagree.
CPUFreq Layered Architecture
devfreq: DVFS Beyond The CPU
CPUFreq only scales CPU cores. Everything else on an SoC that also has its own clock/voltage domain — GPU, memory/DDR interconnect, image signal processor, and similar blocks — is scaled by devfreq, the generic, non-CPU DVFS framework. Architecturally devfreq mirrors CPUFreq closely: a device driver registers with devfreq and supplies an OPP table plus a way to read the device’s current utilization, and a devfreq governor (simple_ondemand and userspace are the most common) decides which OPP to request. Each registered device shows up under /sys/class/devfreq/<device-name>/ with an almost one-to-one mirror of the CPUFreq attribute names (available_frequencies, governor, cur_freq, min_freq, max_freq). If you understand CPU idle and CPUFreq well, devfreq requires almost no additional conceptual investment — it is the same governor/driver/core split applied to a different class of hardware block.
Common Mistakes And Troubleshooting CPU Idle And CPUFreq Issues
- Blaming the wrong layer for latency. A missed real-time deadline is often caused by a deep C-state’s exit latency, not the CPUFreq governor. Check
cpuidle/stateN/latencybefore touching frequency scaling. - Busy-polling drivers defeating idle entirely. A driver that polls a status register in a tight loop instead of sleeping/waiting on an interrupt keeps the run queue non-empty, so the CPU idle governor never even gets invoked — power numbers look wrong for reasons that have nothing to do with governor tuning.
- Assuming ondemand is still the default. Documentation and forum answers written against 4.x-era kernels still describe ondemand as default; on current kernels schedutil is default in most distribution configs, and tuning advice aimed at ondemand’s
sampling_ratesimply does not apply. - Forgetting a policy can span multiple CPUs. Writing to one CPU’s
scaling_governorsilently changes it for every CPU sharing that policy — checkaffected_cpusfirst, especially on big.LITTLE parts with separate policies per cluster. - Confusing scaling_cur_freq with cpuinfo_cur_freq. The former is a request, the latter (when available) is hardware truth; do not assume they match, especially under hardware-managed boosting.
- Disabling a deep C-state globally by accident. The
disableattribute is per-CPU; forgetting to write it to every CPU in a cluster leaves the state selectable elsewhere and the “fix” appears to do nothing.
Best Practices For CPU Idle And CPUFreq Tuning
- Prefer schedutil unless you have a specific, measured reason to use ondemand/conservative — it is where upstream development effort and EAS integration are focused.
- Always validate against real OPP tables from the device tree or ACPI rather than hardcoding frequency/voltage pairs in driver code.
- Use
cpuidle/stateN/disablesparingly and per-workload; disabling a deep state everywhere trades away most of the platform’s power budget. - When writing interrupt handlers or drivers, complete work promptly and go back to sleep/wait — this is what lets CPU idle and CPUFreq do their job correctly.
- Measure before tuning: read
usage/timeunder cpuidle andscaling_cur_freqhistory under cpufreq rather than guessing which governor “feels” faster. - On multi-policy (big.LITTLE) systems, tune each policy independently — a single global governor choice is rarely optimal across heterogeneous cores.
Summary And Key Takeaways
CPU idle and CPUFreq are two independent but complementary kernel subsystems that together implement dynamic power management for the CPU while the rest of the system stays awake. CPU idle chooses how deeply an otherwise-idle core sleeps, guided by governors (menu, TEO, ladder, haltpoll) that weigh each state’s exit latency against its target residency, all exposed through the /sys/devices/system/cpu/cpuidle/ and per-CPU cpuidle/ sysfs trees. CPU hotplug is the coarser, complementary tool for taking an entire core offline. CPUFreq, the kernel’s DVFS engine, chooses how fast a busy core runs, using Operating Performance Points to know which (frequency, voltage) pairs are safe, and delegating the actual decision to a scaling governor — with schedutil now the modern default, superseding ondemand and conservative for most workloads. devfreq extends the exact same governor/driver architecture to non-CPU blocks like the GPU and memory controller. Mastering CPU idle and CPUFreq, and knowing exactly which sysfs knob controls which behavior, is essential groundwork before this free Linux kernel development course moves on to thermal management and system sleep states in later lectures.
Frequently Asked Questions About CPU Idle And CPUFreq
What is the practical difference between CPU idle and CPUFreq?
CPU idle manages power when a core has no work to do, choosing among C-states of increasing depth and wakeup latency. CPUFreq manages power and performance when a core does have work, choosing the clock frequency and voltage it runs at. They run independently and simultaneously on every CPU.
Is ondemand still the default CPUFreq governor?
No. On current mainline kernels and most modern distributions, schedutil is the default governor. ondemand and conservative remain available and fully functional but are considered legacy for general-purpose systems, since schedutil integrates directly with the scheduler’s utilization tracking.
What is the TEO governor and why does it exist alongside menu?
TEO (Timer Events Oriented) is a newer CPU idle governor that leans primarily on the time until the next known timer interrupt rather than menu’s statistical history of past idle durations. It is available on tickless systems as an alternative to menu and several distributions now default to it because of its lower overhead and faster reaction to changing idle patterns.
How do exit latency and target residency affect which C-state is chosen?
The governor predicts how long the CPU is about to be idle, then picks the deepest state whose target residency fits inside that predicted duration and whose exit latency stays within any active PM QoS limit. Picking a state whose target residency exceeds the actual idle time wastes more energy than staying shallow.
Does CPU hotplug replace the need for CPU idle?
No. CPU hotplug and CPU idle solve different time scales. CPU idle makes decisions many times per second for brief idle windows; hotplug removes a core from scheduling entirely for extended periods and is a much heavier-weight operation, typically driven by thermal or power-policy daemons in user space.
What exactly is an Operating Performance Point (OPP)?
An OPP is a validated pairing of CPU (or device) frequency and the supply voltage required to run reliably at that frequency, usually described in the device tree. CPUFreq and devfreq scaling drivers both consult the OPP table rather than assuming arbitrary frequency/voltage combinations are safe.
What is devfreq and how does it relate to CPUFreq?
devfreq is the generic DVFS framework for non-CPU devices such as GPUs and memory interconnects. It mirrors CPUFreq’s architecture — a core, a governor, and a driver per device — and exposes a similar sysfs interface under /sys/class/devfreq/.
Can I disable a specific idle state on only one CPU?
Yes. The disable attribute under each CPU’s cpuidle/stateN/ directory is per-CPU. Writing 1 disables that state only for the CPU whose directory you wrote to; it remains available on every other CPU unless disabled there as well.
Ready To Get Hands-On With CPUFreq?
Continue to the next lecture in this free Linux device drivers course, where we exercise every sysfs knob covered here and write an original CPUFreq notifier kernel module.
Start Hands-On Lecture Browse Full Course Index