Linux Thermal And Sleep States
Lecture 3 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course, covering the thermal framework and an introduction to system sleep states.
Every embedded Linux platform eventually has to answer two very different power questions: how hot is the silicon right now, and how deep can the whole system sleep when nobody is using it? The Linux thermal framework answers the first question by continuously watching temperature sensors and reacting before the hardware damages itself, while the system sleep states subsystem answers the second by defining a ladder of low-power states the entire machine can drop into. This lecture builds a solid mental model of both subsystems against the current mainline kernel, not a decade-old snapshot, so the sysfs paths, kernel APIs, and terminology you learn here match what you will actually find in a 6.x kernel tree.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Working knowledge of C, including pointers and function-pointer (ops) tables.
- Basic familiarity with configuring and building a custom Linux kernel (menuconfig, loadable modules).
- Comfort reading and writing sysfs attributes from a shell.
- Having gone through the previous lecture on CPU Idle and CPUFreq governors in this free Linux kernel development course is strongly recommended, since runtime power scaling and thermal throttling are closely related mechanisms.
Why The Kernel Needs A Dedicated Thermal Framework
Modern SoCs pack several watts of compute into a few square millimeters of silicon. Left unchecked, sustained workloads can push junction temperature past the point where the chip either throttles unpredictably in hardware, corrupts data, or simply dies early. Rather than leaving thermal protection to ad-hoc, driver-specific hacks, the Linux thermal framework centralizes temperature monitoring and cooling decisions into one subsystem that any sensor driver, any cooling mechanism, and any platform can plug into through a small, stable set of interfaces.
The framework deliberately separates “who measures temperature” from “who decides what to do about it” from “who actually reduces power dissipation.” That separation is what lets a single governor algorithm work unmodified across wildly different hardware, from a fanless single-board computer to a multi-core SoC with a PWM-controlled fan.
The Linux Thermal Framework Architecture
The Linux thermal framework is organized around four cooperating building blocks: the thermal zone, the thermal governor, the cooling device, and the thermal core that glues the other three together and exposes them to user space through sysfs.
Thermal Zones And Sensors
A thermal zone represents a piece of hardware whose temperature needs to be tracked, such as a CPU cluster, a GPU, a battery, or a board-level ambient sensor. The thermal zone itself does not measure anything; it delegates that job to a thermal sensor driver, which registers a get_temp() callback that the thermal core polls (or is notified through, on interrupt-driven sensors) to obtain the current reading. Every temperature value that flows through the thermal core is expressed in millidegrees Celsius, so a reading of 45.0°C is represented internally as the integer 45000.
Cooling Devices: Passive Cooling Vs Active Cooling
A cooling device is anything capable of reducing power dissipation once instructed to. The framework recognizes two broad strategies:
- Passive cooling reduces the performance of the offending device instead of adding new hardware. The classic example is throttling CPU frequency and voltage through the same DVFS machinery used by CPUFreq, so the chip simply does less work per second and generates less heat.
- Active cooling engages a dedicated cooling mechanism external to the thing being cooled, such as a GPIO-controlled or PWM-controlled fan, a pump, or a chassis blower. Active cooling recovers thermal headroom without sacrificing performance, at the cost of extra power draw and, often, audible noise.
Every cooling device exposes a small, uniform interface to the thermal core: a maximum throttle state, a current throttle state, and a callback to request a new state. The thermal core does not need to know whether state 3 out of 4 means “70% CPU frequency” or “fan at medium speed” – that mapping is entirely private to the cooling device driver.
Trip Points: Where Policy Meets Hardware Limits
A trip point is a temperature threshold, chosen from hardware datasheet limits, at which some cooling action becomes appropriate. Each trip point carries a type that tells the thermal core what class of response it represents:
- active – engage an active cooling device (fan on).
- passive – engage passive cooling (throttle performance).
- hot – a driver-defined emergency action short of shutdown.
- critical – the zone is about to cause hardware damage; the kernel initiates an emergency shutdown or reboot.
Trip points also carry a hysteresis value, which prevents a temperature that hovers right at the threshold from repeatedly triggering and clearing the same cooling action every polling cycle.
Thermal Governors And Cooling Maps
A thermal governor is the algorithm that decides how aggressively to respond once a trip point is crossed. The kernel ships several, including step_wise (the general-purpose default, which nudges cooling state up or down one step at a time), fair_share, bang_bang, user_space (hands the decision to a userspace daemon through netlink or sysfs), and power_allocator (a PID-style controller that distributes a sustainable power budget across several cooling devices). A cooling map is the binding that says which cooling device should respond to which trip point in which zone, optionally with a weight when several cooling devices share a trip point. On device-tree platforms, cooling maps are declared declaratively in the cooling-maps node; on ACPI or hand-written platform code, the binding is performed programmatically.
Thermal Framework Data Flow
Thermal sysfs Layout At A Glance
The entire thermal subsystem is exposed to user space under a single directory, without needing any dedicated tool:
$ ls /sys/class/thermal/
cooling_device0 cooling_device1 cooling_device2
thermal_zone0 thermal_zone1
Each thermal_zoneN entry corresponds to one registered thermal zone, exposing attributes such as type, temp, policy, and one trip_point_X_temp / trip_point_X_type pair per trip point. Each cooling_deviceN entry exposes type, max_state, and cur_state. The full command walkthrough, including reading real trip points and cooling device statistics, is covered hands-on in the companion examples lecture.
System Power Management: Sleep States Vs Runtime PM
It is easy to conflate runtime power management with system sleep states, but they solve different problems. Runtime PM is per-device and autonomous: an individual driver can suspend its own hardware the moment it notices the device has been idle for a while, and resume it transparently the instant it is needed again, all while the rest of the system, and even the rest of user space, keeps running normally.
System sleep states are global. They put the entire machine, every CPU, every bus, and potentially memory itself, into a coordinated low-power transition. Because only the user, or userspace policy acting on the user’s behalf, actually knows when the whole system is not going to be touched for a while, system sleep transitions are always initiated from user space. The kernel has no notion of “the user stepped away”; it only reacts once something writes the appropriate string into a sysfs control file. This is precisely why so much of working with sleep states is about sysfs and shell commands rather than kernel APIs, a theme this lecture pair follows closely.
The Four Common System Sleep States
The exact sleep states a given platform can offer depend on the underlying hardware and firmware, and they can differ across architectures and even across generations of the same architecture. That said, four sleep states are common enough across platforms that the kernel gives each of them first-class support. Deep hibernation mechanics, image creation, and the restore-kernel handoff are covered in full in the next lecture; here we introduce all four at a conceptual level.
Suspend-to-Idle (Freeze, Sometimes Called S0ix-Style Software Suspend)
This is the lightest, purely software-driven sleep state. User space tasks are frozen, timekeeping is suspended, and every I/O device is pushed into the lowest-power state it supports, so the CPUs can spend nearly all their time in the deepest idle state the CPU Idle framework offers. Because it needs no platform-specific firmware support, suspend-to-idle is always available whenever CONFIG_SUSPEND is enabled, and it is frequently combined with deeper states to shave resume latency.
Standby (Power-On Suspend)
Standby does everything suspend-to-idle does and goes one step further by taking non-boot CPUs completely offline and suspending additional low-level platform functions. It typically saves more energy than suspend-to-idle, at the cost of a somewhat longer, though still modest, resume latency. Support for standby must be explicitly registered by the platform with the core suspend subsystem; it is not automatically available just because CONFIG_SUSPEND is set.
Suspend-to-RAM (Mem, Deep Sleep)
Suspend-to-RAM powers down essentially everything in the system except memory, which is placed into self-refresh so its contents survive. Device and CPU state is saved before power-down and restored afterward. This is the state most people mean when they say “suspend” on a laptop: response latency is noticeably higher than standby but still fast enough to be usable, and the energy savings are substantial because almost nothing besides the RAM refresh circuitry stays powered.
Suspend-to-Disk (Hibernation)
Hibernation offers the deepest power savings of all: a snapshot of memory is written to persistent storage, and then power is cut from nearly the entire system, including memory itself. On resume, a fresh kernel instance boots, locates the saved image, and restores the machine to exactly where it left off. Because it involves writing and later reading a full memory image, hibernation has the longest resume latency of the four states, though it is still faster than a full cold boot and application restart. The mechanics of image creation and the two-kernel restore dance are substantial enough to deserve their own lecture, which follows this one.
Sleep States Compared: Power Savings Vs Resume Latency
| Sleep State | Typical ACPI Analogue | What Stays Powered | Relative Power Savings | Relative Resume Latency |
|---|---|---|---|---|
| Suspend-to-Idle (freeze) | S0 (software idle) | Everything, just deeply idle | Lowest of the four | Fastest |
| Standby | S1 | Core logic, boot CPU | Moderate | Low |
| Suspend-to-RAM (mem) | S3 | RAM in self-refresh only | High | Moderate |
| Suspend-to-Disk (hibernation) | S4 | Nothing (image on disk) | Highest | Slowest (but faster than a cold boot) |
/sys/power/state And /sys/power/mem_sleep Basics
The system sleep sysfs interface lives under /sys/power/. The state file lists which sleep states the running kernel supports; writing one of the listed strings triggers a transition into it.
$ cat /sys/power/state
freeze mem disk standby
The strings freeze, standby, and disk map directly to suspend-to-idle, standby, and hibernation respectively. The string mem is special: what it actually does depends on the currently selected variant in a second file, mem_sleep.
$ cat /sys/power/mem_sleep
[s2idle] deep
The value inside square brackets is the variant that will be used the next time mem is written to state. Possible entries are s2idle (equivalent to suspend-to-idle), shallow (equivalent to standby), and deep (true suspend-to-RAM). Only entries the platform actually supports appear in the file at all, and writing one of the listed strings to mem_sleep changes the selection for subsequent writes of mem. The exact command sequences and expected output for driving these files are demonstrated step by step in the companion examples lecture.
CONFIG_SUSPEND And Why It Matters
System suspend support in the kernel is gated by the CONFIG_SUSPEND configuration option. With it enabled, suspend-to-idle is always usable, since it is a pure software mechanism. Standby and suspend-to-RAM additionally require the underlying platform to register explicit support with the core suspend subsystem; simply enabling CONFIG_SUSPEND does not guarantee those two states exist on a given board. Hibernation is governed by a separate option, CONFIG_HIBERNATION, which itself can only be enabled on architectures that provide the low-level assembly needed to resume execution from a saved image (tracked through ARCH_HIBERNATION_POSSIBLE). We will return to the hibernation-specific configuration requirements, including swap sizing, in the next lecture.
Sleep State Depth Ladder
Common Mistakes And Troubleshooting
- Treating runtime PM as a substitute for system sleep. A device that implements
runtime_suspendis not automatically participating in system-wide suspend; the two mechanisms use related but distinctdev_pm_opscallbacks. - Assuming
memalways means suspend-to-RAM. Ifmem_sleepis not checked first,memmay silently resolve tos2idleon a platform that defaults there, giving far less power savings than expected. - Copy-pasting an old
thermal_zone_device_register()call from an outdated tutorial or book. Current mainline kernels expect trip points to be passed up front throughthermal_zone_device_register_with_trips(), not registered later through per-trip governor callbacks. - Reporting temperature in the wrong unit. The thermal core expects millidegrees Celsius everywhere; a sensor driver that forgets to scale a raw ADC value produces nonsensical readings and trip points that never fire (or always fire).
- Ignoring hysteresis. A trip point with no hysteresis near a noisy sensor can cause cooling actions to flap on and off rapidly.
Best Practices
- Always express and compare temperatures in millidegrees Celsius inside sensor drivers.
- Set a sensible hysteresis on every trip point to avoid oscillation near a threshold.
- Check the bracketed selection in
mem_sleepanddiskbefore scripting any suspend or hibernation automation. - Prefer the current
_with_tripsanddevm_-managed registration APIs for new thermal drivers instead of legacy interfaces. - Test suspend and resume cycles repeatedly, not just once; many sleep-state bugs only surface after several cycles.
- Keep passive cooling responsive enough to prevent thermal runaway, but not so aggressive that performance oscillates visibly.
Summary And Key Takeaways
The Linux thermal framework organizes temperature management into four cooperating pieces: thermal zones that represent hardware to watch, sensors that feed them temperature readings, cooling devices that reduce power dissipation passively or actively, and a thermal core that ties zones, governors, and cooling devices together and exposes the whole picture under /sys/class/thermal/. Trip points, expressed in millidegrees Celsius with an associated type and hysteresis, are where hardware limits become software policy, and cooling maps decide which cooling device answers which trip point.
System sleep states solve an entirely different, system-wide problem: putting the whole machine into a low-power state that is always initiated from user space, never from the kernel itself. The four common states, suspend-to-idle, standby, suspend-to-RAM, and hibernation, trade increasing power savings for increasing resume latency, and are controlled through /sys/power/state and /sys/power/mem_sleep, gated respectively by CONFIG_SUSPEND and, for hibernation, CONFIG_HIBERNATION. With this conceptual foundation in place, the next lecture goes deep on hibernation’s snapshot-and-restore mechanics, and the companion examples lecture puts everything here into your terminal right now.
Frequently Asked Questions
What is the difference between a thermal zone and a thermal sensor?
A thermal zone is the logical representation of a piece of hardware being monitored (a CPU cluster, a battery, a board), registered with the thermal core and exposed under /sys/class/thermal/thermal_zoneN. A thermal sensor is the actual measurement source; its driver supplies the get_temp() callback that the thermal zone calls to obtain a reading.
What is the difference between passive and active cooling in the Linux thermal framework?
Passive cooling reduces performance of the device generating heat, typically by throttling CPU frequency and voltage through the same DVFS mechanism CPUFreq uses. Active cooling engages a separate device, such as a fan or pump, to remove heat without reducing performance, at the cost of extra power draw.
Which sysfs file shows the currently selected suspend-to-RAM variant?
/sys/power/mem_sleep lists the supported variants (s2idle, shallow, deep) with the active one shown inside square brackets. Writing one of the other listed strings changes what happens the next time mem is written to /sys/power/state.
Is suspend-to-idle the same thing as S0ix?
They are closely related but not identical terms. Suspend-to-idle (s2idle, or “freeze”) is the generic, platform-independent kernel mechanism: freeze user space, suspend timekeeping, idle the CPUs as deeply as possible. S0ix is Intel’s platform-specific low-power idle capability that firmware and hardware can layer underneath suspend-to-idle to achieve even deeper power savings on supporting silicon.
Why do system sleep state transitions always have to be started from user space?
Only user space, or the human behind it, actually knows when the system will not be used for a while. The kernel has no built-in concept of “idle for the user’s purposes”; it strictly reacts to a write into /sys/power/state initiated by user space tooling, a display manager, or a power daemon.
What kernel configuration option enables hibernation support?
CONFIG_HIBERNATION must be set, and it can only be enabled on CPU architectures whose port includes the low-level resume code, tracked by ARCH_HIBERNATION_POSSIBLE. This is separate from CONFIG_SUSPEND, which governs the other three sleep states.
Can a platform support suspend-to-RAM without supporting standby?
Yes. Each of standby and suspend-to-RAM must be independently registered with the core suspend subsystem by the platform. A board’s firmware and power sequencing may implement one without the other; only suspend-to-idle is guaranteed whenever CONFIG_SUSPEND is set.
Continue Your Linux Kernel Power Management Journey
Ready to see every sysfs command in this lecture run for real, plus build an original ep_thermal_zone demo driver? Move on to the hands-on examples lecture, and after that, dive into hibernation’s snapshot-and-restore internals as part of this free Linux kernel development course.
Try The Hands-On Examples Next: Hibernation Deep Dive