Linux Kernel Power Management Basics-Free Linux Device Drivers Course

Linux Kernel Power Management Basics

Linux Kernel Power Management Basics

Device Power Management vs System Power Management — the two pillars every driver author must understand before touching a single callback.

Linux kernel power management exists for one blunt reason: batteries are finite and users are impatient. Every embedded board, phone, tablet, or laptop running Linux has to satisfy two contradictory demands at once — be instantly responsive, and sip as little current as possible while doing nothing. This is lecture one of an eleven-part series inside EmbeddedPathashala’s free Linux kernel development course, and its only job is to build the mental model correctly before we write a line of driver code. Get this model wrong, and everything you build on top of it — runtime suspend, wakeup sources, system sleep callbacks — will feel arbitrary instead of obvious. Get it right, and the rest of this series is just filling in the mechanics of ideas you already understand.

The single most important fact to internalize right now: the Linux kernel does not have one power management subsystem. It has two, and they solve genuinely different problems. One manages a single device while the system keeps running normally. The other manages the entire machine when it stops running normally. Confusing these two, or assuming a driver that handles one automatically handles the other, is the most common conceptual mistake new kernel and embedded systems course students make when they first touch struct dev_pm_ops.

What You Will Learn

The performance vs battery-life trade-off Device Power Management (Runtime PM) System Power Management (Sleep States) Why the two subsystems are complementary How runtime-suspended devices interact with system sleep A preview of dev_pm_ops, wakeup sources, and callbacks

Prerequisites

You should be comfortable with basic C — structures, function pointers, and pointer arithmetic — and you should have built a Linux kernel or an out-of-tree module at least once, so terms like menuconfig, insmod, and dmesg are not new to you. Familiarity with the Linux driver model (struct device, struct platform_driver) will help but is not mandatory; we treat it as background for this lecture and will build on it directly in later lectures of this series.

Why Power Management Is a Trade-Off, Not a Feature

It is tempting to think of power management as a bolt-on feature — something a device either “has” or “doesn’t have.” That framing is wrong. Power management is a continuous negotiation the kernel runs on your behalf between two things that cannot both be maximized simultaneously: how fast a device can respond to work, and how little energy it consumes while idle. A CPU running at its highest frequency answers interrupts instantly but burns watts doing nothing. The same CPU clocked down, or parked in a deep idle state, saves enormous energy but takes measurably longer to spin back up to full throughput when a task arrives. The same tension exists for a Wi-Fi radio, a display panel, an eMMC controller, a sensor hub — nearly everything with a clock and a voltage rail.

Because this trade-off cannot be resolved once and for all, Linux kernel power management is built as a set of policies and mechanisms rather than a single switch. The kernel exposes the mechanism — the ability to change a device’s or the system’s power state safely — and leaves the policy of when to actually do it to drivers, subsystems, and ultimately userspace. That separation of mechanism and policy is the design principle that explains almost every API decision you will meet in this series.

The Two Pillars of Linux Kernel Power Management

Officially, the kernel documentation and the driver model split power management responsibilities into exactly two categories, and every power-aware driver ends up dealing with one or both of them.

Device Power Management — Runtime PM

This is the part of Linux kernel power management that acts on one device at a time, while the rest of the system keeps running completely normally. Nothing else pauses. A USB host controller can drop into a low-power link state the moment no device is plugged into it; the moment something is plugged in, it wakes back up — all while your terminal, browser, and every other process on the machine continue executing without interruption. This is why it is also called dynamic power management: the decision is made continuously, on the fly, based on how the device is actually being used right now.

The kernel tracks this per-device state through a usage counter and a runtime status field (RPM_ACTIVE, RPM_SUSPENDED, and the transitional RPM_RESUMING / RPM_SUSPENDING states) living inside struct dev_pm_info. When a driver’s usage counter drops to zero and stays there, the PM core eventually calls the device’s runtime_idle() callback, which typically schedules a runtime_suspend(). The reverse happens the instant something needs the device again. Crucially, this entire mechanism is driven by device activity, not by any decision about the system as a whole.

System Power Management — Sleep States

The second pillar governs the entire machine at once. This is what happens when a laptop lid closes, a phone screen times out and the power button is pressed, or a battery-critical daemon in userspace decides the system must sleep right now to avoid a hard shutdown. Every device in the system is walked, quiesced, and (depending on the target state) has its register state saved, in a strict parent-before-child ordering defined by the device tree in the driver model.

The kernel currently supports up to four sleep states: suspend-to-idle (a pure-software, lightweight state, sometimes called s2idle or freeze), standby (a shallow hardware suspend), suspend-to-RAM (the deep state where nearly everything is powered off except self-refreshing memory), and hibernation (suspend-to-disk, where a full memory snapshot is written to persistent storage and the machine is powered off completely). Critically, the kernel itself does not decide why the system should sleep — a closed lid, an idle timeout, a low battery, or a direct user command are all just different userspace-level reasons that end up writing the same string to the same file. The kernel’s job is only to execute the transition safely and reversibly.

Device PM vs System PM at a Glance

AspectDevice Power Management (Runtime PM)System Power Management (Sleep States)
ScopeOne device at a timeThe entire system together
RunsWhile Linux keeps executing normallyOnly during a suspend / hibernate transition
Decided byThe device driver, from usage countersUserspace policy — lid, idle timer, battery level, user action
Does the kernel care why?Yes — tied to real device activityNo — it only executes the transition it is told to
Core callbacksruntime_suspend, runtime_resume, runtime_idlesuspend, resume, freeze, thaw, poweroff, restore
Userspace entry point/sys/devices/…/power/control/sys/power/state and /sys/power/mem_sleep

Why Neither Subsystem Can Replace the Other

A system that only implements Runtime PM will still burn its battery whenever the user closes the lid or steps away, because Runtime PM never disturbs the CPU scheduler, never stops timekeeping, and never touches devices that report themselves as still in active use. It saves power continuously in the background, but it cannot get anywhere near the savings of a real sleep state, because the system core logic and most peripheral buses stay powered.

Conversely, a system that only implements sleep states and skips Runtime PM entirely wastes power constantly during ordinary use — a Bluetooth radio idling at full power for hours between packets, a camera sensor staying clocked even though no frame has been captured in minutes. Sleep states are heavy-handed and infrequent; Runtime PM is lightweight and continuous. You need both, and the kernel is explicitly designed so that they cooperate rather than compete.

That cooperation is not accidental — the PM core actively checks whether a device is already runtime-suspended before doing any system-sleep work on it. If a device’s prepare() callback reports that it is already sitting in a runtime-suspended state, and every one of its child devices is in the same state, the PM core can skip the suspend(), suspend_late(), and suspend_noirq() phases for it entirely (a mechanism referred to as “direct complete”), and jump straight to the matching complete() callback on resume. In other words, good Runtime PM behavior directly reduces the amount of work a system sleep transition has to do.

Where Runtime PM Meets System Sleep

1. Userspace writes “mem” to /sys/power/state
→2. PM core walks the device tree top-down, calling prepare()
→3. Already runtime-suspended devices may be left exactly as they are
→4. Remaining active devices get suspend() / suspend_late() / suspend_noirq()
→5. The system enters the selected sleep state (s2idle, shallow, or deep)
→6. A wakeup source fires and the system resumes
→7. resume_noirq() / resume_early() / resume() run bottom-up, then complete()

What a Driver Author Will Actually Build: A Forward Look

Everything below is deliberately shallow here — each item gets its own dedicated lecture later in this series. The point right now is just to see where the pieces will slot into the model you just learned.

dev_pm_ops runtime_suspend / runtime_resume / runtime_idle suspend / resume / freeze / thaw / poweroff / restore Wakeup sources Autosuspend delay

Both categories of callback live in the same structure, struct dev_pm_ops, defined in include/linux/pm.h. The system-sleep callbacks occupy one set of function pointers; the three runtime callbacks occupy another. A driver populates whichever ones its hardware actually needs — many drivers only ever implement the runtime trio, letting the generic subsystem-level callbacks handle system sleep by simply resuming the device to full power and then suspending it again afterward.

static const struct dev_pm_ops ep_example_pm_ops = {
	SET_SYSTEM_SLEEP_PM_OPS(ep_example_suspend, ep_example_resume)
	SET_RUNTIME_PM_OPS(ep_example_runtime_suspend,
			    ep_example_runtime_resume,
			    NULL)
};

Wakeup sources are the mechanism that lets a device tell the kernel “I am the reason the system should stop sleeping right now” — a real-time-clock alarm, a network wake-on-LAN packet, or a GPIO tied to a physical button are all classic examples. We will cover device_init_wakeup(), device_may_wakeup(), and the associated sysfs files in depth in a dedicated lecture.

Common Mistakes and Troubleshooting

The mistake we see most often from students new to Linux kernel power management is treating “the device is suspended” as a single, universal fact. It is not — a device can be runtime-suspended while the system is fully awake, and a system can be in suspend-to-RAM while a specific device was left in its own runtime-suspended state throughout the whole transition, never touched by the sleep-state callbacks at all. When debugging, always ask which of the two subsystems you are actually looking at before you start reading register dumps.

A second common error is assuming that implementing only the system-sleep callbacks is “good enough” for a battery-sensitive device. It compiles, it boots, suspend and resume both work — and yet the device drains power continuously for the entire time the system is awake, because nothing ever calls the runtime-suspend path. The two subsystems are additive, not interchangeable; skipping one just means you keep paying for the power that subsystem would have saved.

A third pitfall is forgetting that the userspace policy layer — not the kernel — decides when sleep states happen. If a system never suspends despite an idle timeout configured in a desktop environment or an Android power daemon, the bug is very often in that userspace policy layer or in a wakeup source that never got disabled, not in the kernel’s sleep-state machinery itself.

Best Practices

Design for Runtime PM first, even on hardware you also plan to support with sleep states — it is the mechanism that saves power every single second the system is awake, which for most real devices is the vast majority of their operational lifetime. Keep runtime-suspend and system-suspend callbacks logically separate even when their implementations end up sharing code, because the PM core reasons about them as distinct events with distinct guarantees. Always let a device report an accurate runtime status; a driver that lies about its state (claiming to be suspended while still generating interrupts) breaks the direct-complete optimization for the whole subtree above it in the device hierarchy. And never assume a device’s wakeup capability is symmetric between runtime PM and system sleep — the kernel documentation is explicit that a device may legitimately have different wakeup settings for each.

Summary and Key Takeaways

Linux kernel power management is not one mechanism but a deliberate partnership between two: Device Power Management, also called Runtime PM, which continuously adjusts individual devices while the system runs; and System Power Management, the sleep-state machinery that transitions the entire machine at once, triggered by userspace policy the kernel itself never needs to understand. Neither one substitutes for the other — Runtime PM cannot deliver the deep savings of a real sleep state, and sleep states alone leave enormous amounts of power on the table during ordinary, everyday use. They are wired together deliberately: a device already in runtime suspend can be left untouched through an entire system-sleep cycle, which is exactly why understanding both models — and how they hand off to each other through prepare() and complete() — is the correct starting point before writing a single dev_pm_ops callback. The next lectures in this free Linux kernel development course will build the runtime PM callbacks, the sleep callbacks, and wakeup sources on top of the foundation you just learned.

Frequently Asked Questions

What is the difference between Device Power Management and System Power Management in the Linux kernel?

Device Power Management, also called Runtime PM, adjusts the power state of one device at a time while the rest of the system keeps running. System Power Management, also called Sleep States, transitions the entire machine into a low-power state together, and only happens when userspace requests it.

Is Runtime PM the same thing as suspending the whole system?

No. Runtime PM never stops the CPU scheduler or freezes userspace tasks; it only powers individual devices down when they are idle. A full system suspend, by contrast, freezes tasks, disables interrupts for most devices, and can put the entire machine, including memory, into a much deeper low-power state.

Who decides when a Linux system enters a sleep state?

Userspace decides. The kernel exposes the mechanism through files like /sys/power/state, but the policy — whether that is triggered by a closed laptop lid, an idle timeout, a critical battery level, or a direct user action — lives entirely outside the kernel.

What sleep states does the Linux kernel support?

Up to four, depending on platform support: suspend-to-idle (a lightweight, pure-software state), standby (a shallow hardware suspend), suspend-to-RAM (the deep state, sometimes called S3 on ACPI systems), and hibernation, also known as suspend-to-disk.

Where does dev_pm_ops fit into all of this?

struct dev_pm_ops, defined in include/linux/pm.h, is the single structure that holds both categories of callback: the runtime PM trio (runtime_suspend, runtime_resume, runtime_idle) and the system-sleep set (suspend, resume, freeze, thaw, poweroff, restore, and their noirq/late/early variants).

Can a device stay runtime-suspended through an entire system suspend cycle?

Yes. If a device’s prepare() callback indicates it is already runtime-suspended, and all of its child devices are too, the PM core can skip its suspend and resume callbacks entirely and rely only on the complete() callback afterward — an optimization known as direct complete.

Do I need to implement both Runtime PM and system sleep callbacks in every driver?

Not always. Many drivers only implement the runtime PM callbacks and let generic subsystem-level system-sleep handling take care of the rest by resuming the device to full power before suspend and letting runtime PM re-suspend it afterward. Whether you need both depends on the hardware and the subsystem the device belongs to.

Ready to See This in Action?

Move on to the hands-on companion lecture and learn how to observe Device Power Management and System Power Management on a real, running Linux system.

Continue to Hands-On Lecture Browse the Full Course Index

Leave a Reply

Your email address will not be published. Required fields are marked *