Linux Power Domains And Suspend
Lecture 7 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course, covering the Generic Power Domain framework (genpd) and the full system suspend/resume callback sequence.
Linux power domains group devices that share a common power resource, such as a rail or a reference clock, so the kernel can gate that resource on and off for the whole group at once instead of guessing at each device in isolation. This lecture builds a precise, current-kernel mental model of two closely related subjects: the Generic Power Domain framework, known as genpd, which extends per-device runtime power management to a hardware-defined group of devices; and the high-level system suspend and resume sequence that every device driver’s callbacks participate in, whether or not that driver happens to sit inside a power domain. By the end you will know exactly which callback the PM core chooses to run for a given device, in what order, and why the “noirq” and “_late”/”_early” phase names exist at all.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Having gone through the previous Runtime PM Autosuspend lecture in this free Linux kernel development course, since power domains build directly on the runtime PM status (RPM_ACTIVE/RPM_SUSPENDED) and usage-count model introduced there.
- Basic C, including function pointers and struct-based ops tables (struct dev_pm_ops).
- Comfort reading a device tree source fragment: nodes, phandles, and the
compatibleproperty. - Root access on a Linux system or VM for the companion hands-on examples lecture that follows this one.
What A Power Domain Actually Is
Strip away the kernel data structures for a moment and a power domain is simply a hardware fact: a set of devices that cannot be powered down individually because they share something upstream of them, most commonly a power rail fed by a regulator, or a reference clock that several IP blocks derive their own clocks from. If you turn that shared rail or clock off, every device hanging off it loses power or timing at the same instant, whether you intended to suspend all of them or not. A modern SoC might carve its silicon into a dozen or more of these domains: one covering the CPU cluster, one for the GPU, one for the display pipeline, one for a low-power always-on island that keeps a real-time clock and a handful of wakeup-capable peripherals alive even when the rest of the chip is off. Linux power domains are the kernel’s abstraction over exactly this hardware topology, and getting that abstraction right matters because a driver that thinks it fully controls its own device’s power state, without accounting for the domain it lives in, will get the ordering wrong the moment two devices in the same domain need to be powered down together.
From the kernel’s side, a device’s association with its power domain is tracked through the pm_domain field of struct device, a pointer to a struct dev_pm_domain defined in include/linux/pm.h. That structure carries its own struct dev_pm_ops, the same callback table shape used by buses, classes, and device types, plus a handful of domain-specific hooks: start(), called when a consumer of the domain needs to be brought up; detach(), called when a device leaves the domain; activate() and sync(), called around driver probing; and set_performance_state(), used by domains that also scale voltage or frequency rather than only gating power. A power domain can itself be nested inside a parent domain, in which case it is referred to as a sub-domain, mirroring how real silicon nests smaller power islands inside larger ones.
Generic Power Domain (genpd): Solving The SoC Scaling Problem
Before genpd existed, every SoC vendor that needed power-domain support had two unpleasant options: bolt power-domain logic directly onto their platform bus type, duplicating the same “resume this group of devices before suspending this one” logic across every vendor tree, or skip it and leave whole IP blocks powered needlessly. genpd, implemented in drivers/base/power/domain.c and declared in include/linux/pm_domain.h, is the kernel’s answer: a single, reusable implementation of struct dev_pm_domain that any platform can plug into by describing its domains in the device tree (or via firmware interfaces such as SCMI) instead of writing bespoke C code for every SoC family.
The core genpd data structure is struct generic_pm_domain, which represents one power domain and owns a list of the devices and sub-domains attached to it. Structurally, genpd’s biggest contribution is extending the single-device runtime PM model, usage counts, RPM_ACTIVE/RPM_SUSPENDED status, .runtime_suspend()/.runtime_resume() callbacks, up to the level of a whole group. When every device inside a genpd-managed domain has gone runtime-suspended, and no wakeup-capable device inside it needs to stay armed, genpd’s own governor logic can decide the entire domain, and the shared power rail behind it, is safe to gate off. The individual devices never had to coordinate with each other directly; genpd does that coordination on their behalf by aggregating their runtime PM states. This is precisely the problem complex SoC designs run into as they scale: dozens of IP blocks, each with a perfectly correct per-device runtime PM implementation, still waste power if nothing above them ever asks “is it safe to kill the whole rail now?” genpd is that missing layer.
Device Tree Description: Now A YAML Binding
A PM domain provider (typically a power controller node) and its consumer devices are wired together in the device tree using the standard phandle-and-specifier pattern also used for clocks and regulators. A provider declares how many cells its domain specifier needs with #power-domain-cells, zero if the node represents exactly one domain, one or more if a single power-controller node fans out into multiple domains selected by an index. A consumer device then references its domain with a power-domains property.
One update worth calling out explicitly for anyone still working from older references: the device tree binding for this is no longer the old free-form power-domain.txt plain-text document. Like most bindings across the tree, it has been converted to a machine-checkable YAML schema, Documentation/devicetree/bindings/power/power-domain.yaml, validated against the devicetree meta-schema with tools like dt-validate. This is not a cosmetic change: a malformed power-domains property or a missing #power-domain-cells on a provider node now fails schema validation during the build, catching binding mistakes long before they turn into a device that silently never powers on in the field.
// Power domain provider: one power-controller node exposing two domains
power: power-controller@12340000 {
compatible = "vendor,soc-power-controller";
reg = ;
#power-domain-cells = ;
};
// Consumer: a device that lives inside domain index 0 of that controller
display_pipe: display@1a100000 {
compatible = "vendor,soc-display";
reg = ;
power-domains = ;
};
// Consumer: a device that lives inside domain index 1 of the same controller
gpu: gpu@1a200000 {
compatible = "vendor,soc-gpu";
reg = ;
power-domains = ;
};
Notice that the consumer side is deliberately generic: display_pipe and gpu say nothing about power domains beyond the single power-domains phandle. All of the domain-specific power sequencing logic lives in the provider’s driver, which registers itself with genpd core using of_genpd_add_provider_simple() or of_genpd_add_provider_onecell() depending on whether it exposes one domain or several.
Why Power Domain Callbacks Take Precedence
Here is the rule that matters most in practice: when the PM core needs to run a suspend or resume callback for a device, it does not simply call the bus type’s callback. It follows a strict precedence order, checking each level in turn and stopping at the first one that is populated:
Callback Selection Precedence For Any Given Device
In other words, if a device’s pm_domain pointer is non-NULL, the PM core executes the callback from dev->pm_domain->ops instead of the bus type’s, the class’s, or the device type’s callback, full stop, for every single phase: suspend, suspend_late, suspend_noirq, and their resume counterparts alike. This is exactly what lets a platform with power domains reuse an existing bus type (say, the platform bus) unmodified: the bus doesn’t need to know anything about domains at all, because genpd’s dev_pm_domain callbacks simply intercept the call before the bus ever sees it. The domain-level callback is, of course, still free to turn around and invoke the underlying driver’s own dev->driver->pm methods internally, which is exactly what genpd does after handling its own domain-wide bookkeeping, so a driver author writing an ordinary suspend/resume pair generally does not need to know or care whether a power domain sits above their device at all.
The Full System Suspend/Resume Callback Chain
Suspending or resuming the whole system is not one callback per device, it is several, run in strict phases across every device before the next phase begins. For a normal suspend-to-idle, standby, or suspend-to-RAM transition (hibernation adds extra phases not covered here), the chain looks like this:
System Suspend/Resume Callback Chain
Two ordering guarantees hold throughout this entire chain. First, within any single phase the device hierarchy is walked bottom-up when suspending (children before parents, since a parent may be a bridge or bus controller the child still needs to talk through) and top-down when resuming (parents before children, for the same reason in reverse); prepare and complete are themselves special-cased to run top-down and bottom-up respectively, which is worth remembering because it is the opposite of what you might guess by analogy with suspend/resume. Second, every device finishes one phase before any device starts the next phase; the PM core does not begin issuing suspend_late calls to some devices while others are still in suspend.
What “noirq” And “_late” / “_early” Actually Mean
These names are not arbitrary. The suffix _noirq tells you precisely when that phase runs relative to interrupt handling: by the time suspend_noirq callbacks start, the kernel has already masked interrupt handling system-wide (aside from lines explicitly flagged IRQF_NO_SUSPEND), so a driver’s own interrupt handler is guaranteed not to fire while its suspend_noirq method is running. That guarantee exists for a concrete reason: bus types that let multiple devices share an interrupt vector, PCI being the textbook case, need to know that no other device on that shared line can generate an interrupt mid-transition and confuse a handler that is racing the very suspend logic trying to quiesce it. The matching resume_noirq phase runs before interrupts are unmasked again, so a driver gets a chance to put its device into a state where it can correctly recognize whether it, specifically, is the source of any pending shared interrupt, before that interrupt handling resumes for real.
The _late and _early suffixes exist for a different, complementary reason: splitting “quiesce” from “save state” (and the mirror image, “prepare” from “fully resume”) for devices where doing everything in one callback would be awkward. suspend_late always runs after runtime power management has been disabled for that device, which is a meaningful guarantee: by the time your suspend_late callback executes, nothing can race it with a runtime PM transition for the same device, and for many drivers the very same function pointer used for runtime_suspend is reused directly for suspend_late, since a device that is already runtime-suspended and one that is being system-suspended usually need to end up in the identical low-power register state. resume_early mirrors this on the way back up, before the full resume phase.
High-Level Suspend And Resume Stages
The device-level callback chain above is itself nested inside a small number of higher-level kernel stages that run once per suspend/resume cycle, not once per device. Understanding what each stage is actually for, not just its name, makes it much easier to reason about where a given bug belongs.
| Stage | Direction | Purpose |
|---|---|---|
| Filesystem sync | Suspend | Flush dirty pages to backing storage before anything is frozen, so an unexpected failure partway through suspend does not corrupt on-disk state. |
| PM notifier chain: PM_SUSPEND_PREPARE | Suspend | Give subsystems one last chance to act while user space is still fully running and before any task is frozen. |
| Freeze tasks (freeze_processes()) | Suspend | Freeze freezable user-space processes and kernel threads so nothing can race the device suspend phases that follow; all device-level callbacks run only after this completes. |
| Suspend devices, phase one: dpm_suspend_start() (prepare + suspend) | Suspend | Run the ->prepare and ->suspend device callback chain described above across every registered device. |
| Suspend devices, phase two: dpm_suspend_late() + dpm_suspend_noirq() | Suspend | Finish quiescing every device and disable interrupt handling, completing the suspend_late and suspend_noirq phases. |
| Disable non-boot CPUs | Suspend | Take secondary CPUs offline so only the boot CPU is active when the platform actually enters its low-power mode, since most SoC sleep states are undefined with multiple cores still running. |
| syscore_suspend() | Suspend | Suspend core kernel infrastructure that has no struct device backing it at all, interrupt controllers, timekeeping, some SoC-level clock or PMU blocks, immediately before the final low-power entry. |
| syscore_resume() | Resume | Mirror image of syscore_suspend(); restores this core infrastructure first, before any struct device callback runs, because everything else depends on it. |
| Enable non-boot CPUs | Resume | Bring secondary CPUs back online before device resume callbacks that may depend on multi-core scheduling or per-CPU state. |
| Resume devices, phase one: dpm_resume_noirq() + dpm_resume_early() | Resume | Run the resume_noirq and resume_early device callback chain, still with interrupts masked for the noirq portion. |
| Resume devices, phase two: dpm_resume_end() (resume + complete) | Resume | Run the resume and complete device callback chain, returning every device to full operation. |
| Thaw tasks (thaw_processes()) | Resume | Unfreeze user-space processes and kernel threads, letting the system resume normal scheduling. |
| PM notifier chain: PM_POST_SUSPEND | Resume | Final notification that the system is fully back to the working state, mirroring PM_SUSPEND_PREPARE. |
Common Mistakes And Troubleshooting
- Assuming a driver’s own .suspend() will run even though a power domain is attached. If dev->pm_domain is set, the domain’s ->suspend() runs instead of the bus type’s, not in addition to it at the same precedence level; a driver callback that never gets invoked is often a sign the domain layer is intercepting it, which is expected behavior, not a bug.
- Doing register I/O from a suspend_noirq or resume_noirq callback that depends on an interrupt completing. Interrupt handling is disabled system-wide during the noirq phases, so any operation that would normally complete via an IRQ (a completion-based bus transaction, for instance) must be reworked to poll or must be moved to the suspend_late/resume_early phase instead.
- Forgetting that suspend and resume walk the device tree in opposite directions. Suspend runs bottom-up (children first) while resume runs top-down (parents first); code that implicitly assumes symmetry here will misorder itself relative to a parent bus controller.
- Confusing #power-domain-cells on a provider with the size of the power-domains phandle on a consumer. A consumer’s power-domains property must supply exactly as many argument cells as the referenced provider’s #power-domain-cells declares; a mismatch fails devicetree schema validation against power-domain.yaml before the board even boots.
- Writing a suspend_late callback that does not check the return value of the domain’s own quiescing logic. Because power domain callbacks take precedence, a device driver debugging “why didn’t my suspend_late run” should check whether a domain sits above it and whether that domain’s own callback is failing or returning early.
Best Practices
- Reuse the same function for a device’s runtime_suspend and its suspend_late (and runtime_resume / resume_early) wherever the low-power register state genuinely matches, using SET_LATE_SYSTEM_SLEEP_PM_OPS() to wire this up in one line rather than hand-rolling separate near-duplicate callbacks.
- Keep suspend_noirq/resume_noirq callbacks as short and non-blocking as possible, since interrupt handling for the whole system is masked while they run.
- Describe power domains in the device tree using #power-domain-cells and power-domains exactly as specified in power-domain.yaml, and validate with dt-validate before assuming a binding is correct.
- Let genpd’s own governor decide when to gate a domain off rather than trying to synchronize sibling devices’ power state manually from within a single driver.
- When debugging an unexpected callback ordering, check dev->pm_domain first; power domain callbacks always take precedence over type, class, and bus callbacks for every phase in the chain.
Summary And Key Takeaways
Linux power domains model a real hardware constraint, a shared power rail or clock, as a set of devices that must be gated together, and genpd is the reusable kernel framework that lets a platform describe that grouping once in the device tree, using the now-YAML power-domain.yaml binding, rather than reimplementing group-power logic per SoC. genpd’s real contribution is extending the per-device runtime PM model (usage counts, RPM_ACTIVE/RPM_SUSPENDED) up to the level of an entire domain, so an SoC with dozens of IP blocks does not waste power just because nothing above the individual drivers ever asks whether the whole domain is safe to power down. Whenever a device’s pm_domain pointer is set, that domain’s callbacks take strict precedence over type, class, and bus callbacks for every phase of the system suspend/resume chain: prepare, suspend, suspend_late, suspend_noirq on the way down, and resume_noirq, resume_early, resume, complete on the way back up, with the noirq phases running while interrupts are masked system-wide and the _late/_early phases running after runtime PM has already been disabled for that device. With this chain and the role of power domains both clear, the companion hands-on examples lecture puts it into practice: inspecting real power domains on a running system and building an original driver that demonstrates suspend_late/resume_early ordering with real dmesg output.
Frequently Asked Questions
What exactly makes a set of devices a power domain?
Sharing a power resource that cannot be gated per device, most commonly a power rail or a reference clock. If turning that resource off necessarily affects every device attached to it at once, those devices form a power domain, regardless of whether Linux happens to model that relationship in software yet.
Does every platform need genpd?
No. genpd is only relevant to platforms where multiple devices genuinely share a gateable power resource and where reusing the same driver PM callbacks across different domain configurations is valuable. A simple SoC where every device has its own independent power switch has no real use for it.
Why is the device tree binding now a .yaml file instead of the old .txt format?
Documentation/devicetree/bindings/power/power-domain.yaml replaces the old free-form power-domain.txt document as part of the kernel-wide move to machine-checkable devicetree schemas. It lets tools like dt-validate catch binding mistakes, a missing #power-domain-cells, a malformed power-domains phandle, during the build rather than at runtime on real hardware.
If a device has both a bus type PM callback and a power domain, which one runs?
The power domain’s callback runs, exclusively, for every phase where dev->pm_domain->ops provides one. Power domain callbacks always take precedence over type, class, and bus callbacks; the bus type’s callback for that phase is simply not invoked at all.
Why can’t a suspend_noirq callback wait on an interrupt?
Because interrupt handling is disabled system-wide (except for lines marked IRQF_NO_SUSPEND) for the entire duration of the suspend_noirq and resume_noirq phases. Any operation whose completion normally depends on an IRQ firing must be reworked to poll, or moved earlier into the suspend_late phase where interrupts are still enabled.
What is the difference between suspend_late and suspend_noirq?
suspend_late runs after runtime PM has been disabled for the device but while interrupts are still enabled system-wide; suspend_noirq runs afterward, with interrupts masked. Many drivers put “save the last bit of register state” logic in suspend_late and reserve suspend_noirq for buses that specifically need to worry about shared interrupt vectors.
Does resume always undo suspend in the exact reverse order per device?
Within the phases themselves, yes: resume_noirq mirrors suspend_noirq, resume_early mirrors suspend_late, resume mirrors suspend, complete mirrors prepare. But the device-tree walk direction flips: suspend walks bottom-up (children before parents) while resume walks top-down (parents before children).
Ready To See It In Action?
You now understand what a Linux power domain is, why genpd exists, and the exact order the system suspend/resume callback chain runs in. The next lecture in this free Linux kernel development course puts this into practice with real sysfs and debugfs inspection plus an original demo driver.
Try The Hands-On Examples Back To Course Index