Linux Suspend Resume Sequence Guide
Hands-on companion to the previous power domains lecture: inspect genpd on a real system, then build and load ep_pdemo, an original platform driver demonstrating suspend_late/resume_early ordering, as part of this free Linux kernel development course.
ep_pdemo.c
The previous lecture built the conceptual model of Linux power domains and the system suspend/resume sequence; this lecture puts both in front of you on a real, running kernel. You will first inspect power domain state the way a driver engineer actually does during bring-up, through the genpd debugfs summary and a device tree power-domains property, and then build ep_pdemo, an original, minimal platform driver whose entire purpose is to make the suspend_late and resume_early phases visible in dmesg during a real suspend/resume cycle, so the ordering from the previous lecture stops being an abstract diagram and becomes something you can watch happen.
What You Will Learn
Prerequisites
- Read the companion explanation lecture, “Linux Power Domains And Suspend”, first, this page assumes you already know what genpd is and what the suspend_late/suspend_noirq/resume_early/resume_noirq phases mean.
- Root access on a Linux system or VM where you can build and load kernel modules and trigger a real suspend (a VM with suspend-to-idle support, or a physical board, both work).
- Basic familiarity with the platform_driver structure (probe, remove, of_device_id) from earlier lectures in this series.
Inspecting Power Domains On A Running System
Before writing any code, it is worth knowing how to look at power domain state that already exists on a board you did not write the platform code for. Two places to look: the device tree source that described the domain at build time, and the live genpd debugfs summary that reflects current runtime state.
First, find the power-domains property for a device of interest directly in its device tree source file. This is a plain grep against the .dts/.dtsi files in arch/<your-arch>/boot/dts, no board access required:
$ grep -B2 -A4 "power-domains" arch/arm64/boot/dts/vendor/example-soc.dtsi
display_pipe: display@1a100000 {
compatible = "vendor,soc-display";
reg = ;
power-domains = ;
status = "okay";
};
Here, display_pipe is a consumer of domain index 0 provided by pd_controller. Cross-check that the provider node’s #power-domain-cells matches the number of arguments used in the phandle (one cell here, since one index argument follows the phandle):
$ grep -B1 -A4 "pd_controller:" arch/arm64/boot/dts/vendor/example-soc.dtsi
pd_controller: power-controller@12340000 {
compatible = "vendor,soc-power-controller";
reg = ;
#power-domain-cells = ;
};
Since the binding is now the machine-checkable power-domain.yaml schema rather than free-form text, you can validate a compiled device tree blob against it directly instead of eyeballing the syntax:
$ dt-validate -s Documentation/devicetree/bindings/power/power-domain.yaml \
arch/arm64/boot/dts/vendor/example-soc.dtb
Second, look at live runtime state through genpd’s debugfs interface, available once debugfs is mounted on a kernel built with CONFIG_PM_GENERIC_DOMAINS and CONFIG_DEBUG_FS:
$ sudo mount -t debugfs none /sys/kernel/debug 2>/dev/null
$ cat /sys/kernel/debug/pm_genpd/pm_genpd_summary
domain status children performance
------------------------------------------------------------------------------
pd-display on
/devices/platform/soc/1a100000.display active
pd-gpu off-0
/devices/platform/soc/1a200000.gpu suspended
pd-always-on on (always on)
The status column tells you whether genpd currently believes it is safe to keep the domain’s shared power rail gated on or off, and the indented lines under each domain list the devices genpd is aggregating runtime PM state from to make that decision. A domain showing “off” with every child “suspended” is genpd doing exactly the job described in the previous lecture: nothing under it needs power, so the shared resource above all of them has been gated off, without any of those individual drivers coordinating with each other directly.
Finally, remember that a device’s own runtime PM sysfs attribute, introduced in the runtime PM lectures earlier in this series, is still meaningful even when a power domain sits above it:
$ cat /sys/devices/platform/soc/1a200000.gpu/power/runtime_status
suspended
Meet ep_pdemo: A Suspend_late/Resume_early Ordering Demo
About ep_pdemo
ep_pdemo is an original, deliberately minimal platform driver built for this course. It owns no real hardware registers, an original fictional “saved_state” value stands in for a device register that must be captured before power is removed, so the entire example is about the timing of the callbacks, not about any particular chip. It implements exactly one PM callback pair, suspend_late and resume_early, wired up with the modern SET_LATE_SYSTEM_SLEEP_PM_OPS() macro, and logs a dev_info() line naming each stage as it runs. To make it usable on any bench system without a matching device tree entry, the module registers its own platform_device internally at load time.
Full Driver Source: ep_pdemo.c
The suspend_late/resume_early pair and the macro that wires them into struct dev_pm_ops are highlighted in the walkthrough after the listing.
// SPDX-License-Identifier: GPL-2.0
/*
* ep_pdemo.c - Minimal platform driver demonstrating the suspend_late /
* resume_early phase of the Linux system suspend callback chain, for
* EmbeddedPathashala's free Linux kernel development course.
*
* This is an original teaching example. ep_pdemo does not correspond to
* any real shipping part number or vendor IP block.
*/
#include <linux/module.h>
#include <linux/platform_device.h>
#include <linux/pm.h>
#include <linux/err.h>
struct ep_pdemo_data {
struct device *dev;
u32 saved_state; /* stand-in for a register snapshot */
};
/* ---- suspend_late / resume_early pair -------------------------- */
static int ep_pdemo_suspend_late(struct device *dev)
{
struct ep_pdemo_data *data = dev_get_drvdata(dev);
dev_info(dev, "suspend_late: runtime PM already disabled, IRQs still enabled here\n");
/* Fictional register snapshot -- stands in for real hardware state */
data->saved_state = 0xcafef00d;
dev_info(dev, "suspend_late: state saved (0x%08x), device now quiesced\n",
data->saved_state);
return 0;
}
static int ep_pdemo_resume_early(struct device *dev)
{
struct ep_pdemo_data *data = dev_get_drvdata(dev);
dev_info(dev, "resume_early: restoring saved state 0x%08x\n",
data->saved_state);
dev_info(dev, "resume_early: device ready, waiting for full resume phase\n");
return 0;
}
/*
* SET_LATE_SYSTEM_SLEEP_PM_OPS() wires suspend_late/resume_early into
* .suspend_late/.resume_early AND reuses the same two functions for
* .freeze_late/.thaw_early and .poweroff_late/.restore_early, so this
* driver behaves identically for suspend-to-RAM and hibernation without
* writing three near-duplicate callback pairs by hand.
*/
static const struct dev_pm_ops ep_pdemo_pm_ops = {
SET_LATE_SYSTEM_SLEEP_PM_OPS(ep_pdemo_suspend_late, ep_pdemo_resume_early)
};
/* ---- probe() / remove() ----------------------------------------- */
static int ep_pdemo_probe(struct platform_device *pdev)
{
struct ep_pdemo_data *data;
data = devm_kzalloc(&pdev->dev, sizeof(*data), GFP_KERNEL);
if (!data)
return -ENOMEM;
data->dev = &pdev->dev;
platform_set_drvdata(pdev, data);
dev_info(&pdev->dev, "ep_pdemo probed, ready to observe suspend/resume ordering\n");
return 0;
}
static void ep_pdemo_remove(struct platform_device *pdev)
{
dev_info(&pdev->dev, "ep_pdemo removed\n");
}
static const struct of_device_id ep_pdemo_of_match[] = {
{ .compatible = "ep,pdemo" },
{ }
};
MODULE_DEVICE_TABLE(of, ep_pdemo_of_match);
static struct platform_driver ep_pdemo_driver = {
.driver = {
.name = "ep_pdemo",
.of_match_table = ep_pdemo_of_match,
.pm = pm_ptr(&ep_pdemo_pm_ops),
},
.probe = ep_pdemo_probe,
.remove = ep_pdemo_remove,
};
/*
* No device tree entry is required on a bench system: this module
* registers its own platform_device by name at load time, and unwinds
* it cleanly on unload.
*/
static struct platform_device *ep_pdemo_pdev;
static int __init ep_pdemo_init(void)
{
int ret;
ret = platform_driver_register(&ep_pdemo_driver);
if (ret)
return ret;
ep_pdemo_pdev = platform_device_register_simple("ep_pdemo", -1, NULL, 0);
if (IS_ERR(ep_pdemo_pdev)) {
platform_driver_unregister(&ep_pdemo_driver);
return PTR_ERR(ep_pdemo_pdev);
}
return 0;
}
module_init(ep_pdemo_init);
static void __exit ep_pdemo_exit(void)
{
platform_device_unregister(ep_pdemo_pdev);
platform_driver_unregister(&ep_pdemo_driver);
}
module_exit(ep_pdemo_exit);
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("Minimal platform driver demonstrating suspend_late/resume_early ordering");
MODULE_LICENSE("GPL");
Where ep_pdemo Sits In The Suspend/Resume Sequence
Build And Test Walkthrough
Save the listing above as ep_pdemo.c alongside this Makefile in an empty directory:
obj-m += ep_pdemo.o
KDIR := /lib/modules/$(shell uname -r)/build
PWD := $(shell pwd)
all:
$(MAKE) -C $(KDIR) M=$(PWD) modules
clean:
$(MAKE) -C $(KDIR) M=$(PWD) clean
Build against your running kernel’s headers:
$ make
CC [M] ep_pdemo.o
MODPOST ep_pdemo.mod.c
CC [M] ep_pdemo.mod.o
LD [M] ep_pdemo.ko
Load it. Because the module registers its own platform_device internally, probe() runs immediately with no device tree entry needed:
$ sudo insmod ep_pdemo.ko
$ dmesg | tail -n 1
[ 310.002211] ep_pdemo ep_pdemo.0: ep_pdemo probed, ready to observe suspend/resume ordering
Trigger a real suspend/resume cycle (suspend-to-idle is the safest to test inside a VM; use mem instead of freeze on hardware that supports suspend-to-RAM):
$ cat /sys/power/state
freeze mem
$ echo freeze | sudo tee /sys/power/state
Reading dmesg after the system wakes back up (a timed wakeup, a key press, or any configured wakeup source will bring it back) shows ep_pdemo’s two log lines sitting exactly where the previous lecture’s diagram says suspend_late and resume_early belong, wrapped by the higher-level suspend stages:
$ dmesg | tail -n 20
[ 340.010102] PM: suspend entry (s2idle)
[ 340.030044] Filesystems sync: 0.011 seconds
[ 340.032511] Freezing user space processes
[ 340.033602] Freezing user space processes completed (elapsed 0.001 seconds)
[ 340.033610] Freezing remaining freezable tasks
[ 340.034201] Freezing remaining freezable tasks completed (elapsed 0.001 seconds)
[ 340.040233] printk: Suspending console(s) (use no_console_suspend to debug)
[ 340.098871] ep_pdemo ep_pdemo.0: suspend_late: runtime PM already disabled, IRQs still enabled here
[ 340.098879] ep_pdemo ep_pdemo.0: suspend_late: state saved (0xcafef00d), device now quiesced
[ 340.110044] ACPI: EC: interrupt blocked
[ 340.130511] Disabling non-boot CPUs ...
[ 340.145102] Successfully transitioned to state s2idle
[ 340.160044] Timekeeping suspended for 5.002 seconds
[ 340.170022] Enabling non-boot CPUs ...
[ 340.190019] ep_pdemo ep_pdemo.0: resume_early: restoring saved state 0xcafef00d
[ 340.190026] ep_pdemo ep_pdemo.0: resume_early: device ready, waiting for full resume phase
[ 340.210099] OOM killer enabled.
[ 340.210102] Restarting tasks ... done.
[ 340.220011] PM: suspend exit
Notice that ep_pdemo’s suspend_late line appears after console suspension has already begun but before “Disabling non-boot CPUs”, and its resume_early line appears right after “Enabling non-boot CPUs” but well before user-space tasks are restarted, exactly the ordering established in the previous lecture’s callback chain diagram.
Unload the module and confirm clean teardown:
$ sudo rmmod ep_pdemo
$ dmesg | tail -n 1
[ 400.004112] ep_pdemo ep_pdemo.0: ep_pdemo removed
Common Mistakes And Troubleshooting
- pm_genpd_summary does not exist under /sys/kernel/debug. Confirm debugfs is mounted (
mount | grep debugfs) and that the running kernel was built with CONFIG_PM_GENERIC_DOMAINS; on a platform with no genpd providers registered, the pm_genpd directory itself will not appear at all. - ep_pdemo’s dev_info() lines never show up in dmesg after suspend. Check that the module actually loaded (lsmod | grep ep_pdemo) and that /sys/power/state accepts the state you wrote; some virtualized environments only support “freeze” (suspend-to-idle), not “mem”.
- dt-validate reports an error even though the .dts file looks correct. Re-check that #power-domain-cells on the provider matches the number of specifier cells used in every consumer’s power-domains property; this is the single most common power-domain.yaml validation failure.
- Confusing pm_ptr() with pm_sleep_ptr(). pm_ptr() is used on the .pm struct pointer inside struct platform_driver and compiles it out entirely when CONFIG_PM is disabled; pm_sleep_ptr(), used internally by SET_LATE_SYSTEM_SLEEP_PM_OPS(), compiles a callback out only when CONFIG_PM_SLEEP is disabled while runtime PM support can remain.
- Expecting suspend_late to run while the driver still holds a runtime PM reference. The PM core guarantees runtime PM has already been disabled for the device by the time suspend_late runs; a driver relying on pm_runtime_get()/put() balance inside suspend_late is solving a problem that phase was specifically designed to avoid.
Best Practices
- When a callback genuinely does the same work for suspend-to-RAM and hibernation, use SET_LATE_SYSTEM_SLEEP_PM_OPS() (or the full SET_SYSTEM_SLEEP_PM_OPS()) rather than filling in six near-identical function pointers by hand.
- Log a dev_info() line at the start of every PM callback during bring-up on new hardware; it is the fastest way to confirm the suspend/resume sequence is actually hitting the phases you expect, exactly as ep_pdemo does here.
- Always check debugfs pm_genpd output before assuming a device’s power state is driven purely by its own driver; a domain above it may be the real reason it never fully powers down.
- Validate any new power-domains device tree binding with dt-validate against power-domain.yaml as part of your normal build, not as an afterthought when hardware misbehaves.
- Prefer testing suspend/resume flows with suspend-to-idle (freeze) in a VM first; it exercises the entire device-level callback chain without requiring real platform firmware support for deeper sleep states.
Summary And Key Takeaways
This lecture turned the previous conceptual model of Linux power domains and the suspend/resume sequence into something observable: the genpd debugfs summary and a device tree power-domains property show you real domain state on a running system, while ep_pdemo, an original driver built around a single suspend_late/resume_early callback pair wired up with SET_LATE_SYSTEM_SLEEP_PM_OPS(), makes the exact ordering from the callback chain diagram show up as real dmesg lines during a real suspend/resume cycle. The next lecture in this free Linux kernel development course continues the power management series with wakeup sources and how devices request the system leave a sleep state.
Frequently Asked Questions
Why does ep_pdemo only implement suspend_late/resume_early and not suspend/resume?
To isolate exactly one pair of phases so their position in the suspend/resume sequence is unambiguous in dmesg. A real driver would typically also implement plain suspend/resume for work that does not need to wait until runtime PM is disabled.
Why register the platform_device from inside the module instead of using a device tree entry?
So the example works on any bench Linux system or VM without needing to modify and reflash a device tree blob. On real hardware, the .of_match_table entry (compatible = “ep,pdemo”) would bind automatically to a matching devicetree node instead.
What does SET_LATE_SYSTEM_SLEEP_PM_OPS() actually expand to?
It populates .suspend_late and .resume_early with the two functions given, and additionally reuses them for .freeze_late/.thaw_early and .poweroff_late/.restore_early, so the same two functions cover suspend-to-RAM, suspend-to-idle, and hibernation without separate hibernation-specific callbacks.
Is /sys/kernel/debug/pm_genpd/pm_genpd_summary available on every Linux system?
Only if debugfs is mounted and the kernel has CONFIG_PM_GENERIC_DOMAINS enabled with at least one genpd provider registered. Many desktop x86 systems use ACPI-based power management instead of genpd and will not show this path at all.
What happens if suspend_late returns a nonzero error code?
The PM core aborts the suspend transition and unwinds by resuming the devices that were already suspended, in the same way an error from any other suspend-phase callback does.
Continue Your Linux Kernel Power Management Journey
You have now watched the Linux suspend resume sequence happen in real dmesg output, backed by an original driver you built yourself, as part of this free Linux kernel development course. Continue to the next lecture in the Kernel Power Management series to cover wakeup sources.
Continue The Power Management Series Back To Course Index