Linux System Sleep PM Callbacks
Lecture 8 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course, covering the six generic system sleep callbacks and the dev_pm_ops convenience macros that wire a driver into suspend, suspend-to-idle, and hibernation.
Linux system sleep callbacks are the contract every driver signs up for the moment it wants to survive a suspend, a suspend-to-idle transition, or a full hibernation cycle without leaving hardware in an inconsistent state. This lecture builds a precise, current-kernel understanding of exactly six callbacks, suspend, resume, freeze, thaw, poweroff, and restore, and the small family of macros that populate them into struct dev_pm_ops without forcing you to hand-write six near-identical function pointers every time. Where the previous lecture in this free Linux kernel development course walked the full phase-by-phase callback chain and the role of power domains, this one zooms into a single, sharply scoped question: given a device’s actual sleep requirements, which of these six Linux system sleep callbacks does it need to implement, and which macro should populate them.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Having gone through the previous Power Domains And Suspend lecture and its companion hands-on lecture in this free Linux kernel development course, since this lecture assumes you already know the full prepare/suspend/suspend_late/suspend_noirq phase chain and where a driver’s own callbacks sit inside it.
- Basic C, including function pointers and struct-based ops tables (struct dev_pm_ops).
- Comfort with the platform_driver structure (probe, remove, of_device_id) from earlier lectures in this series.
- Root access on a Linux system or VM for the companion hands-on examples lecture that follows this one.
Why System Sleep And Runtime PM Are Related But Distinct
It is tempting, once a driver already has working runtime_suspend() and runtime_resume() callbacks, to assume system sleep support is free, just point .suspend at runtime_suspend and be done with it. That assumption breaks down in two specific ways. First, runtime PM is a per-device, cooperative decision: the PM core only calls runtime_suspend() when nothing is actively using the device and its usage count has dropped to zero, and it is free to defer or refuse the transition if the device is busy. System sleep is not cooperative in that sense at all: when the administrator, a lid switch, an idle timer, or an ACPI event triggers a system-wide sleep transition, every registered device gets a suspend callback whether it is idle or not, because the entire machine is going down together. A device that is mid-transfer when suspend begins does not get to say “not yet”; the driver’s job is to make that transfer safe to interrupt, not to negotiate the timing.
Second, and just as important, the two mechanisms are gated by different Kconfig options. CONFIG_PM is the umbrella option that enables power management infrastructure generally, including runtime PM. CONFIG_PM_SLEEP is a narrower option layered on top of it, specifically for system-wide sleep transitions, suspend-to-RAM, suspend-to-idle, standby, and hibernation. A kernel can be built with CONFIG_PM=y and CONFIG_PM_SLEEP=n: this describes a real class of embedded and server-like systems that benefit from runtime clock- and power-gating on idle peripherals but never suspend the whole system at all. On such a kernel, every one of the Linux system sleep callbacks in this lecture compiles away to nothing, while runtime PM keeps working normally. That split is precisely why the macros this lecture covers are conditioned on CONFIG_PM_SLEEP specifically, not on CONFIG_PM.
Where the two mechanisms do legitimately meet is in the suspend_late and resume_early phases (covered in the previous lecture), where a device’s runtime PM state has already been forced to the suspended-equivalent register configuration, making it safe, and common practice, to literally reuse the same function pointer for both. That is a deliberate optimization available at one specific phase boundary, not a reason to conflate the two subsystems everywhere else.
The Six Generic Sleep Callbacks
struct dev_pm_ops, declared in include/linux/pm.h, exposes six callbacks whose entire purpose is system sleep, arranged as three suspend/resume-style pairs. Each pair answers a different question about what “the system is going away” actually means for that transition.
.suspend / .resume: RAM-Preserving Sleep
.suspend() runs before the system enters a sleep state in which the contents of main memory are preserved, suspend-to-RAM, suspend-to-idle (s2idle), or standby. Because RAM itself survives untouched, the driver’s own data structures survive with it; .suspend() has exactly one job, make the device quiescent so it performs no further I/O or DMA once the subsystem-level callback returns. It does not need to serialize anything to persistent storage. .resume() undoes exactly that: it brings the device back to a working state, and depending on platform and subsystem there are typically no restrictions on clocks or other resources being available by the time it runs. A driver whose hardware genuinely loses all register state even during this “RAM-preserving” sleep, common on SoCs that cut power to a peripheral’s rail during suspend-to-RAM even though system RAM stays powered, still needs to snapshot and restore that register state here; the “RAM-preserving” guarantee is about main memory, not about every peripheral’s own power rail.
.freeze / .thaw: Building And Rolling Back A Hibernation Image
.freeze() is hibernation-specific and runs while the kernel is constructing an in-memory snapshot that will shortly be written to persistent storage. It is deliberately analogous to .suspend() but with one critical restriction: it must not power down the device or arm it to signal a wakeup event. The reason is structural, at the point .freeze() runs, the machine has not actually lost power yet; it is still the same running “boot kernel” that will, moments later, need those very devices working again to actually write the image to disk. Most subsystems expect .freeze() to save whatever register state .poweroff() will later need, since power genuinely will be cut afterward. .thaw() is the rollback: it runs either after the image has been successfully created (so the boot kernel can carry on and write it out) or after a failed attempt to create one, undoing exactly what .freeze() did so the device is usable again immediately.
.poweroff / .restore: Saving And Restoring Across A Real Power Cycle
.poweroff() runs after the hibernation image has already been safely written to disk, at the point where the machine is genuinely about to be powered off. It performs the same quiescing work as .suspend(), but it does not need to save register state again, .freeze() already captured whatever .restore() will need, because by definition the register contents .poweroff() would see are identical to what .freeze() already recorded moments earlier. .restore() is the true mirror image of .resume(), but under a materially different assumption: the device was actually powered off in between, there is no clock or rail that quietly survived, so .restore() must reprogram the hardware completely from the values saved during .freeze(), never assuming any register retained its prior contents. This runs from an entirely fresh boot of the “restore kernel,” not a continuation of the kernel that hibernated.
Hibernation’s Three Sub-Transitions
The Convenience Macros That Populate struct dev_pm_ops
Writing all six fields by hand for every driver would mean six lines of boilerplate even for the overwhelmingly common case where suspend-to-RAM and hibernation genuinely need identical driver logic. The kernel solves this with a small family of macros in include/linux/pm.h, verified here directly against current mainline source rather than assumed from older documentation.
SET_SYSTEM_SLEEP_PM_OPS: One Pair, All Six Fields
This is still present in current mainline, still gated by CONFIG_PM_SLEEP, and expands to fill all six standard-phase fields (not the _late or _noirq variants) from a single suspend/resume function pair:
#ifdef CONFIG_PM_SLEEP
#define SET_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn) \
SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn)
#else
#define SET_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn)
#endif
/* Where SYSTEM_SLEEP_PM_OPS itself expands to: */
#define SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn) \
.suspend = pm_sleep_ptr(suspend_fn), \
.resume = pm_sleep_ptr(resume_fn), \
.freeze = pm_sleep_ptr(suspend_fn), \
.thaw = pm_sleep_ptr(resume_fn), \
.poweroff = pm_sleep_ptr(suspend_fn), \
.restore = pm_sleep_ptr(resume_fn),
Use it when suspend_fn and resume_fn genuinely produce correct behavior for freeze/thaw and poweroff/restore too, true for most drivers whose only real job is “quiesce and snapshot register state” versus “reprogram and go” regardless of which specific sleep path triggered it.
SET_NOIRQ_SYSTEM_SLEEP_PM_OPS: For The noirq Phase
Also still current, still CONFIG_PM_SLEEP-gated, and fills the six _noirq fields instead:
#define SET_NOIRQ_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn) \
NOIRQ_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn)
/* NOIRQ_SYSTEM_SLEEP_PM_OPS fills: */
.suspend_noirq, .resume_noirq, .freeze_noirq, .thaw_noirq,
.poweroff_noirq, .restore_noirq
Reach for this when your device sits on a bus where interrupt lines are shared, or otherwise needs guaranteed non-racing access to hardware while system-wide interrupt handling is masked, exactly the noirq phase behavior covered in the previous lecture.
SET_LATE_SYSTEM_SLEEP_PM_OPS: For The Late/Early Phase
Fills the six _late/_early fields, for logic that specifically needs to run after runtime PM has already been disabled for the device:
#define SET_LATE_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn) \
LATE_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn)
/* LATE_SYSTEM_SLEEP_PM_OPS fills: */
.suspend_late, .resume_early, .freeze_late, .thaw_early,
.poweroff_late, .restore_early
This is the macro most often paired with reusing your own runtime_suspend()/runtime_resume() functions directly, since by the time suspend_late runs, the PM core guarantees runtime PM transitions for that device cannot race it.
SIMPLE_DEV_PM_OPS And Its Current Replacement, DEFINE_SIMPLE_DEV_PM_OPS
Here is the one place where treating the old book chapter this lecture is based on as gospel would actively mislead you. SIMPLE_DEV_PM_OPS still compiles on current mainline, but the kernel’s own header now marks it explicitly:
/* Deprecated. Use DEFINE_SIMPLE_DEV_PM_OPS() instead. */
#define SIMPLE_DEV_PM_OPS(name, suspend_fn, resume_fn) \
const struct dev_pm_ops __maybe_unused name = { \
SET_SYSTEM_SLEEP_PM_OPS(suspend_fn, resume_fn) \
}
The problem SIMPLE_DEV_PM_OPS has is the __maybe_unused annotation itself: on a CONFIG_PM_SLEEP=n build the old #ifdef-based SET_SYSTEM_SLEEP_PM_OPS leaves the struct’s fields empty, so suspend_fn/resume_fn are never referenced anywhere, and the compiler would warn about unused static functions without that annotation, hence the workaround. The current replacement removes the need for the workaround entirely by changing how the “unused-ness” is expressed:
/*
* Use this if you want to use the same suspend and resume callbacks for
* suspend to RAM and hibernation.
*/
#define DEFINE_SIMPLE_DEV_PM_OPS(name, suspend_fn, resume_fn) \
_DEFINE_DEV_PM_OPS(name, suspend_fn, resume_fn, NULL, NULL, NULL)
#define pm_sleep_ptr(_ptr) PTR_IF(IS_ENABLED(CONFIG_PM_SLEEP), (_ptr))
Because pm_sleep_ptr() is a compile-time conditional expression rather than a preprocessor #ifdef around the whole struct field, the compiler always sees suspend_fn and resume_fn referenced somewhere in the translation unit, whether or not CONFIG_PM_SLEEP is set, so no unused-function warning is possible and __maybe_unused becomes unnecessary. When CONFIG_PM_SLEEP is disabled, PTR_IF resolves the reference to NULL and the now-genuinely-dead function bodies are eliminated by the compiler’s own dead-code elimination rather than by textual preprocessing. Same runtime effect, cleaner mechanism. For drivers whose dev_pm_ops struct must be exported as a symbol for another file to reference, the kernel additionally provides EXPORT_SIMPLE_DEV_PM_OPS() and EXPORT_GPL_SIMPLE_DEV_PM_OPS(), which apply the same SYSTEM_SLEEP_PM_OPS() fields under an EXPORT_SYMBOL wrapper.
Why All These Macros Are Conditioned On CONFIG_PM_SLEEP
The gating is not an implementation accident, it reflects the genuine Kconfig split described earlier. CONFIG_PM_SLEEP is a real, independent option a board can build without: an always-on gateway device, a headless server-class embedded board, or any system whose product requirements simply never include suspending or hibernating the whole machine has no reason to carry suspend/resume/freeze/thaw/poweroff/restore code in its kernel image at all, even though it may still make heavy use of runtime PM for individual peripherals. Every macro in this lecture, old #ifdef-style or new pm_sleep_ptr()-style, exists specifically to let a single driver source file serve both kinds of builds without the driver author writing two versions: on a CONFIG_PM_SLEEP=y kernel the callbacks are wired in and reachable from the suspend/resume/hibernation code paths; on a CONFIG_PM_SLEEP=n kernel the exact same source produces a struct with those fields NULL and the corresponding function bodies gone from the final image, with zero manual #ifdef in the driver itself.
Decision Table: Which Macro For Which Requirement
| If your driver needs… | Use this macro |
|---|---|
| One suspend/resume pair, reused identically for RAM sleep, suspend-to-idle, and hibernation | DEFINE_SIMPLE_DEV_PM_OPS() (current); SIMPLE_DEV_PM_OPS() still works but is marked deprecated in mainline |
| Logic that must run with interrupts masked (shared IRQ lines, bus controllers racing their own handler) | SET_NOIRQ_SYSTEM_SLEEP_PM_OPS(), or DEFINE_NOIRQ_DEV_PM_OPS() when building a fresh struct |
| Logic that should run after runtime PM is already disabled for the device, often reusing runtime_suspend/runtime_resume directly | SET_LATE_SYSTEM_SLEEP_PM_OPS() |
| Fully independent behavior for suspend vs. freeze vs. poweroff (rare, e.g. a device that genuinely must not touch hardware during freeze but does during suspend) | Fill struct dev_pm_ops fields explicitly by hand rather than a single convenience macro |
| Both system sleep and runtime PM callbacks in one struct dev_pm_ops | Current kernel source points you to DEFINE_RUNTIME_DEV_PM_OPS() rather than the deprecated UNIVERSAL_DEV_PM_OPS() |
Common Mistakes And Troubleshooting
- Writing persistent-storage save logic inside .suspend(). RAM-preserving sleep never destroys main memory, so there is nothing to snapshot to disk there; that job belongs to .freeze()/.poweroff() for hibernation specifically.
- Powering down hardware or arming wakeup signaling inside .freeze(). The machine has not lost power yet during image creation; changing power state here breaks the assumption that .thaw() can cleanly roll the device back if image creation aborts.
- Re-saving register state inside .poweroff(). That state was already captured during .freeze(); .poweroff() only needs to quiesce and let power actually be cut.
- Treating .resume() and .restore() as interchangeable. .resume() may run on hardware that quietly retained some state through a shallow sleep; .restore() must assume nothing survived and reprogram everything from the values .freeze() saved.
- Using SIMPLE_DEV_PM_OPS() in new code without checking your target kernel’s pm.h. It still compiles, so nothing forces a rewrite; verify the macro your kernel version actually recommends before treating an older tutorial’s example as current.
Best Practices
- Prefer DEFINE_SIMPLE_DEV_PM_OPS() over SIMPLE_DEV_PM_OPS() for new drivers; it matches what the kernel’s own header now documents as current and avoids __maybe_unused entirely.
- Reuse runtime_suspend()/runtime_resume() for suspend_late()/resume_early() wherever the low-power register state genuinely matches, wiring it with SET_LATE_SYSTEM_SLEEP_PM_OPS() instead of duplicating logic.
- Keep noirq-phase callbacks short and non-blocking; interrupt handling for the entire system is masked while they run.
- Never assume .freeze() and .poweroff() can share power-state-changing code; only the “capture register values” portion is typically shareable.
- When in doubt about a PM macro’s current behavior, check include/linux/pm.h directly for your target kernel version rather than relying on older documentation, exactly as this lecture did.
Summary And Key Takeaways
The six Linux system sleep callbacks, suspend, resume, freeze, thaw, poweroff, and restore, exist because “the system is going to sleep” is not one event but three structurally different transitions: RAM-preserving sleep where main memory survives untouched, hibernation image creation where the device must not lose power yet, and the poweroff/restore pair that brackets an actual power cycle around that saved image. Runtime PM and system sleep share the same struct dev_pm_ops shape but are triggered by entirely different mechanisms and gated by different Kconfig options, CONFIG_PM versus CONFIG_PM_SLEEP, which is exactly why every convenience macro in this lecture, SET_SYSTEM_SLEEP_PM_OPS, SET_NOIRQ_SYSTEM_SLEEP_PM_OPS, SET_LATE_SYSTEM_SLEEP_PM_OPS, and the SIMPLE_DEV_PM_OPS family, is conditioned on CONFIG_PM_SLEEP specifically. Verified directly against current mainline include/linux/pm.h, SIMPLE_DEV_PM_OPS is now flagged deprecated in favor of DEFINE_SIMPLE_DEV_PM_OPS combined with pm_sleep_ptr(), a change worth knowing before copying an older example verbatim. The companion hands-on lecture puts every one of these macros into practice with an original driver you can build and load yourself, watching a full suspend and resume cycle in real dmesg output.
Frequently Asked Questions
Can a driver just point .suspend at its runtime_suspend function and skip writing a separate system sleep callback?
Sometimes, but only when the resulting register state is genuinely identical and you understand that system sleep will invoke it unconditionally regardless of device idle state, unlike runtime PM. It is a valid optimization at the suspend_late/resume_early phase specifically, not a blanket substitution for .suspend/.resume everywhere.
Why does .freeze() exist separately from .poweroff() if both eventually lead to the device being powered off?
Because power is not actually cut at the .freeze() stage. The kernel is still running the pre-hibernation kernel and needs the device usable long enough to write the captured image to disk; only after that succeeds does .poweroff() run and actually remove power.
Is SIMPLE_DEV_PM_OPS actually removed from the kernel?
No, it still compiles as of current mainline, verified directly in include/linux/pm.h. It is marked deprecated in favor of DEFINE_SIMPLE_DEV_PM_OPS, meaning new code should not use it, but existing drivers using it are not broken.
What is the practical difference between the old SET_SYSTEM_SLEEP_PM_OPS and the newer DEFINE_SIMPLE_DEV_PM_OPS?
SET_SYSTEM_SLEEP_PM_OPS only fills fields inside a struct you declare yourself and still relies on an #ifdef CONFIG_PM_SLEEP to decide whether those fields exist at all, which is why callers historically needed __maybe_unused. DEFINE_SIMPLE_DEV_PM_OPS declares the whole struct for you and uses pm_sleep_ptr() so the compiler always sees the functions referenced, letting it drop them via dead-code elimination instead.
Does CONFIG_PM_SLEEP being disabled mean CONFIG_PM is also disabled?
No, the relationship runs the other direction. CONFIG_PM_SLEEP depends on CONFIG_PM, but a kernel can have CONFIG_PM=y with CONFIG_PM_SLEEP=n, giving it full runtime PM support with zero system-wide sleep transition code.
Which macro should a brand-new, simple platform driver reach for first?
DEFINE_SIMPLE_DEV_PM_OPS(), unless the driver has a concrete reason to need noirq-phase or late-phase behavior. It covers the overwhelming majority of drivers where suspend/resume logic is genuinely identical across RAM sleep and hibernation.
Do .thaw() and .resume() ever run for the same transition?
No. .thaw() only runs as part of the hibernation image-creation sub-transition (undoing .freeze()); .resume() only runs after a RAM-preserving sleep. They are mutually exclusive per transition, even though a driver commonly points both at the same underlying function via SET_SYSTEM_SLEEP_PM_OPS() or DEFINE_SIMPLE_DEV_PM_OPS().
Ready To Build A Real Sleep-Aware Driver?
You now know exactly which of the six Linux system sleep callbacks your driver needs and which macro wires them in. The next lecture in this free Linux kernel development course puts it into practice with an original platform driver and real dmesg output across a full suspend/resume cycle.
Try The Hands-On Examples Back To Course Index