Linux Hibernation And Device PM
Lecture 4 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course: suspend-to-disk mechanics, RTC wakeup testing, and the struct dev_pm_ops contract.
Linux hibernation is the one system sleep state that survives a total loss of power. Where suspend-to-RAM keeps your data alive by spending a trickle of current on memory self-refresh, Linux hibernation writes a complete snapshot of memory to persistent storage and then cuts power to almost everything, including RAM itself. This lecture, part of Ravi’s free Linux kernel development course, goes deep on how that snapshot is created and restored, how to drive the process through /sys/power/disk, and how to safely validate suspend-to-RAM wakeup using nothing more exotic than an RTC alarm. We then use hibernation as the doorway into device-driver power management proper: struct device, the PM-relevant fields it carries, and struct dev_pm_ops, the callback contract every power-aware driver has to fill in.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Basic C, including structures, function pointers, and pointer arithmetic.
- Basic Linux kernel build knowledge: menuconfig, building and loading out-of-tree modules.
- Comfort reading and writing sysfs attributes as root.
- Having read the previous lecture, Linux Thermal And Sleep States, is strongly recommended; this lecture assumes you already know the four common sleep states and the basic
/sys/power/stateand/sys/power/mem_sleepinterface from that lecture.
Why Linux Hibernation Still Matters
It is tempting to think of Linux hibernation as a legacy laptop feature from an era before fast SSDs and always-on batteries, but that undersells what it actually offers. Suspend-to-RAM only ever saves you from idle-time power draw; the instant the battery is fully depleted, or someone pulls the plug on a device with no battery at all, everything held in that self-refreshing RAM is gone. Hibernation is the only system sleep state the kernel supports that survives a genuine, total loss of power, because by the time power is actually removed, the entire state of the machine is already sitting safely on disk. That property still matters for long-idle embedded appliances that cannot justify a battery just to keep RAM alive, for systems where a scheduled maintenance window needs to guarantee zero power draw, and for kernel developers who use hibernation’s test_resume mode purely as a stress test for device driver power management code, which is exactly the code this lecture is building toward.
Testing Suspend-to-RAM Wakeup With An RTC Alarm
Before diving into hibernation, it is worth locking down a practical habit that makes every sleep-state experiment in this series safer: never trigger a system sleep transition without first arming a wakeup source you fully control. The Real Time Clock is the most predictable one available on almost any platform, because it needs no external hardware, no network, and no human intervention. Every RTC exposed through the RTC class framework publishes a wakealarm attribute under /sys/class/rtc/rtcN/:
$ cat /sys/class/rtc/rtc0/wakealarm
# echo +30 > /sys/class/rtc/rtc0/wakealarm
An empty read means no alarm is currently armed. Writing a string prefixed with + arms the alarm that many seconds from now, relative to the RTC’s own clock, which is exactly what you want when scripting an unattended suspend-to-RAM test: arm the alarm for, say, thirty seconds out, then write mem to /sys/power/state, and the machine should come back on its own without you touching a keyboard, a network link, or a power button. This is the same mechanism, at a lower level, that rtcwake wraps for you; understanding the raw sysfs attribute matters because not every embedded target ships that userspace tool, and because the identical alarm mechanism is what wakes a machine back up from hibernation when the platform’s low-power state (rather than a plain power-off) is selected. The full hands-on walkthrough, including a script that reads the current RTC time before arming the alarm, is in the companion examples lecture.
Linux Hibernation Mechanics: From Snapshot To Swap
Hibernation is structurally different from every other sleep state, because it is not really one transition at all. Entering it takes three distinct system state changes, and leaving it takes two, and a second, independent kernel instance is involved along the way.
The Three-Phase Path Into Hibernation
When hibernation is triggered, the kernel first stops ordinary system activity (freezing user space tasks, much like any other sleep transition) and builds a snapshot image of memory. Building that snapshot cannot simply be a raw memcpy, because the memory being copied is also the memory the copying code itself needs to run in; the hibernation core solves this with pre-allocated memory bitmaps that track which pages belong to the image and works around the chicken-and-egg problem by keeping enough free pages set aside up front. Once the image exists as a self-consistent snapshot, the system moves into a state where devices are reactivated just long enough to write that image out to persistent storage, because writing to a disk or an eMMC part requires a working block driver, DMA, and interrupts, none of which were available while the snapshot was being taken with everything quiesced. Only after the image has actually landed in storage does the kernel proceed to the last phase: quiesce every device one final time and cut power, either through firmware-assisted low-power entry or a hard shutdown.
Where The Image Actually Lives: Swap And image_size
A Linux hibernation image is written to swap space, not to an arbitrary file on an ordinary filesystem, because swap is the one storage area the kernel can address directly at a low level without depending on a mounted, journaled filesystem being available during the fragile write-out phase. In practice this means a machine intended to hibernate needs a swap partition (or, on more recent setups, a properly configured swap file with a resolvable physical offset) at least as large as the memory actually in use at hibernation time, and the kernel’s boot parameters or resume tooling need to know which swap device and offset to look at when searching for a saved image after a reboot. The size of the image itself is tunable: /sys/power/image_size accepts a best-effort byte limit that the hibernation core tries to respect, defaulting to roughly two-fifths of available RAM, and writing 0 forces the smallest image the kernel can manage to produce, which trades a longer creation time (more compression and page reclaim effort) for less required swap space.
The Two-Kernel Resume: Restore Kernel Meets Image Kernel
Resuming from Linux hibernation is where the mechanism earns its reputation for complexity. Control cannot simply jump back into the old, hibernated kernel instance, because that instance no longer exists in memory; power was removed. Instead, the platform firmware runs the boot loader exactly as it would for any cold boot, and that boot loader starts a fresh kernel instance, referred to as the restore kernel. The restore kernel locates the saved image in swap, loads it into memory, and then does something unusual: it stops all activity, overwrites its own running memory with the contents of the saved image, and jumps into a small trampoline area that was itself saved as part of the image, handing control to what is now called the image kernel, the original hibernated kernel brought back to life. From that trampoline onward, the image kernel is responsible for restoring devices to their pre-hibernation state and letting user space run again. The restore kernel does not need to be identical to the image kernel; it only needs enough driver support to reach the storage device holding the image, which is why a minimal initramfs is often sufficient to drive the resume path even on systems with a much larger, fuller kernel actually doing the hibernating.
Linux Hibernation Entry And Exit Flow
Controlling Hibernation Behavior Through /sys/power/disk
The /sys/power/disk file governs what the kernel does immediately after a hibernation image has been created and saved. Reading it lists every option the running kernel supports, with the currently selected one shown in square brackets:
$ cat /sys/power/disk
[platform] shutdown reboot suspend test_resume
Each option changes what happens after the image is safely on disk:
- platform — put the system into a special firmware-assisted low-power state (ACPI S4, for example) instead of a plain power-off, which can unlock extra wakeup sources such as a keyboard key or a laptop lid switch. Only available when the platform provides this mechanism.
- shutdown — power the system off completely, the simplest and most universally supported option.
- reboot — reboot immediately after saving the image, mainly useful for diagnostics since it lets you watch the resume path happen right away without physically power-cycling anything.
- suspend — hybrid sleep: after saving the image, the kernel actually enters the suspend variant selected through
mem_sleeprather than powering off. If the system wakes normally from that suspend state, the saved hibernation image is simply discarded and everything continues; only if the suspend attempt fails to bring the system back does the saved image get used to restore state, giving you suspend-to-RAM’s speed with hibernation’s safety net. - test_resume — a pure diagnostic mode: load the just-created image back in immediately and run the full resume path as though a real restore kernel had booted, without ever actually powering anything off. This is invaluable when debugging a driver’s hibernation callbacks, since it exercises freeze/thaw/restore logic on demand with no reboot required.
Writing one of the listed strings selects it for the next hibernation triggered by writing disk to /sys/power/state. The exact command sequence, including a full reboot-mode hibernation cycle with realistic kernel log output, is demonstrated step by step in the companion examples lecture.
CONFIG_HIBERNATION And ARCH_HIBERNATION_POSSIBLE
Linux hibernation support is gated by the CONFIG_HIBERNATION kernel configuration option, but unlike CONFIG_SUSPEND, this option cannot simply be turned on for every architecture. Resuming from hibernation requires low-level, architecture-specific assembly that can safely overwrite running kernel memory with a saved image and then jump into it without corrupting the very code doing the jumping, and not every CPU architecture port includes that code. The kernel tracks which architectures have implemented it through the ARCH_HIBERNATION_POSSIBLE Kconfig symbol; CONFIG_HIBERNATION can only be selected on architectures where that symbol is set. On architectures where it is not set, /sys/power/disk and the disk string in /sys/power/state simply will not exist, regardless of how the rest of the kernel is configured.
Linux Hibernation Vs Suspend-to-RAM: A Direct Comparison
| Aspect | Suspend-to-RAM (mem / deep) | Linux Hibernation (disk) |
|---|---|---|
| Power draw while asleep | Low, but nonzero (RAM self-refresh) | Effectively zero (or platform S4 minimum) |
| Survives total power loss | No — losing power loses all state | Yes — state is already on persistent storage |
| Storage requirement | None beyond RAM itself | Swap space at least as large as the in-use image |
| Typical resume latency | Roughly one to a few seconds | Several seconds to tens of seconds (I/O bound) |
| Kernel instances involved | One (the same kernel resumes itself) | Two (a restore kernel hands off to the image kernel) |
| Governing config option | CONFIG_SUSPEND + platform mem_sleep support | CONFIG_HIBERNATION + ARCH_HIBERNATION_POSSIBLE |
| Best suited for | Short pauses: laptop lid close, quick idle | Long idle, no-battery systems, driver PM stress testing |
From System Sleep To Device Drivers: struct device And Power Management
Everything covered so far in this lecture has been system-wide: the whole machine moves through freeze, image creation, and power-off together. But none of that would work if individual devices did not cooperate, and that cooperation is coordinated through fields the Linux driver model already attaches to every struct device. Conceptually, a handful of fields matter directly for power management:
struct device {
struct device *parent; /* PM walks the device tree bottom-up
* to suspend, top-down to resume */
struct bus_type *bus; /* may supply bus->pm callbacks */
struct device_driver *driver; /* driver->pm is where dev_pm_ops lives */
struct dev_pm_info power; /* runtime status, wakeup source, flags */
struct dev_pm_domain *pm_domain; /* PM domain ops, if present, win
* over bus/type/class/driver ops */
/* ... many other fields omitted ... */
};
The parent pointer is what lets the PM core enforce a hard ordering rule: a device can never be suspended before all of its children are already suspended, and can never be resumed after any of its children have been resumed, because the driver model’s device hierarchy is defined to mirror real hardware bus topology. The power field, of type struct dev_pm_info, is where the PM core tracks everything about that device’s current power state: whether it is runtime-suspended, whether it can wake the system up, what driver PM flags it was probed with, and more. The pm_domain pointer exists for platforms where several devices share a power rail or clock and must be transitioned together; if it is set, its callbacks take precedence over everything else. Most driver authors, though, will spend the overwhelming majority of their time in exactly one place: the pm member of struct device_driver, which is where struct dev_pm_ops plugs in.
struct dev_pm_ops: The Contract Every PM-Aware Driver Fills In
Device power management operations, whether declared at the bus, class, PM domain, or individual driver level, are all implemented by populating one shared object type: struct dev_pm_ops, defined in include/linux/pm.h. This is the actual current definition from mainline kernel documentation, and every field you see below is a function pointer taking nothing but the affected struct device *:
struct dev_pm_ops {
int (*prepare)(struct device *dev);
void (*complete)(struct device *dev);
int (*suspend)(struct device *dev);
int (*resume)(struct device *dev);
int (*freeze)(struct device *dev);
int (*thaw)(struct device *dev);
int (*poweroff)(struct device *dev);
int (*restore)(struct device *dev);
int (*suspend_late)(struct device *dev);
int (*resume_early)(struct device *dev);
int (*freeze_late)(struct device *dev);
int (*thaw_early)(struct device *dev);
int (*poweroff_late)(struct device *dev);
int (*restore_early)(struct device *dev);
int (*suspend_noirq)(struct device *dev);
int (*resume_noirq)(struct device *dev);
int (*freeze_noirq)(struct device *dev);
int (*thaw_noirq)(struct device *dev);
int (*poweroff_noirq)(struct device *dev);
int (*restore_noirq)(struct device *dev);
int (*runtime_suspend)(struct device *dev);
int (*runtime_resume)(struct device *dev);
int (*runtime_idle)(struct device *dev);
};
Notice what changed from the pre-4.x calling convention some older references still describe: every one of these callbacks takes just struct device *dev, not a legacy pm_message_t state argument. Do not copy suspend/resume signatures that carry a second parameter from outdated material; they belong to a deprecated, non-dev_pm_ops interface the kernel documentation explicitly calls out as legacy.
The System-Sleep Callback Family
prepare() and complete() bookend every kind of system sleep transition; prepare() mainly exists to stop new child devices from being registered mid-transition, and it is the only PM phase where the device hierarchy is walked top-down instead of bottom-up. suspend() and resume() handle suspend-to-idle, standby, and suspend-to-RAM. freeze() and thaw() are hibernation-specific: they run while the memory snapshot is being created and then reactivate the device just long enough for the image to be written out, without ever fully powering the device down. poweroff() and restore() are also hibernation-specific, covering the final quiesce before power is cut and the corresponding bring-up in the image kernel after resume. Because freeze/poweroff and suspend are so often functionally identical for a given driver, the kernel provides the SET_SYSTEM_SLEEP_PM_OPS() convenience macro, which points suspend, freeze, and poweroff at the same function, and resume, thaw, and restore at another single function. That macro, and exactly how it changes what you see in dmesg during a real hibernation cycle, is the centerpiece of this lecture’s hands-on demo driver.
The _noirq, _late, And _early Variants
Each phase above also has _late/_early and _noirq siblings that run at progressively more constrained points during the transition. The plain callback runs first; _late variants (like suspend_late) continue that same work, typically after runtime PM has already been disabled for the device; and _noirq variants run last, after the kernel has disabled IRQ handling entirely (except for handlers explicitly flagged IRQF_NO_SUSPEND). Bus types that share a single interrupt vector across multiple devices, PCI being the textbook example, generally need the _noirq phase so a device is not left able to field a shared interrupt raised by a sibling device that already finished suspending. Most simple drivers implement none of these split-phase variants and rely purely on the plain suspend/resume pair.
The Three Runtime PM Callbacks: A Preview
The last three entries, runtime_suspend(), runtime_resume(), and runtime_idle(), are structurally part of the same struct but conceptually a different subsystem: Runtime Power Management, which lets a single device suspend and resume itself autonomously while the rest of the system, and the rest of user space, keeps running normally. This lecture introduces them only at survey level, as entries in the contract you will eventually fill in, because implementing real runtime-suspend logic, usage counters, and autosuspend delays properly is substantial enough to be the entire subject of the next lecture in this series.
Device Power States: A D0-D3 Mental Model
Device low-power states are not standardized the way system sleep states are; one device might only ever support “on” and “off,” while another exposes a dozen intermediate levels. That said, it helps enormously to keep a simple four-level mental model borrowed from PCI power management terminology, since a lot of subsystem documentation and driver code still speaks in these terms even outside PCI itself. D0 is full power: the device is completely operational. D1 and D2 are intermediate, vendor- and bus-defined low-power states where some functionality is reduced but the device can typically still generate a wakeup signal and return to D0 relatively quickly. D3 is the deepest state; PCI further splits it into D3hot (power is still applied but the device is otherwise fully quiesced, context may be lost) and D3cold (power is physically removed from the device, so absolutely everything about its state must be rebuilt on the way back to D0). When a driver’s suspend() callback runs during a system-wide sleep transition, its job is conceptually to move the device toward D3; when resume(), thaw(), or restore() runs, its job is to bring the device back to D0, and the amount of state that survived the trip in between is exactly what determines how much reinitialization work that callback actually has to do.
Common Mistakes And Troubleshooting
- Assuming CONFIG_HIBERNATION implies hibernation is supported everywhere. The option can only be enabled on architectures where
ARCH_HIBERNATION_POSSIBLEis set; the low-level resume trampoline code is architecture-specific and not automatically available. - No swap, or swap too small, then wondering why hibernation fails. The image is written to swap; without a swap partition or properly configured swap file at least as large as the in-use memory footprint, hibernation cannot complete.
- Testing hibernation over SSH without an RTC alarm or a physical way to power the machine back on. Once power is cut, there is no network link left to reconnect to; always keep local or remote power control available, or lean on
test_resumemode while developing. - Copying an old suspend()/resume() signature that takes a pm_message_t state argument. Current
struct dev_pm_opscallbacks take onlystruct device *dev; that extra parameter belongs to the deprecated legacy PM interface. - Forgetting freeze() and poweroff() exist and only implementing suspend(). If a driver relies on manual struct initialization instead of
SET_SYSTEM_SLEEP_PM_OPS()and only fills insuspend/resume, its device will not be properly quiesced or restored during a hibernation cycle, only during suspend-to-RAM. - Not using pm_trace when a hibernation or suspend transition hangs mysteriously. Writing
1to/sys/power/pm_tracerecords the last suspend/resume event point into RTC memory across a hard reboot, which is often the fastest way to identify which driver’s callback never returned.
Best Practices
- Always arm an RTC wakealarm before testing any deeper sleep state on hardware you cannot physically reach.
- Use hibernation’s
test_resumemode while developing driver PM callbacks, so you can iterate without a real reboot cycle each time. - Size swap generously if hibernation is a real requirement; a too-small swap area is one of the most common causes of hibernation failing partway through image creation.
- Prefer
SET_SYSTEM_SLEEP_PM_OPS()for drivers whose suspend and hibernation handling are identical, rather than hand-populating six nearly duplicate callback pointers. - Treat the runtime_* callbacks as a separate concern from system sleep while you are still learning; understand the contract’s shape first, implement real runtime PM logic once that shape is second nature.
Summary And Key Takeaways
Linux hibernation is fundamentally different from every other system sleep state because it survives total power loss: the kernel builds a memory snapshot, reactivates devices just long enough to write that snapshot to swap, and then powers down, with /sys/power/disk controlling exactly what “powers down” means (platform, shutdown, reboot, hybrid suspend, or a diagnostic test-only resume). Resuming is a two-kernel affair, a restore kernel handing off to the original image kernel through a saved trampoline, gated end to end by CONFIG_HIBERNATION and architecture support tracked through ARCH_HIBERNATION_POSSIBLE. None of this works without individual devices cooperating, which is exactly what struct dev_pm_ops exists to standardize: a single, well-defined contract of prepare/complete, suspend/resume, freeze/thaw, poweroff/restore, their _late and _noirq siblings, and the three runtime PM callbacks that the next lecture in this free Linux kernel development course will implement for real. With hibernation mechanics and the shape of dev_pm_ops now solid, the companion examples lecture puts an RTC-driven suspend test and a full hibernation cycle in reboot mode on your own terminal, alongside an original driver that shows exactly where this contract plugs into a real driver structure.
Frequently Asked Questions
Why does Linux hibernation need swap space specifically, not just any file?
The hibernation image is written by low-level kernel code during a fragile phase of the transition, before a full filesystem stack can safely be relied upon. Swap space (a partition, or a swap file with a resolvable physical offset) is the one storage area the kernel can address directly for this purpose.
What is the difference between freeze() and suspend() in struct dev_pm_ops?
suspend() runs for suspend-to-idle, standby, and suspend-to-RAM transitions and generally puts the device into a genuine low-power state. freeze() is hibernation-specific: it quiesces the device while the memory snapshot is created, but should not fully power the device down or enable wakeup signaling, since the device will be briefly reactivated again for thaw() shortly afterward.
How do I test my driver’s hibernation callbacks without rebooting every time?
Write test_resume to /sys/power/disk before triggering hibernation. The kernel creates the image, then immediately runs the full resume path against it as if a restore kernel had booted, without actually powering anything off, which is much faster to iterate on during development.
Is CONFIG_HIBERNATION available on every CPU architecture?
No. It can only be enabled on architectures where ARCH_HIBERNATION_POSSIBLE is set, because resuming from hibernation requires architecture-specific low-level code capable of safely overwriting running memory with a saved image and jumping into it.
Do suspend() and resume() in current kernels take a pm_message_t state argument?
No. In the current struct dev_pm_ops contract, every callback, including suspend() and resume(), takes only struct device *dev. The extra state parameter belongs to an older, deprecated legacy PM interface that predates dev_pm_ops.
What does SET_SYSTEM_SLEEP_PM_OPS actually do?
It is a convenience macro that populates suspend, freeze, and poweroff with the same suspend function pointer, and resume, thaw, and restore with the same resume function pointer, which is exactly right for drivers whose device handling is identical across suspend-to-RAM and hibernation.
What is the practical difference between D3hot and D3cold?
Both are the deepest low-power state in the D0-D3 model, but D3hot still has power applied to the device even though it is fully quiesced, while D3cold means power has been physically removed. A device coming back from D3cold generally needs full reinitialization, since nothing about its prior state survived.
Ready To Put This Into Your Terminal?
You now understand how Linux hibernation builds and restores a memory snapshot, how to steer it through /sys/power/disk, and the full shape of struct dev_pm_ops. Continue to the hands-on examples lecture to run a real RTC-triggered suspend, a full hibernation cycle in reboot mode, and an original ep_hibernate_demo driver, part of this free Linux kernel development course.
Try The Hands-On Examples Next: Runtime PM Implementation