Implementing Linux Runtime PM-Free Linux Device Drivers Training Online

Implementing Linux Runtime PM

Implementing Linux Runtime PM

Part 5 of an 11-lecture series on Linux kernel power management — a free Linux kernel development course from EmbeddedPathashala covering the .runtime_suspend, .runtime_resume, and .runtime_idle callback trio that every well-behaved device driver implements.

3 Callbacks
runtime_suspend / runtime_resume / runtime_idle
Lecture 5 of 11
Kernel Power Management Series
Verified Against
Mainline 6.x Kernel Sources

Linux runtime power management is the mechanism that lets a single device — a sensor, a codec, a storage controller — drop into a low-power state the instant nobody is using it, entirely independent of whether the system as a whole is awake, asleep, or hibernating. In the previous lecture of this series we covered hibernation and the broader system sleep framework; this lecture narrows the focus down to the driver-level contract that makes per-device runtime suspend and resume possible: the .runtime_suspend(), .runtime_resume(), and .runtime_idle() callback trio defined inside struct dev_pm_ops, plus the handful of helper calls a driver’s probe() function uses to switch runtime PM on for a device in the first place.

Everything here is checked against the current mainline kernel source — include/linux/pm.h, include/linux/pm_runtime.h, and Documentation/power/runtime_pm.rst — rather than copied from any older textbook, because the macro the kernel expects you to use for wiring up these callbacks has genuinely changed since the 4.19-era material this series is modernizing.

Key Terms In This Lecture

runtime_suspend runtime_resume runtime_idle dev_pm_ops DEFINE_RUNTIME_DEV_PM_OPS pm_runtime_enable pm_runtime_set_active pm_runtime_get_noresume pm_runtime_put usage_count RPM_ACTIVE / RPM_SUSPENDED

What You Will Learn

Why runtime PM is per-device, not system-wide The exact trigger condition for each of the three callbacks What return values mean to the PM core SET_RUNTIME_PM_OPS vs DEFINE_RUNTIME_DEV_PM_OPS Why UNIVERSAL_DEV_PM_OPS is now deprecated The enable/disable lifecycle at probe() time pm_runtime_get_noresume() vs pm_runtime_get_sync() Common bring-up mistakes and how to spot them

Prerequisites

  • Comfortable reading and writing basic C, including function pointers and structs.
  • Basic familiarity with either I2C client drivers or platform drivers (probe/remove, device tree matching).
  • Having read the previous lecture in this series on Hibernation and Device Power Management — this lecture assumes you already know the difference between suspend-to-RAM, suspend-to-disk, and the system-wide dev_pm_ops suspend/resume callbacks.
  • A Linux build environment capable of compiling an out-of-tree kernel module against your running kernel’s headers.

What Is Linux Runtime Power Management?

Every struct device the kernel knows about carries an embedded struct dev_pm_info named power, and inside that structure sits the entire runtime PM bookkeeping: a status field (enum rpm_status, one of RPM_ACTIVE, RPM_SUSPENDED, RPM_RESUMING, or RPM_SUSPENDING), a usage_count atomic, a child_count atomic, and a disable_depth counter. None of this has anything to do with the system-wide suspend/resume/freeze/restore callbacks you studied in the hibernation lecture. A device can be runtime-suspended while the system is fully awake and busy doing other work, and a device can (and usually should) stay runtime-suspended straight through a system suspend/resume cycle if nothing forces it to wake — that interaction between runtime PM and system sleep is exactly why the .suspend_late()/.resume_early() callback pair often point at the very same functions as .runtime_suspend()/.runtime_resume().

The PM core decides whether a device is eligible for a runtime suspend using two independent counters. The usage_count atomic tracks how many active users currently hold a reference on the device — every pm_runtime_get*() call increments it, every pm_runtime_put*() call decrements it. The child_count atomic tracks how many of the device’s children are themselves currently active, unless the parent has set power.ignore_children. Only when both counters read zero, and the device’s runtime_status is RPM_ACTIVE, will the PM core consider running .runtime_idle() and, following that, .runtime_suspend().

One detail that surprises a lot of driver authors the first time: the runtime PM callback the core actually invokes is looked up through a priority chain — PM domain, then device type, then device class, then bus type — before it ever falls back to the driver’s own dev_pm_ops via dev->driver->pm. For a typical I2C or platform driver you write yourself, none of the higher-priority levels usually apply, so the callbacks you register on your own driver’s struct device_driver.pm field are what actually runs. But it is worth knowing that a bus type (I2C, SPI, platform) can, in principle, intercept and even completely replace your driver’s runtime PM behavior.

Runtime PM Status State Machine

RPM_ACTIVE, usage_count and child_count both 0→runtime_idle() runs
runtime_idle() returns 0 (or is NULL)→suspend request queued
RPM_ACTIVE→runtime_suspend() runs→RPM_SUSPENDING
runtime_suspend() returns 0→RPM_SUSPENDED
runtime_suspend() returns -EBUSY / -EAGAIN→stays RPM_ACTIVE, safe to retry later
device access requested (get/resume)→RPM_SUSPENDED→runtime_resume() runs→RPM_RESUMING
runtime_resume() returns 0→RPM_ACTIVE again

The Linux Runtime Power Management Callback Contract

struct dev_pm_ops carries exactly three members dedicated to runtime PM: runtime_suspend, runtime_resume, and runtime_idle, each a plain int (*)(struct device *dev) function pointer. The PM core’s helper functions in drivers/base/power/runtime.c are the only code that is supposed to call these directly — a driver never invokes its own .runtime_suspend() by hand, it goes through pm_runtime_suspend(), pm_runtime_idle(), or one of their many variants so that the core’s internal locking and mutual-exclusion guarantees apply.

.runtime_suspend()

This callback only ever runs while the device’s runtime_status is RPM_ACTIVE and both counters are zero. Its job is to make the device stop talking to the CPU and RAM entirely — mask interrupts that aren’t needed for wakeup, quiesce any DMA in flight, and, if it makes sense for the hardware, cut a clock or a regulator. Returning 0 tells the core the device is now genuinely suspended and moves it to RPM_SUSPENDED. Returning -EBUSY or -EAGAIN means “not right now” — the device stays RPM_ACTIVE and the core is free to try again later; this is the correct return value when, for example, an I2C transaction is mid-flight and cannot be safely interrupted. Returning any other negative error code is treated as fatal: the PM core will refuse to run any further runtime PM helpers for that device until something explicitly forces its status back with pm_runtime_set_active() or pm_runtime_set_suspended().

.runtime_resume()

The mirror image — only ever runs while runtime_status is RPM_SUSPENDED. It must put the device back into full working order: restore power, wait out any settling time the hardware datasheet specifies, and reprogram any registers that were lost when power was cut. Once it returns 0 the core marks the device RPM_ACTIVE and considers it fully operational — callers are entitled to assume I/O will work immediately after a successful pm_runtime_resume() call returns. A nonzero return here is fatal in the same sense as with runtime_suspend(): the PM core stops trusting the device’s runtime PM state until it is corrected explicitly.

.runtime_idle()

This one is not a suspend at all — think of it as a question the core is asking your driver: “usage_count and child_count both just hit zero, are you fine with me suspending this device now?” It fires only when the device is RPM_ACTIVE and both counters read zero. If .runtime_idle() is left unimplemented (NULL) or returns 0, the PM core proceeds to attempt a runtime suspend on your behalf (respecting an autosuspend delay if one has been configured — that mechanism is the subject of the next lecture in this series). Returning any nonzero value tells the core to leave the device alone for now; a driver might do this if it has just queued its own delayed work to suspend later on a custom schedule. Negative error codes returned from .runtime_idle() are simply ignored by the core.

CallbackTrigger ConditionRequired BehaviorReturn Value Meaning
.runtime_suspend()Status is RPM_ACTIVE; usage_count == 0 and child_count == 0 (or ignore_children set)Fully quiesce the device: stop I/O and DMA, optionally cut power/clock, enable remote wakeup if applicable0 = suspended; -EBUSY/-EAGAIN = stay active, retry later; other negative = fatal PM error
.runtime_resume()Status is RPM_SUSPENDEDRestore full power and register state; device must be usable the instant this returns0 = active and operational; negative = fatal PM error
.runtime_idle()Status is RPM_ACTIVE; usage_count and child_count both just reached 0Decide whether a suspend attempt should follow; typically just a policy check0 or NULL = core proceeds to suspend; nonzero = core leaves device alone; negative values ignored

Registering Runtime Power Management Callbacks: Which Macro Should You Use?

The old-style way of populating the three runtime members of struct dev_pm_ops is still present in include/linux/pm.h and still compiles: the SET_RUNTIME_PM_OPS() macro (built on top of the lower-level RUNTIME_PM_OPS() macro) simply assigns .runtime_suspend, .runtime_resume, and .runtime_idle inside a manually written dev_pm_ops initializer, guarded by CONFIG_PM.

static const struct dev_pm_ops ep_prox01_pm_ops = {
	SET_RUNTIME_PM_OPS(ep_prox01_runtime_suspend,
			    ep_prox01_runtime_resume,
			    ep_prox01_runtime_idle)
};

That is still valid C and still exists in the header verbatim as of the 6.x series checked for this lecture. But for the common case — a driver that wants its system-sleep suspend/resume to simply reuse the runtime suspend/resume logic — the kernel now provides a single, more complete macro in include/linux/pm_runtime.h: DEFINE_RUNTIME_DEV_PM_OPS(). It builds the entire const struct dev_pm_ops for you in one line, and it wires the system-sleep .suspend/.resume pointers to pm_runtime_force_suspend() and pm_runtime_force_resume() automatically, so a device that is already runtime-suspended when the system goes to sleep does not get needlessly woken up and resuspended.

static DEFINE_RUNTIME_DEV_PM_OPS(ep_prox01_pm_ops,
				  ep_prox01_runtime_suspend,
				  ep_prox01_runtime_resume,
				  ep_prox01_runtime_idle);

Verified directly against the mainline header comments while researching this lecture: the older UNIVERSAL_DEV_PM_OPS() macro — the one many 4.19-era drivers, and the source material this series is modernizing, used to share a single suspend/resume pair across system sleep and runtime PM — now carries an explicit comment in include/linux/pm.h reading “Deprecated. You most likely don’t want this macro. Use DEFINE_RUNTIME_DEV_PM_OPS() instead.” SET_RUNTIME_PM_OPS() itself has not been deprecated and remains useful as a building block when you need a hand-rolled dev_pm_ops with different system-sleep and runtime callbacks, but for a brand-new driver in 2026, DEFINE_RUNTIME_DEV_PM_OPS() is the macro to reach for by default. Once the struct exists, attach it to your bus driver with the pm_ptr() helper, which quietly compiles to NULL on kernels built without CONFIG_PM instead of requiring you to hand-write #ifdef guards:

static struct i2c_driver ep_prox01_driver = {
	.driver = {
		.name = "ep_prox01",
		.pm   = pm_ptr(&ep_prox01_pm_ops),
	},
	.probe  = ep_prox01_probe,
	.remove = ep_prox01_remove,
};

The Enable/Disable Lifecycle: Wiring Linux Runtime Power Management at probe()

Runtime PM is disabled for every device by default. The disable_depth field inside dev_pm_info starts at 1, and every runtime PM helper function checks it before doing anything; while it is nonzero, calls like pm_runtime_get_sync() simply fail with -EACCES rather than invoking your callbacks. pm_runtime_enable() decrements disable_depth by one, and only once it reaches zero does the PM core start honoring runtime PM requests for that device. pm_runtime_disable() is the mirror call, and it is nestable — every pm_runtime_disable() must be balanced by exactly one later pm_runtime_enable(), which is why the counts must not drift across error paths in probe().

The second piece of default state that trips people up: every device’s initial runtime_status is RPM_SUSPENDED, regardless of what the hardware is actually doing. If your device is genuinely powered up and functional the moment probe() runs — which is common, since a bootloader or a regulator framework may have already turned it on — you must tell the PM core the truth with pm_runtime_set_active() before calling pm_runtime_enable(). If you skip this step, the core’s internal bookkeeping believes the device is already suspended, so the very first .runtime_idle() check will conclude there is nothing to do, and .runtime_suspend() may never run even once — silently leaking power for the entire time the driver is loaded.

A third helper, pm_runtime_get_noresume(), rounds out the “first look” at this lifecycle. It bumps usage_count by exactly one without going through the resume path at all — no callback runs, it is a pure counter increment. Drivers typically call it right after pm_runtime_set_active()/pm_runtime_enable(), while the device is known to already be active, purely to hold a reference open during the rest of probe(). That prevents an unlucky asynchronously-queued idle check from suspending the device out from under you while you are still writing configuration registers. Once probe() finishes its hardware setup, a matching pm_runtime_put() drops that reference; if it brings usage_count back to zero, the core queues an asynchronous .runtime_idle() check via pm_request_idle(), which is what ultimately allows the device to drop into low power once the driver is done needing it awake.

Deeper get/put discipline around actual I/O paths — and the autosuspend delay mechanism that avoids needlessly bouncing a device between active and suspended on every access — is deliberately out of scope here; that is the entire subject of the next lecture in this series. This lecture only covers wiring the lifecycle up once, at probe() and remove() time.

probe() Runtime PM Lifecycle Order

1. pm_runtime_set_active(dev)→tell the core the hardware is really on right now
2. pm_runtime_enable(dev)→disable_depth reaches 0, callbacks now live
3. pm_runtime_get_noresume(dev)→usage_count = 1, no callback invoked
4. finish register / hardware init→safe: device cannot suspend while held
5. pm_runtime_put(dev)→usage_count = 0, idle check queued
6. runtime_idle() → 0 → runtime_suspend() runs→device now RPM_SUSPENDED

Common Mistakes and Troubleshooting

  • Forgetting pm_runtime_enable(). Every helper function keeps returning -EAGAIN/-EACCES forever. Nothing crashes, the device simply never participates in runtime PM, and this is easy to miss unless you check /sys/devices/…/power/runtime_status.
  • Leaving the initial status as suspended when hardware is actually powered on at boot. The first idle check believes the device is already suspended and skips the real power-down path, so .runtime_suspend() may never fire even once.
  • Returning 0 from .runtime_suspend() when the device genuinely could not quiesce — for example, an outstanding I2C transaction. The core marks it suspended anyway, and the next access will fail because the driver’s own internal state assumed hardware that is no longer reachable.
  • Doing slow, blocking work inside .runtime_idle() itself instead of treating it as a cheap policy check and letting .runtime_suspend() do the real work.
  • Copying UNIVERSAL_DEV_PM_OPS() from an old book or blog post. It still compiles, but it is explicitly marked deprecated in the current kernel headers in favor of DEFINE_RUNTIME_DEV_PM_OPS().
  • Mismatched pm_runtime_enable()/pm_runtime_disable() counts across probe() error paths and remove(), leaving disable_depth permanently nonzero after an unload/reload cycle.
  • Forgetting pm_ptr() (or an equivalent #ifdef) around the .pm field, which produces unused-function warnings on kernels built without CONFIG_PM.

Best Practices

  • Report the true initial hardware state with pm_runtime_set_active() (or leave it suspended, if that is truthful) before calling pm_runtime_enable() — never guess.
  • Keep .runtime_suspend() and .runtime_resume() symmetric: resume should undo exactly what suspend did, nothing more, nothing less.
  • Return -EBUSY or -EAGAIN honestly when a device cannot suspend right now instead of silently pretending success.
  • Default to DEFINE_RUNTIME_DEV_PM_OPS() so your system-sleep behavior stays consistent with your runtime PM behavior through pm_runtime_force_suspend()/pm_runtime_force_resume().
  • Treat every pm_runtime_get*() like a lock acquisition — pair it with exactly one pm_runtime_put*() on every code path, including error paths.
  • Watch /sys/devices/…/power/runtime_status, runtime_active_time, and runtime_suspended_time during bring-up rather than relying on dmesg alone.
  • Keep the work done inside .runtime_suspend()/.runtime_resume() minimal — they run from the pm_wq workqueue and block other pending PM operations on the same device while executing.

Summary and Key Takeaways

Linux runtime power management gives every device driver a self-contained, per-device suspend/resume contract that operates completely independently of system-wide sleep states. struct dev_pm_ops exposes exactly three hooks for it — .runtime_suspend(), .runtime_resume(), and .runtime_idle() — each with a precise trigger condition and a return-value contract the PM core relies on to keep its internal RPM_ACTIVE/RPM_SUSPENDED bookkeeping correct. The macro used to register these callbacks has moved on since the older textbooks: DEFINE_RUNTIME_DEV_PM_OPS() is now the default choice, SET_RUNTIME_PM_OPS() remains available as a lower-level building block, and UNIVERSAL_DEV_PM_OPS() is explicitly deprecated in current kernel headers. At the driver level, the lifecycle begins in probe() with pm_runtime_set_active(), pm_runtime_enable(), and a pm_runtime_get_noresume()/pm_runtime_put() bracket around any remaining hardware setup, and ends in remove() with pm_runtime_disable(). The deeper discipline of pairing get/put calls around real I/O paths, and the autosuspend delay mechanism that keeps a device from thrashing between power states, is the subject of the next lecture. Understanding this callback contract solidly is the foundation the rest of this free Linux kernel development course builds on for everything else in the power management series.

Frequently Asked Questions

What is the difference between runtime PM and system suspend?

System suspend (suspend-to-RAM) and hibernation (suspend-to-disk) act on the whole machine at once, driven by struct dev_pm_ops members like .suspend()/.resume() and .freeze()/.restore(), covered in the previous lecture. Linux runtime power management instead acts on one device at a time, triggered purely by that device’s own usage_count and child_count reaching zero, and it can happen many times per second while the system stays fully awake.

What should .runtime_suspend() return if the device is busy?

Return -EBUSY or -EAGAIN. Both leave the device’s runtime_status at RPM_ACTIVE and tell the PM core it is safe to attempt the suspend again later. Any other negative error code is treated as a fatal PM error and stops further runtime PM helpers from running for that device until the status is corrected explicitly.

Why does .runtime_idle() usually just return 0?

Because in most drivers there is no extra policy decision to make — once usage_count and child_count both hit zero, the device really is idle and should suspend. Returning 0 (or leaving the callback NULL entirely) simply lets the PM core proceed straight to .runtime_suspend().

Is SET_RUNTIME_PM_OPS still valid to use in current kernels?

Yes. It is checked directly against mainline include/linux/pm.h and it still exists and compiles. It has not been marked deprecated. However, for the common case of sharing suspend/resume behavior between runtime PM and system sleep, DEFINE_RUNTIME_DEV_PM_OPS() from include/linux/pm_runtime.h is now the recommended, more complete macro.

What happens if I forget to call pm_runtime_enable() in probe()?

Nothing crashes. Every runtime PM helper function checks power.disable_depth first, and since it starts at 1, the helpers simply return -EAGAIN or -EACCES forever. Your .runtime_suspend()/.runtime_resume()/.runtime_idle() callbacks will never be invoked, and the device silently never participates in runtime PM.

Why call pm_runtime_set_active() before pm_runtime_enable() in probe()?

Because every device’s initial runtime_status defaults to RPM_SUSPENDED regardless of the device’s real hardware state. If the hardware is actually already powered on when probe() runs, pm_runtime_set_active() corrects the core’s bookkeeping to match reality before runtime PM helpers start being honored.

Can .runtime_suspend() and .runtime_resume() run at the same time?

No. The PM core guarantees these callbacks are mutually exclusive for a given device — it will never run .runtime_suspend() and .runtime_resume() concurrently, nor two instances of .runtime_suspend() concurrently. .runtime_idle() is the one exception: it can run alongside .runtime_suspend() or .runtime_resume(), although it will not start while either of the other two is already in progress.

Do I need pm_runtime_get_noresume() in every driver’s probe()?

Only when the device is already active at probe() time and you want to hold a reference open while you finish hardware initialization, so an asynchronously-queued idle check cannot suspend the device mid-setup. If a driver’s hardware genuinely starts in a suspended, powered-off state, this call is unnecessary.

See This Callback Trio In A Real Driver

The next page in this pair implements every callback and lifecycle call discussed here inside a complete, original ep_prox01 I2C driver — with a build walkthrough and real dmesg output showing suspend, resume, and idle transitions. It’s part of our free Linux device drivers course, built entirely from current mainline kernel APIs.

See The ep_prox01 Driver Example Back To Course Index

Leave a Reply

Your email address will not be published. Required fields are marked *