Linux Runtime PM Autosuspend Guide
Lecture 6 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course, covering synchronous versus asynchronous runtime PM operations and the runtime PM autosuspend mechanism.
Every runtime PM call a driver makes is either synchronous or asynchronous, and knowing which one you are calling, and from where, is the difference between a driver that behaves correctly under an interrupt handler and one that deadlocks or sleeps where it must not. Layered on top of that distinction sits runtime PM autosuspend, the mechanism that keeps a device powered for a short grace period after the last access instead of suspending it the instant it goes idle. This lecture builds a precise, current-kernel mental model of both: how the power core decides whether to run a request inline or hand it to a workqueue, and how runtime PM autosuspend turns a naive suspend-on-idle policy into something that survives real, bursty device traffic without thrashing the hardware.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Basic C, including function pointers and struct-based ops tables.
- I2C and platform driver basics: probe/remove, struct device, and how a driver binds to a struct device.
- Having gone through the previous Runtime PM Implementation lecture in this free Linux kernel development course is strongly recommended, since it introduces pm_runtime_enable(), the basic runtime_suspend/runtime_resume callback pair, and the ep_sensor driver this lecture continues to build on.
- Root access on a Linux system or VM where you can build and load kernel modules.
Why Runtime PM Needs Both Synchronous And Asynchronous Paths
A driver rarely touches its hardware from just one context. A sysfs read, an ioctl, or a probe routine all run in process context, where sleeping is fine and a driver can simply wait for a device to wake up before continuing. An interrupt handler, a tasklet, or any other atomic context is a different story: sleeping there is a bug, full stop, yet the driver may still need to tell the PM core that the device just became busy or just went idle. The runtime PM core resolves this tension by offering every core operation, resume, suspend, and idle notification, in two flavors: a synchronous one that blocks the caller until the underlying transition finishes, and an asynchronous one that only records the request and returns immediately, leaving the actual work to run later on a dedicated workqueue.
Synchronous Runtime PM Operations
Synchronous helpers, most commonly pm_runtime_get_sync() (or its safer modern counterpart pm_runtime_resume_and_get(), which additionally drops the usage count automatically if resume fails) and pm_runtime_put_sync(), invoke the underlying suspend or resume machinery directly, on the calling thread, and do not return until that machinery has actually run. If the device was suspended and you call pm_runtime_get_sync(), your thread blocks inside the PM core while the registered runtime_resume callback executes, and only continues once the device is confirmed active. This is exactly the behavior you want in probe(), in a sysfs show/store handler, or anywhere else your code is about to touch hardware registers and simply cannot proceed until the device is guaranteed to be powered and clocked.
The cost of that guarantee is that synchronous calls must run in a context where blocking is legal. Calling pm_runtime_get_sync() from inside a hardware interrupt handler, a spinlock-protected section, or any other atomic context is a real bug: at best it triggers a “scheduling while atomic” warning, at worst it deadlocks the system.
Asynchronous Runtime PM Operations And The PM Workqueue
Asynchronous helpers such as pm_runtime_get(), pm_runtime_put(), and pm_runtime_idle() never run the resume, suspend, or idle callback on the calling thread at all. Instead, calling one of them causes the PM core to record what needs to happen and defer the actual execution. Concretely, the sequence looks like this:
Asynchronous Runtime PM Request Path
The field that drives this is power.request, a member of the device’s struct dev_pm_info whose type is enum rpm_request: values such as RPM_REQ_IDLE, RPM_REQ_SUSPEND, RPM_REQ_RESUME, and RPM_REQ_AUTOSUSPEND tell the deferred work function which action to carry out once it finally runs. The work itself is power.work, a struct work_struct initialized during pm_runtime_init() with pm_runtime_work() as its function; queuing it schedules that function to run on one of the kernel’s worker threads rather than on your calling thread.
Why Async Helpers Are Safe From Atomic And IRQ Context
Because the asynchronous helpers do nothing more than write a couple of fields and schedule work, they never sleep, never take a mutex that a resume callback might also need, and never block waiting on hardware. Everything that could sleep, the actual suspend or resume callback, runs later, on a normal kernel thread with a normal process context, where sleeping is perfectly legal. This is precisely why pm_runtime_get(), pm_runtime_put(), and friends are documented as safe to call from an interrupt handler: calling one from an IRQ context is equivalent to saying “handle this once you get a chance,” and the workqueue mechanism guarantees that chance will come, just not on your interrupt stack.
The Autosuspend Mechanism
Plain runtime PM, without autosuspend, suspends a device the moment its usage count drops to zero and the idle callback (or the default idle handling) decides to proceed with a suspend. For devices that are cheap to power up and down, that is fine. For devices where powering off and back on again costs real time, a sensor that needs tens of milliseconds to stabilize after power-up, a codec that must re-negotiate a clock, a storage controller that has to re-spin media, suspending immediately after every single access is actively harmful. A driver polled once a second by user space would spend more time resuming and suspending than it ever spends actually transferring data, and the device would wear through more power-state transitions than it would under a policy that just kept it on a little longer.
Runtime PM autosuspend solves this by inserting a deliberate grace period between “usage count reached zero” and “suspend actually happens.” It is important to be precise about what runtime PM autosuspend does not mean: the device does not suspend itself automatically the instant it is idle. Instead, a timer is armed, and only when that timer expires without the device being touched again does a suspend request actually get submitted. If the device is accessed again before the timer fires, the timer is simply pushed back out, and the suspend never happens at all for that idle window.
The Autosuspend Timer And power.last_busy
The timer backing runtime PM autosuspend is the device’s power.suspend_timer field, set up alongside the rest of the PM state in pm_runtime_init(). Its expiration is computed from two other fields: power.last_busy, a timestamp in jiffies recording the last time the device was known to be active, and power.autosuspend_delay, the configured grace period in milliseconds. Conceptually, the next expiration is last_busy + msecs_to_jiffies(autosuspend_delay). Every time a driver calls pm_runtime_mark_last_busy(), that timestamp is refreshed to “now,” which in turn pushes the timer’s effective expiration further into the future.
Autosuspend Timer Lifecycle
Enabling And Configuring Runtime PM Autosuspend
A driver opts into runtime PM autosuspend by calling pm_runtime_use_autosuspend() on its device, preferably before the device is registered, which sets power.use_autosuspend to true. From that point on, the driver should consistently use the _autosuspend() family of helpers rather than the plain ones, so the grace period is actually respected. The length of the grace period itself is set with pm_runtime_set_autosuspend_delay(), in milliseconds, and can also be observed and changed at runtime from user space through the /sys/devices/.../power/autosuspend_delay_ms attribute, without recompiling or reloading the driver. A value of 0 means suspend essentially as soon as the device is idle, while a negative value disables autosuspend entirely for that device, falling back to the immediate-suspend behavior of plain runtime PM.
The companion helper pm_schedule_suspend() can also arm the same timer directly, with an explicit delay in milliseconds passed as an argument. When called this way, the explicit delay takes precedence over whatever is currently stored in power.autosuspend_delay for that one request, which is useful for one-off cases where a driver knows a particular idle period should be treated differently from its usual policy.
Runtime PM Autosuspend Helper Function Reference
The table below summarizes the runtime PM helper functions most relevant to synchronous versus asynchronous behavior and to runtime PM autosuspend specifically, all declared in include/linux/pm_runtime.h and implemented in drivers/base/power/runtime.c.
| Function | Sync/Async | What It Does |
|---|---|---|
| pm_runtime_get_sync() | Synchronous | Increments the usage count and resumes the device inline, blocking until resume completes; does not automatically drop the count on failure. |
| pm_runtime_resume_and_get() | Synchronous | Same as pm_runtime_get_sync(), but automatically drops the usage count if the resume fails, avoiding a common reference-counting bug. |
| pm_runtime_put_sync() | Synchronous | Decrements the usage count and, if it reaches zero, runs the idle/suspend path inline before returning. |
| pm_runtime_get() / pm_runtime_put() | Asynchronous | Adjust the usage count and queue a resume/idle request on the PM workqueue instead of running it inline; safe from atomic context. |
| pm_runtime_put_noidle() | Neither (no callback) | Decrements the usage count only, without triggering any idle notification; typically used in remove() right before pm_runtime_disable(). |
| pm_runtime_use_autosuspend() | Setup call | Enables the autosuspend mechanism for this device by setting power.use_autosuspend, so the *_autosuspend() helpers take effect. |
| pm_runtime_dont_use_autosuspend() | Setup call | Disables autosuspend for this device, reverting to immediate-suspend semantics on the next idle notification. |
| pm_runtime_set_autosuspend_delay() | Setup call | Sets power.autosuspend_delay in milliseconds, the grace period used to compute the suspend timer’s expiration. |
| pm_runtime_mark_last_busy() | Bookkeeping | Updates power.last_busy to the current time, which is what the autosuspend timer’s expiration is calculated from. |
| pm_runtime_put_autosuspend() | Asynchronous | Decrements the usage count and, if it reaches zero, arms/queues an autosuspend request honoring power.autosuspend_delay, without blocking the caller. |
| pm_runtime_put_sync_autosuspend() | Synchronous | Same intent as pm_runtime_put_autosuspend(), but evaluated on the calling thread rather than deferred to the workqueue. |
| pm_schedule_suspend() | Asynchronous | Arms the suspend timer with an explicit delay in milliseconds, overriding power.autosuspend_delay for that one request. |
The get_sync / mark_last_busy / put_autosuspend Pattern
Nearly every mainline sensor and IIO driver that supports runtime PM autosuspend follows the same three-step shape around a hardware access, whether that access happens inside probe(), a sysfs show handler, or an IIO read_raw() callback:
/* 1. Guarantee the device is powered before touching it */
ret = pm_runtime_resume_and_get(dev);
if (ret < 0)
return ret;
/* 2. Talk to the hardware */
value = ep_sensor_read_register(data, EP_SENSOR_REG_DATA);
/* 3. Mark activity, then request an autosuspend instead of an immediate one */
pm_runtime_mark_last_busy(dev);
pm_runtime_put_autosuspend(dev);
Step one guarantees the device is active before any register is touched, blocking if necessary since this runs in process context. Step three is what makes runtime PM autosuspend actually useful in practice: without the call to pm_runtime_mark_last_busy() immediately beforehand, the autosuspend timer would be computed from a stale timestamp, and without pm_runtime_put_autosuspend() instead of a plain pm_runtime_put() or pm_runtime_put_sync(), the device would suspend immediately rather than waiting out the configured grace period. The full driver built around this exact pattern, including probe() and remove(), is in the companion examples lecture.
Callback Execution Order: Idle, Suspend, And Resume Interactions
The runtime PM core follows a small set of general rules about which callback runs when, and drivers that fight these rules end up with hard-to-reproduce bugs. First, the idle callback (runtime_idle, or the framework’s default handling when a driver does not supply one) only ever runs when the usage count has dropped to zero; as long as any code path holds an outstanding pm_runtime_get()/get_sync() reference, neither idle nor suspend processing is even attempted. Second, when a driver has enabled runtime PM autosuspend and does not override runtime_idle, the default behavior is to treat “going idle” as “request an autosuspend,” which is exactly the queuing step shown in the timer lifecycle diagram above, rather than an unconditional immediate suspend. Third, a resume request always takes precedence over a pending suspend or autosuspend request for the same device: if a new pm_runtime_get() arrives while an autosuspend timer is still counting down, or even after a suspend request has been queued but not yet executed, the resume wins and the suspend is effectively cancelled. Finally, suspend and resume callbacks for a single device are always serialized with respect to each other by the PM core’s internal locking, so a driver’s runtime_suspend and runtime_resume implementations never need to worry about being invoked concurrently for the same device, though they absolutely can be invoked concurrently across different devices, including child devices in a PM hierarchy, which is a separate topic from what this lecture covers.
Common Mistakes And Troubleshooting
- Calling pm_runtime_get_sync() (or resume_and_get()) from an interrupt handler or spinlock-protected section. This is the single most common atomic-context bug in runtime PM code; use the asynchronous pm_runtime_get() there instead and let the workqueue do the blocking work later.
- Forgetting pm_runtime_mark_last_busy() before pm_runtime_put_autosuspend(). Without refreshing power.last_busy first, the timer computes its expiration from a stale timestamp, which can make the device appear to suspend far sooner (or, less commonly, later) than the configured delay suggests.
- Mixing plain put() calls with an autosuspend-enabled device. Once pm_runtime_use_autosuspend() has been called, calling pm_runtime_put() or pm_runtime_put_sync() instead of the _autosuspend() variants defeats the whole purpose, since those calls do not consult power.autosuspend_delay at all.
- Leaking a usage count reference. Every pm_runtime_get()/get_sync() must be matched by exactly one put(); an unmatched get leaves the usage count above zero forever, so the device never idles, never autosuspends, and silently stays powered.
- Ignoring the return value of pm_runtime_get_sync(). A negative return means the resume failed, and continuing to access hardware registers anyway will produce nonsensical reads; prefer pm_runtime_resume_and_get() so a failed resume does not also leave the usage count stuck incremented.
- Setting an unrealistically small autosuspend delay. A delay far shorter than the device’s actual power-up/power-down latency reintroduces the exact thrashing runtime PM autosuspend was meant to prevent, just with extra bookkeeping overhead on top.
Best Practices For Runtime PM Autosuspend
- Call pm_runtime_use_autosuspend() and pm_runtime_set_autosuspend_delay() during probe(), before the device is registered with its subsystem, so no access can race the setup.
- Always pair every hardware access with pm_runtime_mark_last_busy() immediately before pm_runtime_put_autosuspend(), never after.
- Prefer pm_runtime_resume_and_get() over the older pm_runtime_get_sync() in new code, so a failed resume cannot leave a stray usage count reference behind.
- Never call the synchronous helpers from interrupt handlers, tasklets, or any spinlock-held region; use the asynchronous variants there.
- Pick an autosuspend delay based on the device’s real power-up latency plus some margin, not an arbitrary round number, and document why.
- In remove(), resume the device with a synchronous get, drop the reference with pm_runtime_put_noidle(), then call pm_runtime_disable(), so no autosuspend timer can fire against a device mid-teardown.
Summary And Key Takeaways
Synchronous and asynchronous runtime PM operations exist to serve two very different calling contexts: synchronous helpers block the calling thread and run the underlying transition inline, which is correct in process context but unsafe in atomic context; asynchronous helpers instead record the request in the device’s power.request field and queue power.work on the PM workqueue, letting pm_runtime_work() carry out the real transition later on a thread where sleeping is safe, which is exactly why they can be called from an interrupt handler. Runtime PM autosuspend builds on top of this by inserting a configurable grace period, tracked through power.last_busy and power.suspend_timer, between a device going idle and it actually being suspended, which is essential for hardware where power-state transitions are expensive. Drivers enable it with pm_runtime_use_autosuspend() and pm_runtime_set_autosuspend_delay(), and then consistently pair every access with pm_runtime_mark_last_busy() followed by pm_runtime_put_autosuspend(). With sync versus async operation and runtime PM autosuspend now both firmly in place conceptually, the companion examples lecture puts the entire pattern into a complete, original driver you can build and load yourself.
Frequently Asked Questions
What is the actual difference between pm_runtime_get() and pm_runtime_get_sync()?
pm_runtime_get() only records the request and schedules it on the PM workqueue, returning immediately without blocking; pm_runtime_get_sync() runs the resume inline on the calling thread and does not return until the device is confirmed active. The plain get() variant is safe from atomic/IRQ context, while get_sync() is not.
Does runtime PM autosuspend mean the device suspends itself automatically?
No. The name is slightly misleading. It means a timer, armed by pm_runtime_put_autosuspend() and refreshed by pm_runtime_mark_last_busy(), delays the suspend request by a configurable grace period, power.autosuspend_delay, rather than the device suspending immediately on idle or suspending on its own without a request ever being submitted.
Where does the autosuspend delay come from, and can it be changed without recompiling?
The driver sets an initial value with pm_runtime_set_autosuspend_delay() in milliseconds, but it can also be read and overwritten live from user space through the /sys/devices/…/power/autosuspend_delay_ms sysfs attribute for that device, without touching the driver’s source code.
Is it safe to call pm_runtime_get() from an interrupt handler?
Yes, that is exactly what the asynchronous helpers are for. They only set fields and queue work; the actual resume or suspend callback runs later on a PM workqueue thread in normal process context, not on your interrupt stack.
What happens if I forget to call pm_runtime_mark_last_busy() before pm_runtime_put_autosuspend()?
The autosuspend timer’s expiration is computed from power.last_busy, so if that timestamp was never refreshed after the most recent access, the timer may expire based on stale data, which can suspend the device sooner than the configured autosuspend_delay actually intends.
What is the difference between pm_runtime_put_autosuspend() and pm_runtime_put_sync_autosuspend()?
Both request an autosuspend that honors the configured delay rather than an immediate suspend. pm_runtime_put_autosuspend() is asynchronous and returns immediately; pm_runtime_put_sync_autosuspend() evaluates the request on the calling thread instead of deferring it to the PM workqueue.
Can a resume request cancel an autosuspend that is already queued?
Yes. A resume request always takes precedence over a pending suspend or autosuspend request for the same device, so a fresh pm_runtime_get() arriving before the autosuspend timer fires effectively cancels that pending suspend.
Continue Your Linux Kernel Power Management Journey
Ready to see synchronous versus asynchronous runtime PM and runtime PM autosuspend built into a complete, original driver? Move on to the hands-on examples lecture, where the ep_sensor driver gets full autosuspend support with real dmesg output, as part of this free Linux kernel development course.
Try The Hands-On Examples Next: Power Domains And Suspend Sequence