Linux Driver Wakeup Events Guide
Lecture 10 of the Linux Kernel Power Management series, part of Ravi’s free Linux kernel development course, covering how a driver reports Linux driver wakeup events with pm_wakeup_event() and how the kernel accounts for them.
Linux driver wakeup events are how a device tells the power management core “something just happened that the system needs to be awake to handle,” whether that event arrives while the system is still running or in the middle of pulling itself out of suspend. The previous lecture in this series introduced the wakeup source object itself, the thing a device attaches to dev->power.wakeup once userspace opts it in. This lecture goes one level deeper: the exact runtime mechanics of reporting a wakeup event from inside a real interrupt handler with pm_wakeup_event(), what the kernel does internally when that call happens, how device_may_wakeup() lets a driver honor the user’s policy choice, and the full set of sysfs and debugfs statistics you will actually use when a device either refuses to let the system sleep or refuses to wake it back up.
Key Terms In This Lecture
What You Will Learn
By the end of this lecture you will be able to:
Prerequisites
- Having read the previous Wakeup Source Fundamentals lecture in this series, since this lecture assumes you already know what a struct wakeup_source is, how device_init_wakeup() attaches one to a device, and how the /sys/devices/…/power/wakeup attribute toggles it.
- Comfort with Linux interrupt handling: requesting an IRQ, top-half/threaded-handler basics, and enable_irq_wake()/disable_irq_wake().
- Having read the earlier suspend/resume callback chain lecture in this series, since suspend() and resume() device callbacks are where device_may_wakeup() gets checked.
- Root access on a Linux system or VM for the companion hands-on examples lecture that follows this one.
What Actually Happens When You Call pm_wakeup_event()
A driver reports a wakeup event by calling pm_wakeup_event(struct device *dev, unsigned int msec), declared in include/linux/pm_wakeup.h. In current mainline kernels this is a thin, always-inline wrapper: it calls pm_wakeup_dev_event(dev, msec, false), which in turn locks dev->power.lock and calls pm_wakeup_ws_event() on the device’s attached wakeup source. The trailing boolean is the “hard” flag; pm_wakeup_event() always passes false, reserving the true case for a separate helper, pm_wakeup_hard_event(), used when an event needs to forcibly abort a suspend already in progress and pull the system out of suspend-to-idle immediately. For the overwhelming majority of drivers, an interrupt line that should be able to bring the system out of sleep, a power button, a network card seeing a magic packet, a touch controller seeing a tap, pm_wakeup_event() is the call you want, and it is safe to call directly from interrupt context because every lock it takes is an IRQ-safe spinlock.
Calling pm_wakeup_event() on a device that has never been made wakeup-capable, or whose wakeup source has not been created because the user disabled it through sysfs, is harmless: dev->power.wakeup is simply NULL, and pm_wakeup_dev_event() returns immediately without doing anything. This is exactly why the “activate a wakeup source” step described below is entirely conditional on that source object existing in the first place, which ties directly into why device_may_wakeup() matters in your suspend/resume callbacks, covered later in this lecture.
pm_wakeup_event() Call Path
The msec Argument: Immediate vs Timed Deactivation
The single most consequential decision when wiring up a Linux driver wakeup event is what you pass for msec, because it decides how the kernel closes out the “no suspend” window that event opens.
Pass 0 and the wakeup source is activated and deactivated in the same call: the PM core reports the event, immediately increments its statistics, and then immediately deactivates the source again before pm_wakeup_event() even returns. This is correct when the event itself is the entire story, nothing downstream needs extra time to process it before the system is allowed to go back to sleep. A momentary “poke” that just needs to abort an in-progress suspend transition, or confirm the system is awake for an instant, fits this case.
Pass a nonzero value and the source stays active for at least that many milliseconds. Internally, pm_wakeup_ws_event() converts msec to jiffies and arms (or re-arms, if a later expiration would push it further out) a per-source kernel timer via mod_timer(). Until that timer fires, the “in progress” wakeup event counter stays nonzero, which is precisely the signal pm_wakeup_pending() and the suspend path check before allowing a transition into a sleep state to actually complete. This is the case that matters for hardware wakeup events: a button press or network packet is detected by the IRQ handler, but the work that needs to happen because of it, waking up a userspace daemon, running a bottom-half, letting an input event actually get delivered, takes real wall-clock time. Give that window too little time and the system can race back to sleep before the event has been fully handled; give it a large fixed value out of caution and you have effectively disabled a legitimate suspend for that duration on every single event. A driver can also close the window early and explicitly with pm_relax() if it can positively confirm processing finished before the timer would have expired anyway; that call is safe to make even after pm_wakeup_event() already scheduled a timer, whichever happens first wins.
Inside The Wakeup Source Statistics Model
Every struct wakeup_source carries a small set of counters and ktime_t fields that the PM core updates on every activation and deactivation, and understanding this update sequence is what makes the sysfs attributes in the next section legible rather than opaque numbers.
On activation (the source transitions from idle to active, which happens at most once per “no suspend” window even if pm_wakeup_event() is called repeatedly while it is already active): the source’s active flag is set, its active_count is incremented, and its last_time timestamp is set to the current monotonic clock. Separately, every single call to pm_wakeup_event(), active or not, increments event_count, and if the kernel is currently in the middle of checking whether wakeup events should abort a suspend attempt, it also increments wakeup_count, the field the wakeup_abort_count sysfs attribute exposes.
On deactivation, whether triggered by the msec timer expiring or by an explicit pm_relax() call, the elapsed active duration is computed from last_time to now, added into total_time, and compared against the source’s current max_time, updating it if this activation ran longer than any before it. last_time is then reset to the current moment, marking either “when this source last went active” or, once deactivated, effectively “when this source was last touched.” If deactivation happened specifically because the timer expired rather than because something explicitly relaxed it, expire_count is incremented as well, a signal that the msec value you chose is the thing actually closing the window, not application-level processing finishing early.
Wakeup Source Activation And Deactivation
device_may_wakeup(): Respecting User Space Wakeup Policy
A device being physically capable of generating a wakeup signal and a device being permitted to do so right now are two separate facts, and the kernel keeps them as two separate checks. device_can_wakeup(dev) answers only the hardware-capability question, it reflects dev->power.can_wakeup, set once at probe time (typically via device_init_wakeup() or devm_device_init_wakeup()) and rarely touched again. device_may_wakeup(dev) answers the policy question that matters at suspend time: it returns true only when the device is both capable and currently has an attached wakeup source, meaning dev->power.can_wakeup is set AND dev->power.wakeup is non-NULL. That second condition is exactly what toggling the /sys/devices/…/power/wakeup attribute controls: writing “enabled” calls device_wakeup_enable(), which creates and attaches a fresh wakeup_source object; writing “disabled” calls device_wakeup_disable(), which detaches and destroys it. So device_may_wakeup() is, in effect, a live read of that sysfs attribute from kernel context.
This is why a correctly written suspend() and resume() pair checks device_may_wakeup() before calling enable_irq_wake() or disable_irq_wake(), rather than calling either helper unconditionally. There are two concrete failure modes if you skip the check. First, calling enable_irq_wake() on every suspend regardless of user policy arms an interrupt line as a system wakeup source even when the user explicitly disabled that behavior through sysfs, silently overriding their choice and, on some interrupt controllers, keeping a rail powered that the user expected to go fully quiet. Second, and more subtle: enable_irq_wake()/disable_irq_wake() calls must balance, exactly one disable for every enable on a given IRQ. If a driver calls enable_irq_wake() unconditionally in suspend() but the user flips the wakeup attribute to “disabled” between one suspend/resume cycle and the next, an unconditional disable_irq_wake() in resume() will under-run that balance and the kernel will log a “Unbalanced IRQ wake disable” warning, or, worse, leave the wake-enable count on that IRQ line permanently elevated. Checking device_may_wakeup() symmetrically in both callbacks keeps the enable/disable calls paired exactly to what was actually armed.
device_may_wakeup() In The Suspend/Resume Pair
Debugging Wakeup Sources: sysfs Attributes And What Each One Tells You
Once a device’s wakeup source exists, the kernel exposes its live statistics as a set of read-only files directly under that device’s /sys/devices/…/power/ directory. Every one of these attributes disappears entirely if the device is not wakeup-capable, and reads back empty (not zero, empty) if the device is capable but currently disabled through the wakeup attribute, a distinction worth remembering when a script parsing these files gets an unexpected blank line instead of a number.
| sysfs attribute | What it measures | When you’d check it |
|---|---|---|
| power/wakeup | The policy switch itself: “enabled” or “disabled”. Writable by userspace. | First thing to check when a device seemingly does nothing at all on suspend/resume; if this reads “disabled”, device_may_wakeup() will always be false. |
| power/wakeup_count | Total number of times pm_wakeup_event()/pm_wakeup_event()-style reporting has fired for this device (event_count). | Confirming the IRQ handler is actually being reached and is calling pm_wakeup_event() at all. |
| power/wakeup_active_count | Number of times the source transitioned from idle to active (active_count), i.e. how many distinct “no suspend” windows it opened. | Comparing against wakeup_count to see how many raw event reports actually resulted in a fresh activation versus landing on an already-active source. |
| power/wakeup_abort_count | Number of times an event from this device may have aborted a suspend transition already in progress (the internal wakeup_count field). | Diagnosing “the system tried to suspend but immediately came back” — a high value here points squarely at this device as the culprit. |
| power/wakeup_expire_count | Number of times this source’s activation was ended by its msec timer expiring rather than by an explicit pm_relax(). | Tuning the msec value passed to pm_wakeup_event(); consistently hitting expire rather than an explicit relax suggests the timeout, not real processing time, is what is holding sleep off. |
| power/wakeup_active | Instantaneous 1 or 0: is a wakeup event from this device being processed right now. | Live polling while a suspend attempt is stuck, to catch a source that is active at the exact moment you’re investigating. |
| power/wakeup_total_time_ms | Cumulative milliseconds this source has spent active, summed across every activation (total_time). | Quantifying the total sleep-blocking cost of a device over an uptime period, useful for battery-life regressions. |
| power/wakeup_max_time_ms | The single longest continuous activation this source has ever had (max_time). | Catching one pathological, unusually long wake-lock-style hold rather than many short ones. |
| power/wakeup_last_time_ms | The monotonic clock reading, in milliseconds, of the last time this source was activated (last_time), a timestamp, not a duration. | Confirming recency: has this device signaled anything at all recently, and does that line up with when a suspend attempt actually failed. |
| power/wakeup_prevent_sleep_time_ms | Total time this source has been actively blocking opportunistic transitions to sleep under CONFIG_PM_AUTOSLEEP (prevent_sleep_time). | Android-style or autosleep-enabled systems where a device is suspected of preventing the whole system from ever reaching an idle sleep state. |
For a system-wide view instead of one device at a time, debugfs exposes every registered wakeup source in a single table at /sys/kernel/debug/wakeup_sources, formatted directly from the same struct wakeup_source fields: name, active_count, event_count, wakeup_count, expire_count, active_since, total_time, max_time, last_change, and prevent_suspend_time. This is almost always the faster starting point when you do not yet know which device is at fault, scan the whole table for the largest total_time or the most recent last_change, and only then drill into that specific device’s individual sysfs attributes for finer detail.
Common Mistakes And Troubleshooting
- Passing msec=0 for an event that needs asynchronous processing. The wakeup source deactivates before the event has actually been handled downstream, and the system is free to race back into suspend before, for example, a userspace daemon has had a chance to act on the event.
- Forgetting device_init_wakeup()/devm_device_init_wakeup() in probe(). Without it, dev->power.can_wakeup is never set, the power/wakeup sysfs attribute never appears at all, and pm_wakeup_event() silently does nothing because there is no wakeup source to attach the event to.
- Calling enable_irq_wake()/disable_irq_wake() unconditionally instead of guarding with device_may_wakeup(). This both overrides explicit user policy and risks an unbalanced enable/disable pair if the policy changes between a suspend and the matching resume.
- Reading an empty string from a wakeup_* attribute and assuming it means zero. Empty means the device currently has wakeup disabled via the power/wakeup attribute; the file only reports real numbers once a wakeup source is actually attached.
- Confusing wakeup_abort_count with wakeup_expire_count. Abort count is about a suspend transition being aborted because of this event; expire count is about this source’s own timer running out. A device can have a high expire_count while never once aborting anything.
- Choosing a huge fixed msec “to be safe.” Every event then holds the system awake for that full duration regardless of how quickly it was actually processed, showing up as inflated wakeup_total_time_ms and, on autosleep-enabled systems, wakeup_prevent_sleep_time_ms.
Best Practices
- Call pm_wakeup_event(dev, msec) as close to the top of the IRQ handler as possible, before any slower processing, so the “no suspend” window opens immediately rather than after a delay that has already let the system start suspending.
- Choose msec based on genuine downstream processing latency, not a guess; if a bottom half or input event delivery reliably completes within, say, 100ms, do not pass 2000.
- Always pair every device_may_wakeup() check in suspend() with the identical check in resume(), so enable_irq_wake()/disable_irq_wake() calls stay balanced across policy changes.
- Prefer /sys/kernel/debug/wakeup_sources for a first-pass, system-wide scan before drilling into any single device’s sysfs attributes.
- Watch wakeup_prevent_sleep_time_ms specifically on autosleep-enabled builds; it is the one attribute purpose-built to catch a device that is quietly vetoing opportunistic sleep.
- Use pm_relax() explicitly when a driver can positively confirm event processing has finished, rather than relying solely on the msec timeout to close every window.
Summary And Key Takeaways
Reporting a Linux driver wakeup event correctly comes down to three things working together: pm_wakeup_event(dev, msec) from the IRQ handler, where msec decides whether the “no suspend” window closes instantly or stays open for a bounded, tunable duration; device_may_wakeup() checked symmetrically in suspend() and resume() so enable_irq_wake()/disable_irq_wake() calls always match what the user actually opted into through the power/wakeup sysfs attribute; and the wakeup source statistics, event_count, active_count, wakeup_count, expire_count, total_time, max_time, last_time, and prevent_sleep_time, that the PM core updates on every activation and deactivation and that surface as the wakeup_* sysfs attributes and the debugfs wakeup_sources table. Together these give you everything needed to answer both classic wakeup-source bug reports: “why won’t this device let the system sleep” and “why didn’t this device wake the system up.” The next lecture in this free Linux kernel development course closes out the chapter by covering IRQF_NO_SUSPEND and pulling the whole power management chapter together.
Frequently Asked Questions
Does pm_wakeup_event() work if called from a threaded IRQ handler instead of the primary handler?
Yes. pm_wakeup_event() only takes IRQ-safe spinlocks internally, so it is safe from both hard-IRQ context and a threaded handler’s process context. Many real drivers call it from the threaded half specifically because that is where the actual event data has been read out.
What is the difference between pm_wakeup_event() and pm_stay_awake()/pm_relax()?
pm_wakeup_event(dev, msec) both reports the event and schedules its own deactivation, either immediately (msec=0) or via a timer (msec>0). pm_stay_awake()/pm_relax() is the lower-level pair for events whose processing lifetime the driver can track explicitly: stay_awake opens the window, and the driver calls relax itself once done, with no timer involved at all.
Why does device_may_wakeup() check dev->power.wakeup instead of just dev->power.can_wakeup?
Because can_wakeup only reflects hardware capability, set once at probe time. Whether the device is actually permitted to wake the system right now is a live userspace policy decision, represented by whether a wakeup_source object is currently attached, which is exactly what the power/wakeup sysfs attribute controls.
What happens if I call pm_wakeup_event() on a device where wakeup is disabled via sysfs?
Nothing happens. dev->power.wakeup is NULL in that state, and pm_wakeup_dev_event() returns immediately without touching any statistics. This is intentional: the call is always safe to make unconditionally from an IRQ handler regardless of current policy.
Why is wakeup_last_time_ms a timestamp and not a duration like wakeup_total_time_ms?
It is a direct millisecond reading of the monotonic clock at the moment the source was last activated, not an elapsed time. Comparing it against the current monotonic clock (available from a tool like /proc/uptime scaled appropriately) tells you how long ago the last event actually fired.
Can a device’s wakeup_expire_count go up without wakeup_abort_count ever increasing?
Yes, and it is common. expire_count only tracks how this source’s own activation window closed (timer expiry vs explicit relax). abort_count only increments when an event happens to land during an active suspend-check window and is judged to have aborted that specific transition. A device can fire, time out, and close its window many times while the system is fully awake, with abort_count staying at zero the entire time.
Is /sys/kernel/debug/wakeup_sources always available?
Only if debugfs is mounted and CONFIG_DEBUG_FS is enabled in the running kernel. The file itself is registered unconditionally by the wakeup sources framework at postcore_initcall time, so if debugfs is mounted anywhere, it will be present at that path.
Ready To Wire Up A Real Wakeup IRQ?
You now understand exactly what pm_wakeup_event() does internally, how device_may_wakeup() enforces user policy, and how to read every wakeup-source statistic in sysfs and debugfs, as part of this free Linux kernel development course. The companion hands-on lecture builds a complete IRQ-driven driver around all of it.
Try The Hands-On Examples Back To Course Index