Linux IRQF_NO_SUSPEND Flag Explained-Free Linux Device Drivers Course

Linux IRQF_NO_SUSPEND Flag Explained
PREV_LEC | NEXT_LEC

Linux IRQF_NO_SUSPEND Flag Explained

The final lecture in our free Linux kernel development course power management series — how one interrupt flag keeps a handler alive through the entire suspend-resume cycle, and why it is one of the most misused flags in the IRQ subsystem.

Lecture 21 of 22

Series: Linux Kernel PM

Final PM Lecture

The IRQF_NO_SUSPEND flag is a single bit in the Linux kernel’s interrupt request flags that decides whether a device’s interrupt handler is allowed to keep running while the rest of the system is asleep. It sounds like a small detail, but get it wrong and you either lose timer ticks during suspend-to-idle, or you crash the kernel by letting a handler touch hardware that has already been powered off. This lecture closes out our Linux Kernel Power Management series by digging into exactly what IRQF_NO_SUSPEND does at the code level in the current mainline kernel, how it differs from enable_irq_wake(), why mixing it with IRQF_SHARED is dangerous, and how everything we covered across this free Linux kernel development course power management series — runtime PM, system sleep, CPU idle/frequency scaling, and wakeup sources — fits together in a real driver.

What You Will Learn

IRQF_NO_SUSPEND internals suspend_device_irqs() flow no_suspend_depth counter Chained irqchip exemption enable_irq_wake() vs IRQF_NO_SUSPEND dev_pm_set_wake_irq() IRQF_SHARED danger IRQF_COND_SUSPEND Series-wide PM synthesis

Prerequisites

This lecture assumes you have already gone through the previous lecture pair on Driver Wakeup Events (enable_irq_wake(), device_init_wakeup(), and wakeup source accounting in sysfs/debugfs). Ideally you have followed this entire free Linux kernel development course power management series from the beginning — Runtime PM, the PM core and dev_pm_ops, system sleep states, CPUFreq, CPUIdle, and Thermal — since this lecture treats all of that as background and spends its time on synthesis rather than repeating fundamentals. A working knowledge of request_irq(), IRQ handler context, and basic suspend/resume callback ordering is required.

The Suspend IRQ Problem IRQF_NO_SUSPEND Solves

When the kernel suspends the system, it does not just flip devices into low-power states and stop. Part of the process is a dedicated step, implemented by suspend_device_irqs(), that walks every registered interrupt descriptor and disables it. This happens after every device’s ->prepare, ->suspend, and ->suspend_late callbacks have already run, and it is the mechanism that opens the so-called noirq window — the period during which device interrupts are not supposed to fire at all. The reasoning is straightforward: once a device has entered its low-power state, there is no legitimate reason for its interrupt line to trigger, and if some driver has not suspended cleanly yet, it is safer to block its interrupts than to let a handler run against hardware in an unknown state.

The problem is that not every interrupt in the system belongs to a device that is going to sleep. Some interrupt sources are structural to the kernel itself and must keep functioning throughout the entire suspend-resume cycle, including the noirq phase and the window during which non-boot CPUs are taken offline and later brought back online. This is exactly the gap that the IRQF_NO_SUSPEND flag closes.

What IRQF_NO_SUSPEND Does at the Code Level

IRQF_NO_SUSPEND is a request flag passed to request_irq() / request_threaded_irq() (bit value 0x00004000 in current mainline include/linux/interrupt.h). When an interrupt action is installed, irq_pm_install_action() increments a per-descriptor counter called no_suspend_depth on the underlying irq_desc if the flag is set. Later, when the system goes to sleep, suspend_device_irqs() iterates every IRQ descriptor and calls an internal helper that begins with this exact check:

if (!desc->action || irq_desc_is_chained(desc) ||
    desc->no_suspend_depth)
        return false;

If no_suspend_depth is non-zero for that descriptor, the function returns immediately without disabling the IRQ or marking it suspended. In plain terms: the interrupt line is simply left alone. It is never masked, never disabled, and its handler is invoked exactly as it would be during normal operation — right through the noirq phase, right through CPU hotplug for suspend, and all the way to resume_device_irqs() on the way back up. IRQF_NO_SUSPEND does not add any special resume-time logic of its own; there is nothing to re-enable because the IRQ was never touched in the first place.

One detail worth knowing: the composite flag IRQF_TIMER is itself defined as (__IRQF_TIMER | IRQF_NO_SUSPEND | IRQF_NO_THREAD). That means every interrupt registered as a kernel timer interrupt already carries IRQF_NO_SUSPEND by definition — driver authors never need to add it manually for core timer/clockevent interrupts, because the timer subsystem bakes it in at the flag-definition level.

Suspend-Resume IRQ Timeline

1. Running →

dpm_suspend(): ->prepare, ->suspend, ->suspend_late for all devices

2. ↓

suspend_device_irqs() runs — noirq window begins

3. ↓

Normal IRQs: disabled and masked, handlers do NOT run

4. ↓

Wakeup-armed IRQs (enable_irq_wake): left enabled, but IRQD_WAKEUP_ARMED — first trigger disables the line and calls pm_system_irq_wakeup(), handler does NOT run

5. ↓

IRQF_NO_SUSPEND IRQs: never touched — handler keeps running normally

6. ↓

Non-boot CPUs taken offline, ->suspend_noirq callbacks run

7. ↓

System sleep state entered (S2Idle / Suspend-to-RAM)

8. ↓

Resume: CPUs online, ->resume_noirq callbacks run

9. ↓

resume_device_irqs() re-enables normal and wakeup IRQs — noirq window ends

10. ↓

->resume_early, ->resume, ->complete for all devices

Why Some Interrupts Need IRQF_NO_SUSPEND and Most Do Not

The kernel documentation for this area is explicit that legitimate IRQF_NO_SUSPEND users are a small, special-purpose set: timer interrupts (already covered via IRQF_TIMER), inter-processor interrupts (IPIs) used for CPU coordination during the hotplug dance, and a handful of platform-specific interrupts that the SoC genuinely cannot function without even while “asleep.” A well-known real example is a hardware watchdog kick or keepalive interrupt on platforms where suspend-to-idle keeps the SoC clocked and only idles the CPUs — if the watchdog is not serviced during that window, it can reset the board mid-suspend. Another real case is chained or cascaded interrupt controllers: a GPIO expander or secondary interrupt controller that fans out several other device interrupts, including ones used as system wakeup sources, needs its own dispatch logic to keep running so that the wakeup interrupts behind it can be recognized and routed at all.

Interestingly, chained interrupt controllers get a second, independent form of protection in the kernel: the same check inside suspend_device_irqs() also tests irq_desc_is_chained(desc), and chained descriptors are skipped regardless of whether IRQF_NO_SUSPEND was ever requested on them. This is a structural exemption for the parent dispatcher IRQ, not a substitute for the flag — any individual leaf device sitting behind that cascaded controller that needs its own handler to keep executing during the noirq window must still request IRQF_NO_SUSPEND explicitly.

Every other interrupt — the overwhelming majority in any system — should not use this flag. A network card, a storage controller, a touchscreen, a regular GPIO-based sensor: none of these have any business running their interrupt handler while their device is powered down or clocked off. Letting their handlers fire anyway is not a convenience, it is a way to touch registers or DMA buffers of hardware that a driver’s own ->suspend() callback has already put to sleep.

IRQF_NO_SUSPEND vs enable_irq_wake() and dev_pm_set_wake_irq()

This is the distinction the old chapter-summary material only mentions in passing, and it is the single most common point of confusion around IRQF_NO_SUSPEND. These are two unrelated mechanisms that happen to sit next to each other in the same header file:

  • IRQF_NO_SUSPEND controls whether an interrupt line stays enabled and actively invoking its handler during suspend. It has nothing to do with whether that interrupt can bring the system out of sleep.
  • enable_irq_wake() (and the higher-level dev_pm_set_wake_irq() / dev_pm_set_dedicated_wake_irq() helpers built on top of it) mark an IRQ as a system wakeup source. They flip on platform-specific wakeup routing logic so that a signal on that line can abort an in-progress suspend or bring the system back from sleep — but the handler itself is explicitly not invoked while the IRQ is in this “armed” state.

The kernel’s own interrupt PM code makes this split obvious. A wakeup-armed IRQ goes through irq_pm_handle_wakeup(): on the first trigger it clears the armed flag, marks the IRQ as suspended and pending, disables it, and calls pm_system_irq_wakeup() to notify the PM core — the handler function itself never runs at that point. An IRQF_NO_SUSPEND IRQ, by contrast, is never disabled and its handler runs exactly as if the system were fully awake. Two completely different code paths, two completely different jobs.

AspectIRQF_NO_SUSPENDenable_irq_wake() / dev_pm_set_wake_irq()
PurposeKeep an IRQ line enabled and its handler running throughout the suspend-resume cycleMark an IRQ as capable of waking the system from a sleep state
Effect on suspend_device_irqs()IRQ is skipped entirely — never disabled, never maskedIRQ is armed (IRQD_WAKEUP_ARMED); on trigger it is disabled and the event is reported to the PM core
Does the handler run during suspend?Yes, normally, every time it firesNo — the first trigger aborts/ends suspend instead of invoking the driver’s handler
Set whereFlag passed to request_irq() / request_threaded_irq()Called from a driver’s ->suspend()/->resume() PM ops, usually gated by device_may_wakeup()
Common mistakeAssuming it also wakes the system from sleep — it does notAssuming the handler still executes on the wakeup trigger — it generally does not

Why IRQF_SHARED Plus IRQF_NO_SUSPEND Is Dangerous

IRQF_NO_SUSPEND is a property of the entire interrupt descriptor, not of one handler on a shared line. Look again at the kernel check: desc->no_suspend_depth is incremented once for every installed action that requests the flag, and suspend_device_irqs() tests the descriptor-wide counter, not a per-handler flag. That means if even one driver sharing a physical IRQ line sets IRQF_NO_SUSPEND, the entire line is exempted from suspension — and every other handler chained onto that same line, including ones belonging to devices that were correctly and fully suspended moments earlier, will still be invoked when the line triggers.

Concretely: imagine two devices sharing one physical interrupt line. Device A is an always-on companion chip that legitimately needs IRQF_NO_SUSPEND. Device B is a regular peripheral behind an I2C or SPI bus that has already been clocked off and had its regulator dropped by its own ->suspend() callback. If Device A’s driver requests the shared line with IRQF_SHARED | IRQF_NO_SUSPEND, then the moment that line fires during the noirq window, Device B’s handler runs too — and promptly tries to touch a bus or register that is no longer powered. Depending on the platform this manifests as a silent hang, a bus timeout splat, or in the worst case a full system lockup that is extremely painful to reproduce and debug, because it only happens on shared-IRQ hardware and only during suspend.

Mainline kernel documentation is direct about this: using IRQF_NO_SUSPEND and IRQF_SHARED together should be avoided. There is exactly one sanctioned narrow exception, described in the next section.

IRQF_COND_SUSPEND: The One Safe Exception

Current mainline kernels define a companion flag, IRQF_COND_SUSPEND, precisely for the rare legitimate case where a wakeup-capable driver must share a line with an IRQF_NO_SUSPEND user. It is not a free pass — it comes with strict obligations. A driver requesting IRQF_COND_SUSPEND must be able to tell a genuine wakeup event apart from a spurious interrupt caused by the shared line firing for the other device, it must still call enable_irq_wake() so the IRQ functions as a real wakeup source, and it must report genuine wakeup events itself via pm_system_wakeup() instead of relying on the normal armed-IRQ machinery. If a driver cannot meet all three of those requirements, IRQF_COND_SUSPEND is not a safe substitute for redesigning the interrupt topology — most real drivers should simply avoid sharing an IRQF_NO_SUSPEND line at all.

Common Mistakes and Troubleshooting

  • Treating IRQF_NO_SUSPEND as a wakeup flag. It only keeps a handler running; it never wakes a sleeping system. Use enable_irq_wake()/dev_pm_set_wake_irq() for that job.
  • IRQF_SHARED + IRQF_NO_SUSPEND on a line with suspendable peers. This is the anti-pattern covered in the companion example lecture — every handler on the line keeps firing, even for devices that are properly asleep.
  • Doing slow or blocking work in an IRQF_NO_SUSPEND hard-IRQ handler. The noirq window and CPU-hotplug transition are latency-sensitive; a handler that blocks here can stall suspend/resume for the whole system.
  • Forgetting that chained irqchips are already exempt. Adding IRQF_NO_SUSPEND to a chained parent controller is usually redundant — the kernel already skips chained descriptors via irq_desc_is_chained(). The flag is for leaf handlers that specifically need to keep running.
  • Assuming resume automatically re-arms things. IRQF_NO_SUSPEND IRQs need no resume-time re-enable logic because they were never disabled — if your driver adds one anyway “just in case,” it is a sign the flag was misunderstood.

Best Practices

  • Default to leaving IRQF_NO_SUSPEND off. Only add it when you can name the specific reason the handler must run during the noirq window.
  • Never combine IRQF_NO_SUSPEND with IRQF_SHARED unless every handler on that line meets the strict IRQF_COND_SUSPEND requirements.
  • Keep IRQF_NO_SUSPEND handlers short, non-blocking, and free of dependencies on subsystems that may themselves be suspended (clocks, regulators, buses).
  • Use enable_irq_wake()/dev_pm_set_wake_irq() for “get me out of sleep,” and IRQF_NO_SUSPEND for “keep running while asleep” — never conflate the two in driver documentation or code comments.
  • Prefer isolating an always-on interrupt on its own line where the hardware allows it, so IRQF_SHARED never becomes a temptation in the first place.
  • When touching a chained/cascaded controller, verify with cat /proc/interrupts and driver source whether the exemption already comes from the chained-descriptor path before adding the flag manually.

Series Wrap-Up: The Four Pillars of Linux Kernel Power Management

This closes our Linux Kernel Power Management series, and it is worth stepping back to see how everything connects. Across this free Linux kernel development course, the power management subsystem breaks down into four pillars that solve different problems at different scopes:

The Four Pillars of Linux Kernel Power Management

Runtime PM

Per-device, dynamic — pm_runtime_get/put, autosuspend, idles a single device while the system stays fully awake

System Sleep

Whole-system states — suspend-to-idle, suspend-to-RAM, hibernation, driven by dev_pm_ops ->suspend/->resume and the noirq phases

CPU Idle / CPUFreq / Thermal

Dynamic power interfaces — idle-state selection between instructions, frequency/voltage scaling under load, thermal-driven throttling

Wakeup Sources

Getting back out of sleep — device_init_wakeup(), enable_irq_wake(), IRQF_NO_SUSPEND for the rare handler that must never stop

Runtime PM answers “can this one device nap while everything else keeps working?” System sleep answers “can the whole machine nap together?” CPUIdle, CPUFreq, and Thermal answer “how much power does this device or CPU need right now, moment to moment, even while fully awake?” And wakeup sources — including the narrow, carefully-scoped role of IRQF_NO_SUSPEND — answer “how does anything get back out of sleep at all?” A real production driver almost never uses just one of these. A typical modern peripheral driver enables runtime PM in its probe function so the device autosuspends when idle, implements SIMPLE_DEV_PM_OPS or DEFINE_SIMPLE_DEV_PM_OPS for system-wide suspend/resume, calls device_init_wakeup() so user space can opt the device into being a wakeup source via /sys/devices/.../power/wakeup, and — only if the hardware genuinely demands it — marks one specific interrupt IRQF_NO_SUSPEND rather than sprinkling it across every IRQ the driver owns. Understanding where each pillar’s responsibility starts and stops is what separates a driver that merely “supports power management” from one that actually saves power without introducing suspend/resume bugs.

Summary and Key Takeaways

IRQF_NO_SUSPEND keeps a specific interrupt line enabled and its handler actively invoked across the entire suspend-resume cycle, including the noirq phases and CPU hotplug for suspend — it is implemented in mainline as a per-descriptor no_suspend_depth counter checked inside suspend_device_irqs(). It has nothing to do with waking the system from sleep; that job belongs to enable_irq_wake() and the higher-level dev_pm_set_wake_irq() helpers, which use an entirely separate armed/pending code path. Legitimate uses of IRQF_NO_SUSPEND are narrow — timer interrupts (already covered via IRQF_TIMER), IPIs, watchdog-kick style interrupts, and leaf handlers behind chained controllers that must keep dispatching — while the vast majority of interrupts should never carry it. Combining IRQF_NO_SUSPEND with IRQF_SHARED is dangerous because the flag applies to the whole IRQ descriptor: every handler on that line keeps firing, even for devices that suspended correctly, risking access to powered-down hardware. IRQF_COND_SUSPEND exists as a narrow, obligation-heavy exception for genuinely unavoidable sharing.

Zooming out across the whole series: Linux kernel power management rests on four pillars — Runtime PM for per-device dynamic idling, System Sleep for whole-machine states, CPU Idle/CPUFreq/Thermal for continuous power/performance tuning, and Wakeup Sources for getting back out of sleep. A well-behaved production driver combines several of these deliberately, never accidentally, and IRQF_NO_SUSPEND is the sharp, narrow tool reserved for the rare interrupt that truly cannot afford to stop.

Frequently Asked Questions

What does the IRQF_NO_SUSPEND flag actually do in the Linux kernel?

It tells suspend_device_irqs() to leave that interrupt descriptor completely untouched during system suspend, so the handler keeps being invoked normally through the noirq window and CPU hotplug phase, instead of being disabled like most device interrupts.

Does IRQF_NO_SUSPEND make an interrupt a wakeup source?

No. It only keeps the handler running while the system is going to or coming out of sleep. It provides no guarantee that the interrupt will abort suspend or wake a sleeping system — that requires enable_irq_wake() or dev_pm_set_wake_irq().

Why shouldn’t most interrupt handlers use IRQF_NO_SUSPEND?

Because most devices are genuinely powered down, clock-gated, or otherwise inactive during suspend. Letting their handler run anyway means it may touch registers, buses, or memory belonging to hardware that is not in a safe state to be accessed.

What happens if IRQF_NO_SUSPEND is combined with IRQF_SHARED?

The flag applies to the entire shared IRQ line, not just the handler that requested it. Every handler registered on that line will keep firing during suspend, including handlers for devices that were correctly suspended, which can crash or hang the system.

Is there any safe way to share an IRQ between a wakeup device and an IRQF_NO_SUSPEND device?

Only via IRQF_COND_SUSPEND, and only if the wakeup driver can distinguish genuine wakeup events from spurious ones, calls enable_irq_wake(), and reports wakeup itself with pm_system_wakeup(). Without meeting all three conditions, sharing should be avoided entirely.

How is IRQF_NO_SUSPEND different from IRQCHIP_MASK_ON_SUSPEND?

IRQF_NO_SUSPEND is a request flag set by a device driver on an individual interrupt action. IRQCHIP_MASK_ON_SUSPEND is a chip-level flag telling the IRQ core that a particular interrupt controller hardware needs its non-wakeup lines physically masked at suspend because it lacks its own wakeup-configuration facility. They operate at different layers of the IRQ stack.

Do chained/cascaded interrupt controllers need IRQF_NO_SUSPEND?

Usually not on the parent dispatcher itself — the kernel already skips chained descriptors in suspend_device_irqs() regardless of flags. Individual leaf devices behind that controller that need their own handler to keep running during noirq must still request the flag explicitly.

What’s the difference between suspend-to-RAM and suspend-to-idle regarding IRQF_NO_SUSPEND?

The flag behaves the same way in both: the handler keeps running. The practical difference is that suspend-to-idle idles CPUs in a loop waiting for interrupts, so an IRQF_NO_SUSPEND interrupt firing there simply brings a CPU out of idle without causing a system wakeup, whereas on full suspend-to-RAM it runs during the noirq window before hardware is fully powered down.

Power Management Series Complete — Onward to PCI Device Drivers

This wraps up every lecture in the Linux Kernel Power Management series of our free Linux kernel development course. Next, we move into PCI device drivers — enumeration, BARs, MSI/MSI-X interrupts, and how PCI devices plug into the same driver model you have been using throughout this free Linux device drivers course.

Start PCI Device Drivers Review the Full PM Series Index
PREV_LEC | NEXT_LEC

Leave a Reply

Your email address will not be published. Required fields are marked *