Linux Watchdog Driver Basics
A free Linux kernel development course lecture on the watchdog subsystem, struct watchdog_device, and struct watchdog_info
What You Will Learn
- What a linux watchdog device driver actually protects against, and why every production embedded board needs one
- How the kernel represents a watchdog with
struct watchdog_deviceand the modern watchdog core framework - The role of
timeout,pretimeout, and the internalstatusbitmask - What
struct watchdog_infoexposes to user space throughWDIOC_GETSUPPORT - How to write and register a minimal, original software watchdog driver on a modern kernel
- How to test it from user space and read the resulting kernel log output
Prerequisites
- Working knowledge of C and basic pointer/struct usage
- Familiarity with loadable kernel modules (
insmod/rmmod,module_init/module_exit) - A recent mainline kernel tree (6.x) for header references — this lecture targets the current watchdog core, not any specific old release
- Basic idea of what a hardware timer/counter is (helpful, not mandatory)
Why a Watchdog Exists
Every embedded system eventually hangs — a driver deadlocks, a userspace process spins forever holding a lock, or a bug in your control loop stalls the main thread. On a desktop this just means an annoyed user. On a robot arm, a medical pump, or a remote sensor node sitting on a pole two hundred kilometers from the nearest engineer, a hang that never recovers is a real problem.
A watchdog solves this with a very simple idea: a timer counts down in the background, and some piece of software (usually a userspace daemon) must “pet” it — reset the countdown — before it reaches zero. If the countdown ever reaches zero, the watchdog assumes the system is unhealthy and forces a hardware reset. No pet, no mercy.
The Linux kernel doesn’t implement watchdog logic itself — it provides a watchdog core framework that any hardware watchdog IP (or even a software-emulated one) plugs into, and a uniform /dev/watchdogN character device interface that user space uses regardless of the underlying chip.
struct watchdog_device: The Core Representation
Every registered watchdog in the kernel is represented by one struct watchdog_device. This is the object your driver fills in and hands to the watchdog core — it is not something you invent the layout of, only something you populate correctly. The important fields, in plain terms:
| Field | Meaning |
|---|---|
id | Watchdog ID assigned by the core at registration time — you never set this yourself |
parent | The underlying struct device (your platform/I2C/etc. device) |
info | Pointer to a struct watchdog_info describing identity and capability flags |
ops | Pointer to your driver’s start/stop/ping/set_timeout callbacks |
timeout | Current timeout in seconds — the countdown length |
pretimeout | Seconds before the real timeout at which an early warning interrupt fires (0 = disabled) |
min_timeout / max_timeout | Valid timeout range the hardware supports; if left 0 the core skips range validation and your driver must enforce it |
driver_data | Your private context pointer — set with watchdog_set_drvdata(), read back with watchdog_get_drvdata() |
status | Bitmask of internal state flags (see below) |
On pretimeout: this is the part beginners consistently get backwards. pretimeout is not “seconds until the warning” — it’s “seconds before the real timeout that the warning fires.” If timeout = 60 and pretimeout = 10, the pretimeout interrupt fires at the 50-second mark, giving you 10 seconds to dump diagnostics or panic cleanly before the hardware reset lands. A pretimeout interrupt is typically wired as a non-maskable interrupt (NMI) precisely because the system may already be too unhealthy to service a normal IRQ.
The status Bitmask
The status field tracks internal driver state. The flags a modern watchdog driver actually deals with:
| Flag | Meaning |
|---|---|
WDOG_ACTIVE | Set while the watchdog is currently running/armed |
WDOG_NO_WAY_OUT | The “nowayout” feature — once armed, it can never be stopped from user space, only reset can save you. Set via watchdog_set_nowayout() |
WDOG_STOP_ON_REBOOT | Ensures the watchdog is disarmed on an orderly reboot so it doesn’t fire mid-shutdown |
WDOG_HW_RUNNING | Tells the core the hardware watchdog is already ticking (e.g. it started at boot before your driver probed) — set with set_bit() on your start success path |
WDOG_STOP_ON_UNREGISTER | Stops the watchdog automatically when the driver is unregistered — enable with watchdog_stop_on_unregister() |
struct watchdog_info: What User Space Sees
While watchdog_device is kernel-internal, struct watchdog_info is part of the stable user-space ABI, defined in the UAPI headers. It’s what a userspace watchdog daemon receives back from a WDIOC_GETSUPPORT ioctl call on /dev/watchdog:
options is a capability bitmask — it tells the caller what the device can actually do. The most commonly checked flag is WDIOF_SETTIMEOUT: if it’s set, your driver must implement a set_timeout callback in its watchdog_ops, because the hardware genuinely supports a configurable countdown length rather than a single fixed value.
Building an Original Software Watchdog Driver
Below is a small, original software-emulated watchdog module — ep_soft_wdt — built entirely on a kernel hrtimer, so it needs no special hardware to try out in QEMU or on any Linux box. It registers through the watchdog core exactly the way a real hardware driver would.
// ep_soft_wdt.c — original software watchdog demo for EmbeddedPathashala
#include <linux/module.h>
#include <linux/watchdog.h>
#include <linux/hrtimer.h>
#include <linux/platform_device.h>
#define EP_WDT_DEFAULT_TIMEOUT 20 /* seconds */
struct ep_wdt_priv {
struct hrtimer timer;
struct watchdog_device wdd;
};
static enum hrtimer_restart ep_wdt_fire(struct hrtimer *t)
{
struct ep_wdt_priv *priv = container_of(t, struct ep_wdt_priv, timer);
pr_emerg("ep_soft_wdt: timeout expired, forcing emergency restart\n");
emergency_restart();
return HRTIMER_NORESTART;
}
static int ep_wdt_start(struct watchdog_device *wdd)
{
struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);
ktime_t period = ktime_set(wdd->timeout, 0);
hrtimer_start(&priv->timer, period, HRTIMER_MODE_REL);
set_bit(WDOG_HW_RUNNING, &wdd->status);
return 0;
}
static int ep_wdt_stop(struct watchdog_device *wdd)
{
struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);
hrtimer_cancel(&priv->timer);
return 0;
}
static int ep_wdt_ping(struct watchdog_device *wdd)
{
struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);
ktime_t period = ktime_set(wdd->timeout, 0);
hrtimer_forward_now(&priv->timer, period);
return 0;
}
static int ep_wdt_set_timeout(struct watchdog_device *wdd, unsigned int t)
{
wdd->timeout = t;
return ep_wdt_ping(wdd);
}
static const struct watchdog_ops ep_wdt_ops = {
.owner = THIS_MODULE,
.start = ep_wdt_start,
.stop = ep_wdt_stop,
.ping = ep_wdt_ping,
.set_timeout = ep_wdt_set_timeout,
};
static const struct watchdog_info ep_wdt_info = {
.identity = "ep_soft_wdt",
.options = WDIOF_SETTIMEOUT | WDIOF_KEEPALIVEPING,
};
static struct ep_wdt_priv ep_priv;
static int __init ep_wdt_init(void)
{
int ret;
hrtimer_init(&ep_priv.timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
ep_priv.timer.function = ep_wdt_fire;
ep_priv.wdd.info = &ep_wdt_info;
ep_priv.wdd.ops = &ep_wdt_ops;
ep_priv.wdd.timeout = EP_WDT_DEFAULT_TIMEOUT;
ep_priv.wdd.min_timeout = 1;
ep_priv.wdd.max_timeout = 300;
watchdog_set_drvdata(&ep_priv.wdd, &ep_priv);
watchdog_set_nowayout(&ep_priv.wdd, false);
watchdog_stop_on_reboot(&ep_priv.wdd);
ret = watchdog_register_device(&ep_priv.wdd);
if (ret) {
pr_err("ep_soft_wdt: registration failed (%d)\n", ret);
return ret;
}
pr_info("ep_soft_wdt: registered, default timeout %ds\n",
EP_WDT_DEFAULT_TIMEOUT);
return 0;
}
static void __exit ep_wdt_exit(void)
{
watchdog_unregister_device(&ep_priv.wdd);
}
module_init(ep_wdt_init);
module_exit(ep_wdt_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original software watchdog demo");
Build it against your kernel tree’s headers with a minimal Makefile:
obj-m += ep_soft_wdt.o
all:
make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules
clean:
make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean
Load it and confirm the device node showed up:
$ make
$ sudo insmod ep_soft_wdt.ko
$ dmesg | tail -n 3
[ 812.331420] ep_soft_wdt: registered, default timeout 20s
$ ls /dev/watchdog*
/dev/watchdog /dev/watchdog0
Now exercise it from user space with the standard watchdog tools, or a two-line test program that opens the node and pets it in a loop:
#include <fcntl.h>
#include <unistd.h>
int main(void)
{
int fd = open("/dev/watchdog", O_WRONLY);
for (;;) {
write(fd, "\0", 1); /* pet the dog */
sleep(5);
}
}
Stop running that program and wait past the timeout — you’ll see the module’s pr_emerg line just before the board resets, confirming the countdown really is enforced by the kernel, not just cosmetic.
Common Mistakes and Troubleshooting
- Forgetting WDOG_HW_RUNNING when the hardware auto-starts at boot — some SoCs’ hardware watchdogs are already counting down before Linux even boots. If your
probe()doesn’t detect and reflect that with this flag, the watchdog core assumes it’s idle and a userspace daemon may never open the device in time. - Setting nowayout unconditionally — great for production safety, terrible for development, since you can’t stop the watchdog once armed except by rebooting. Gate it behind a Kconfig or module parameter during bring-up.
- Ignoring min_timeout/max_timeout — if you leave both at 0, the framework does no range checking, and a userspace program can request a nonsensical timeout your hardware can’t represent, silently truncating it.
- Assuming pretimeout counts from “now” — as covered above, it’s relative to the real timeout, not to the moment it was set. Getting this backwards means your “early warning” interrupt fires either far too early or after the real reset already happened.
Best Practices
- Always implement
pingeven for “fire and forget” hardware — some watchdog IP requires the exact keepalive sequence on every refresh, not just a register write. - Use
watchdog_init_timeout()to let device tree/module parameters override your compiled-in default timeout instead of hardcoding it. - Prefer registering through the managed
devm_watchdog_register_device()variant in platform drivers so cleanup on probe failure is automatic. - Keep
min_timeout/max_timeoutaccurate — userspace daemons likewatchdogdread these to pick a safe keepalive interval automatically.
Security consideration: a watchdog device node should never be world-writable in production — anything that can write to /dev/watchdog can also deliberately starve it (by simply not petting it) to force a denial-of-service reboot loop.
Summary and Key Takeaways
- The watchdog core gives every hardware or software watchdog a uniform
/dev/watchdogNinterface viastruct watchdog_deviceandwatchdog_ops. pretimeoutis relative to the real timeout, not to “now” — a frequent source of bugs.struct watchdog_infois the stable UAPI contract exposed to user space throughWDIOC_GETSUPPORT.- Flags like
WDOG_NO_WAY_OUTandWDOG_HW_RUNNINGcontrol safety and boot-time behavior and are easy to get wrong.
This lecture covered the data model. The next lecture in this free linux device drivers course walks through the full watchdog user-space interface — ioctls, /dev/watchdog semantics, and the watchdogd daemon workflow.
Watchdog Driver FAQ
What is a Linux watchdog device driver used for?
It forces an automatic hardware reset if software stops refreshing its countdown timer, recovering embedded systems from hangs and kernel panics without human intervention.
What’s the difference between timeout and pretimeout?
timeout is when the hardware reset actually fires. pretimeout is how many seconds before that reset an early warning interrupt is raised, giving software a last chance to log diagnostics.
What does the nowayout feature do?
Once a watchdog with WDOG_NO_WAY_OUT set is started, it cannot be stopped from user space at all — only a hardware reset ends the countdown, which is intentional for production safety.
Why does WDIOF_SETTIMEOUT matter to driver authors?
If this capability flag is advertised in watchdog_info.options, the driver’s watchdog_ops must implement a working set_timeout callback, since user space is allowed to rely on it.
Can I test a watchdog driver without real watchdog hardware?
Yes — a software-emulated watchdog built on an hrtimer, like the ep_soft_wdt example in this lecture, registers through the exact same watchdog core APIs as real hardware drivers.
What happens if min_timeout and max_timeout are both 0?
The watchdog core performs no range validation on requested timeouts, so the driver itself becomes responsible for rejecting values the hardware can’t actually support.
Is this lecture part of a free Linux kernel development course?
Yes — this is part of EmbeddedPathashala’s ongoing free Linux kernel development course, which also covers character devices, platform drivers, DMA, PCI, and more, all updated for modern kernels.
Continue the Free Embedded Linux Course
Next up: the watchdog user-space interface, ioctls, and building a real keepalive daemon.
Next Lecture Browse Full Course