Linux Watchdog Driver Basics-Free Linux Device Drivers Course

Linux Watchdog Driver Basics

A free Linux kernel development course lecture on the watchdog subsystem, struct watchdog_device, and struct watchdog_info

Chapter 13 · Lecture 1
Kernel 6.x APIs
~20 min read
linux watchdog device driver free linux kernel development course free linux device drivers course free embedded linux course watchdog_device struct

What You Will Learn

  • What a linux watchdog device driver actually protects against, and why every production embedded board needs one
  • How the kernel represents a watchdog with struct watchdog_device and the modern watchdog core framework
  • The role of timeout, pretimeout, and the internal status bitmask
  • What struct watchdog_info exposes to user space through WDIOC_GETSUPPORT
  • How to write and register a minimal, original software watchdog driver on a modern kernel
  • How to test it from user space and read the resulting kernel log output

Prerequisites

  • Working knowledge of C and basic pointer/struct usage
  • Familiarity with loadable kernel modules (insmod/rmmod, module_init/module_exit)
  • A recent mainline kernel tree (6.x) for header references — this lecture targets the current watchdog core, not any specific old release
  • Basic idea of what a hardware timer/counter is (helpful, not mandatory)

Why a Watchdog Exists

Every embedded system eventually hangs — a driver deadlocks, a userspace process spins forever holding a lock, or a bug in your control loop stalls the main thread. On a desktop this just means an annoyed user. On a robot arm, a medical pump, or a remote sensor node sitting on a pole two hundred kilometers from the nearest engineer, a hang that never recovers is a real problem.

A watchdog solves this with a very simple idea: a timer counts down in the background, and some piece of software (usually a userspace daemon) must “pet” it — reset the countdown — before it reaches zero. If the countdown ever reaches zero, the watchdog assumes the system is unhealthy and forces a hardware reset. No pet, no mercy.

The Linux kernel doesn’t implement watchdog logic itself — it provides a watchdog core framework that any hardware watchdog IP (or even a software-emulated one) plugs into, and a uniform /dev/watchdogN character device interface that user space uses regardless of the underlying chip.

struct watchdog_device: The Core Representation

Every registered watchdog in the kernel is represented by one struct watchdog_device. This is the object your driver fills in and hands to the watchdog core — it is not something you invent the layout of, only something you populate correctly. The important fields, in plain terms:

watchdog_device — key fields
struct watchdog_device { int id; struct device *parent; const struct watchdog_info *info; const struct watchdog_ops *ops; unsigned int bootstatus; unsigned int timeout; unsigned int pretimeout; unsigned int min_timeout; unsigned int max_timeout; void *driver_data; unsigned long status; … };
FieldMeaning
idWatchdog ID assigned by the core at registration time — you never set this yourself
parentThe underlying struct device (your platform/I2C/etc. device)
infoPointer to a struct watchdog_info describing identity and capability flags
opsPointer to your driver’s start/stop/ping/set_timeout callbacks
timeoutCurrent timeout in seconds — the countdown length
pretimeoutSeconds before the real timeout at which an early warning interrupt fires (0 = disabled)
min_timeout / max_timeoutValid timeout range the hardware supports; if left 0 the core skips range validation and your driver must enforce it
driver_dataYour private context pointer — set with watchdog_set_drvdata(), read back with watchdog_get_drvdata()
statusBitmask of internal state flags (see below)

On pretimeout: this is the part beginners consistently get backwards. pretimeout is not “seconds until the warning” — it’s “seconds before the real timeout that the warning fires.” If timeout = 60 and pretimeout = 10, the pretimeout interrupt fires at the 50-second mark, giving you 10 seconds to dump diagnostics or panic cleanly before the hardware reset lands. A pretimeout interrupt is typically wired as a non-maskable interrupt (NMI) precisely because the system may already be too unhealthy to service a normal IRQ.

The status Bitmask

The status field tracks internal driver state. The flags a modern watchdog driver actually deals with:

FlagMeaning
WDOG_ACTIVESet while the watchdog is currently running/armed
WDOG_NO_WAY_OUTThe “nowayout” feature — once armed, it can never be stopped from user space, only reset can save you. Set via watchdog_set_nowayout()
WDOG_STOP_ON_REBOOTEnsures the watchdog is disarmed on an orderly reboot so it doesn’t fire mid-shutdown
WDOG_HW_RUNNINGTells the core the hardware watchdog is already ticking (e.g. it started at boot before your driver probed) — set with set_bit() on your start success path
WDOG_STOP_ON_UNREGISTERStops the watchdog automatically when the driver is unregistered — enable with watchdog_stop_on_unregister()

struct watchdog_info: What User Space Sees

While watchdog_device is kernel-internal, struct watchdog_info is part of the stable user-space ABI, defined in the UAPI headers. It’s what a userspace watchdog daemon receives back from a WDIOC_GETSUPPORT ioctl call on /dev/watchdog:

watchdog_info — UAPI structure
struct watchdog_info { __u32 options; __u32 firmware_version; __u8 identity[32]; };

options is a capability bitmask — it tells the caller what the device can actually do. The most commonly checked flag is WDIOF_SETTIMEOUT: if it’s set, your driver must implement a set_timeout callback in its watchdog_ops, because the hardware genuinely supports a configurable countdown length rather than a single fixed value.

Building an Original Software Watchdog Driver

Below is a small, original software-emulated watchdog module — ep_soft_wdt — built entirely on a kernel hrtimer, so it needs no special hardware to try out in QEMU or on any Linux box. It registers through the watchdog core exactly the way a real hardware driver would.

// ep_soft_wdt.c — original software watchdog demo for EmbeddedPathashala
#include <linux/module.h>
#include <linux/watchdog.h>
#include <linux/hrtimer.h>
#include <linux/platform_device.h>

#define EP_WDT_DEFAULT_TIMEOUT  20   /* seconds */

struct ep_wdt_priv {
    struct hrtimer     timer;
    struct watchdog_device wdd;
};

static enum hrtimer_restart ep_wdt_fire(struct hrtimer *t)
{
    struct ep_wdt_priv *priv = container_of(t, struct ep_wdt_priv, timer);

    pr_emerg("ep_soft_wdt: timeout expired, forcing emergency restart\n");
    emergency_restart();
    return HRTIMER_NORESTART;
}

static int ep_wdt_start(struct watchdog_device *wdd)
{
    struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);
    ktime_t period = ktime_set(wdd->timeout, 0);

    hrtimer_start(&priv->timer, period, HRTIMER_MODE_REL);
    set_bit(WDOG_HW_RUNNING, &wdd->status);
    return 0;
}

static int ep_wdt_stop(struct watchdog_device *wdd)
{
    struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);

    hrtimer_cancel(&priv->timer);
    return 0;
}

static int ep_wdt_ping(struct watchdog_device *wdd)
{
    struct ep_wdt_priv *priv = watchdog_get_drvdata(wdd);
    ktime_t period = ktime_set(wdd->timeout, 0);

    hrtimer_forward_now(&priv->timer, period);
    return 0;
}

static int ep_wdt_set_timeout(struct watchdog_device *wdd, unsigned int t)
{
    wdd->timeout = t;
    return ep_wdt_ping(wdd);
}

static const struct watchdog_ops ep_wdt_ops = {
    .owner       = THIS_MODULE,
    .start       = ep_wdt_start,
    .stop        = ep_wdt_stop,
    .ping        = ep_wdt_ping,
    .set_timeout = ep_wdt_set_timeout,
};

static const struct watchdog_info ep_wdt_info = {
    .identity = "ep_soft_wdt",
    .options  = WDIOF_SETTIMEOUT | WDIOF_KEEPALIVEPING,
};

static struct ep_wdt_priv ep_priv;

static int __init ep_wdt_init(void)
{
    int ret;

    hrtimer_init(&ep_priv.timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
    ep_priv.timer.function = ep_wdt_fire;

    ep_priv.wdd.info    = &ep_wdt_info;
    ep_priv.wdd.ops     = &ep_wdt_ops;
    ep_priv.wdd.timeout = EP_WDT_DEFAULT_TIMEOUT;
    ep_priv.wdd.min_timeout = 1;
    ep_priv.wdd.max_timeout = 300;

    watchdog_set_drvdata(&ep_priv.wdd, &ep_priv);
    watchdog_set_nowayout(&ep_priv.wdd, false);
    watchdog_stop_on_reboot(&ep_priv.wdd);

    ret = watchdog_register_device(&ep_priv.wdd);
    if (ret) {
        pr_err("ep_soft_wdt: registration failed (%d)\n", ret);
        return ret;
    }

    pr_info("ep_soft_wdt: registered, default timeout %ds\n",
            EP_WDT_DEFAULT_TIMEOUT);
    return 0;
}

static void __exit ep_wdt_exit(void)
{
    watchdog_unregister_device(&ep_priv.wdd);
}

module_init(ep_wdt_init);
module_exit(ep_wdt_exit);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala original software watchdog demo");

Build it against your kernel tree’s headers with a minimal Makefile:

obj-m += ep_soft_wdt.o

all:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules

clean:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean

Load it and confirm the device node showed up:

$ make
$ sudo insmod ep_soft_wdt.ko
$ dmesg | tail -n 3
[  812.331420] ep_soft_wdt: registered, default timeout 20s
$ ls /dev/watchdog*
/dev/watchdog  /dev/watchdog0

Now exercise it from user space with the standard watchdog tools, or a two-line test program that opens the node and pets it in a loop:

#include <fcntl.h>
#include <unistd.h>

int main(void)
{
    int fd = open("/dev/watchdog", O_WRONLY);
    for (;;) {
        write(fd, "\0", 1);   /* pet the dog */
        sleep(5);
    }
}

Stop running that program and wait past the timeout — you’ll see the module’s pr_emerg line just before the board resets, confirming the countdown really is enforced by the kernel, not just cosmetic.

Common Mistakes and Troubleshooting

  • Forgetting WDOG_HW_RUNNING when the hardware auto-starts at boot — some SoCs’ hardware watchdogs are already counting down before Linux even boots. If your probe() doesn’t detect and reflect that with this flag, the watchdog core assumes it’s idle and a userspace daemon may never open the device in time.
  • Setting nowayout unconditionally — great for production safety, terrible for development, since you can’t stop the watchdog once armed except by rebooting. Gate it behind a Kconfig or module parameter during bring-up.
  • Ignoring min_timeout/max_timeout — if you leave both at 0, the framework does no range checking, and a userspace program can request a nonsensical timeout your hardware can’t represent, silently truncating it.
  • Assuming pretimeout counts from “now” — as covered above, it’s relative to the real timeout, not to the moment it was set. Getting this backwards means your “early warning” interrupt fires either far too early or after the real reset already happened.

Best Practices

  • Always implement ping even for “fire and forget” hardware — some watchdog IP requires the exact keepalive sequence on every refresh, not just a register write.
  • Use watchdog_init_timeout() to let device tree/module parameters override your compiled-in default timeout instead of hardcoding it.
  • Prefer registering through the managed devm_watchdog_register_device() variant in platform drivers so cleanup on probe failure is automatic.
  • Keep min_timeout/max_timeout accurate — userspace daemons like watchdogd read these to pick a safe keepalive interval automatically.

Security consideration: a watchdog device node should never be world-writable in production — anything that can write to /dev/watchdog can also deliberately starve it (by simply not petting it) to force a denial-of-service reboot loop.

Summary and Key Takeaways

  • The watchdog core gives every hardware or software watchdog a uniform /dev/watchdogN interface via struct watchdog_device and watchdog_ops.
  • pretimeout is relative to the real timeout, not to “now” — a frequent source of bugs.
  • struct watchdog_info is the stable UAPI contract exposed to user space through WDIOC_GETSUPPORT.
  • Flags like WDOG_NO_WAY_OUT and WDOG_HW_RUNNING control safety and boot-time behavior and are easy to get wrong.

This lecture covered the data model. The next lecture in this free linux device drivers course walks through the full watchdog user-space interface — ioctls, /dev/watchdog semantics, and the watchdogd daemon workflow.

Watchdog Driver FAQ

What is a Linux watchdog device driver used for?

It forces an automatic hardware reset if software stops refreshing its countdown timer, recovering embedded systems from hangs and kernel panics without human intervention.

What’s the difference between timeout and pretimeout?

timeout is when the hardware reset actually fires. pretimeout is how many seconds before that reset an early warning interrupt is raised, giving software a last chance to log diagnostics.

What does the nowayout feature do?

Once a watchdog with WDOG_NO_WAY_OUT set is started, it cannot be stopped from user space at all — only a hardware reset ends the countdown, which is intentional for production safety.

Why does WDIOF_SETTIMEOUT matter to driver authors?

If this capability flag is advertised in watchdog_info.options, the driver’s watchdog_ops must implement a working set_timeout callback, since user space is allowed to rely on it.

Can I test a watchdog driver without real watchdog hardware?

Yes — a software-emulated watchdog built on an hrtimer, like the ep_soft_wdt example in this lecture, registers through the exact same watchdog core APIs as real hardware drivers.

What happens if min_timeout and max_timeout are both 0?

The watchdog core performs no range validation on requested timeouts, so the driver itself becomes responsible for rejecting values the hardware can’t actually support.

Is this lecture part of a free Linux kernel development course?

Yes — this is part of EmbeddedPathashala’s ongoing free Linux kernel development course, which also covers character devices, platform drivers, DMA, PCI, and more, all updated for modern kernels.

Continue the Free Embedded Linux Course

Next up: the watchdog user-space interface, ioctls, and building a real keepalive daemon.

Next Lecture Browse Full Course

Leave a Reply

Your email address will not be published. Required fields are marked *