Every time a machine comes back up after an unexpected reset, one question matters more than any other: did the watchdog cause this? A power-supply brownout, an overheating SoC, and a hung kernel thread that finally missed its keep-alive all leave the board in the same physical state — rebooted — but they are very different problems to fix. This lecture is part of our free linux kernel development course and walks through exactly how a user-space program asks the kernel “why did we reboot?” using the watchdog subsystem’s status ioctls.
What You Will Learn
- The difference between
WDIOC_GETSTATUSandWDIOC_GETBOOTSTATUS - Why old miscellaneous-device watchdog drivers and new generic-framework drivers answer these ioctls differently
- How the watchdog core builds the status bitmask internally, conceptually, without copying any book source
- How to write a small diagnostic tool that decodes the boot status into a human-readable reason
- Common mistakes when interpreting status flags across driver generations
Prerequisites
You should already be comfortable with opening /dev/watchdogN, arming it, and pinging it with WDIOC_KEEPALIVE, along with reading capability flags via WDIOC_GETSUPPORT — all covered earlier in this watchdog series. Basic familiarity with ioctl() and bitmask flags is assumed.
Two Questions, Two ioctls
The watchdog uAPI exposes two related but distinct queries. WDIOC_GETSTATUS asks “what is the watchdog’s current internal status right now” — things like whether a keep-alive ping has arrived recently, or whether the magic-close sequence was seen. WDIOC_GETBOOTSTATUS asks a different, more historical question: “what caused the *last* reset the system went through.” Confusing the two is one of the most common mistakes when building board-bring-up diagnostics, because on some hardware they happen to report the same bits, which hides the bug until you move to different silicon.
| ioctl | Question it answers | Typical use |
|---|---|---|
WDIOC_GETSTATUS | What is happening right now? | Runtime health monitoring, is a ping overdue |
WDIOC_GETBOOTSTATUS | Why did we reboot last time? | Post-mortem / crash-reason logging at boot |
Old Drivers vs the Generic Watchdog Framework
Not every watchdog driver in the tree answers these ioctls the same way, and understanding why matters when you are debugging a board-support package that predates the generic framework.
- Legacy miscellaneous-device drivers implement their own
file_operationsand their own.unlocked_ioctlhandler from scratch. Some only bother to supportWDIOC_GETSTATUS, reporting the raw content of a hardware status register. A few also supportWDIOC_GETBOOTSTATUS, and whether that returns the same value asGETSTATUSor a genuinely different, pre-parsed value depends entirely on how that individual driver author wrote theswitchstatement in their ioctl handler. There is no guarantee either way — you have to read that specific driver’s source. - Generic-framework drivers (any driver built on
struct watchdog_deviceandstruct watchdog_ops) don’t implement ioctl handling themselves at all. The watchdog core’s shared character-device code handles both ioctls uniformly, so behavior is consistent across every driver that uses the framework — which is effectively every actively maintained driver upstream today.
How the Core Builds the Status Value
For a generic-framework driver, WDIOC_GETBOOTSTATUS is the simplest case: the core just returns the value already sitting in the driver’s watchdog_device.bootstatus field, which the driver itself populated at probe time — usually by reading a reset-cause register and translating it into one or more WDIOF_* flags such as WDIOF_CARDRESET, WDIOF_POWERUNDER, or WDIOF_OVERHEAT.
WDIOC_GETSTATUS is slightly more involved. If the driver supplies an optional .status callback in its watchdog_ops, the core calls straight into the driver and hands the returned value to user space unmodified — the driver is trusted to know its own hardware best. If no .status callback exists, the core falls back to synthesizing a value itself: it masks bootstatus down to only the flags that make sense as “current status” bits, then layers in two software-tracked bits that have nothing to do with hardware at all — a WDIOF_MAGICCLOSE bit if the device is currently configured to allow the magic-close disarm sequence, and a WDIOF_KEEPALIVEPING bit that is set (and atomically cleared) if a keep-alive ping has arrived since the last time anyone asked. That keep-alive bit is a one-shot “have you pinged since I last checked” signal, which is exactly why it gets cleared on read.
Reading the Status from User Space
Once you understand what the two ioctls mean, using them is a two-line affair:
int flags = 0;
if (ioctl(fd, WDIOC_GETBOOTSTATUS, &flags) == -1) {
perror("WDIOC_GETBOOTSTATUS");
exit(EXIT_FAILURE);
}
/* flags now holds a bitmask of WDIOF_* reasons for the last reset */
Swap in WDIOC_GETSTATUS with the same call shape to ask about current state instead of boot history. From here you decode flags exactly the same way you decoded capability flags from WDIOC_GETSUPPORT earlier in this course — bit by bit, against the WDIOF_* constants.
Original Demo: ep_wdt_bootcheck
Below is a small original tool — not adapted from any book — that a board bring-up engineer could genuinely drop into a boot script. It reads the boot status once, translates it into a readable sentence, and exits with a status code a systemd unit or init script can act on.
/* ep_wdt_bootcheck.c
* Reports the reason for the last system reset, as seen by the watchdog.
* Build: gcc -Wall -O2 -o ep_wdt_bootcheck ep_wdt_bootcheck.c
* Run: ./ep_wdt_bootcheck /dev/watchdog0
*/
#include <stdio.h>
#include <stdlib.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <linux/watchdog.h>
struct ep_reason {
unsigned int flag;
const char *text;
};
static const struct ep_reason ep_reasons[] = {
{ WDIOF_OVERHEAT, "device overheated" },
{ WDIOF_FANFAULT, "cooling fan failed" },
{ WDIOF_EXTERN1, "external reset line 1 fired" },
{ WDIOF_EXTERN2, "external reset line 2 fired" },
{ WDIOF_POWERUNDER, "power supply under-voltage" },
{ WDIOF_CARDRESET, "watchdog card issued reset" },
{ WDIOF_POWEROVER, "power supply over-voltage" },
};
int main(int argc, char *argv[])
{
const char *dev = (argc > 1) ? argv[1] : "/dev/watchdog0";
int fd = open(dev, O_RDWR);
int flags = 0;
size_t i;
int matched = 0;
if (fd < 0) {
perror("open");
return EXIT_FAILURE;
}
if (ioctl(fd, WDIOC_GETBOOTSTATUS, &flags) == -1) {
perror("WDIOC_GETBOOTSTATUS");
close(fd);
return EXIT_FAILURE;
}
printf("ep_wdt_bootcheck: boot status raw = 0x%08x\n", flags);
for (i = 0; i < sizeof(ep_reasons) / sizeof(ep_reasons[0]); i++) {
if (flags & ep_reasons[i].flag) {
printf(" -> %s\n", ep_reasons[i].text);
matched = 1;
}
}
if (!matched)
printf(" -> normal power-on / no watchdog-related reset recorded\n");
close(fd);
return EXIT_SUCCESS;
}
$ sudo ./ep_wdt_bootcheck /dev/watchdog0
ep_wdt_bootcheck: boot status raw = 0x00000004
-> external reset line 1 fired
Boot Status Decision Flow
Common Mistakes and Troubleshooting
- Treating GETSTATUS and GETBOOTSTATUS as interchangeable — they answer different questions; only test on real hardware where a driver author has clearly differentiated them, don’t assume from one SoC’s behavior.
- Ignoring the .status callback vs core-synthesized path — a driver with its own
.statusop may report bits the generic fallback never would; read the specific driver if numbers look unexpected. - Forgetting that WDIOF_KEEPALIVEPING is one-shot — calling
WDIOC_GETSTATUStwice in a row can legitimately show the ping bit clear on the second call even though the daemon is healthy, because reading it clears it. - Not checking the ioctl return value — a driver may not implement
WDIOC_GETBOOTSTATUSat all and return-ENOTTY; always check for errors before trustingflags.
Best Practices
- Log boot status early in your init sequence, before anything else can obscure the cause of the previous reset.
- Prefer
WDIOC_GETBOOTSTATUSfor post-mortem logging; reserveWDIOC_GETSTATUSfor runtime health checks in a monitoring daemon. - Persist the decoded reason to non-volatile storage or a remote log server if the device is headless — the local kernel log itself may not survive a hard power cycle.
Summary / Key Takeaways
WDIOC_GETSTATUSreports current watchdog state;WDIOC_GETBOOTSTATUSreports why the last reset happened.- Legacy drivers implement these ioctls independently and inconsistently; generic-framework drivers get uniform behavior from the watchdog core.
- The core can either delegate to a driver’s
.statuscallback or synthesize a value frombootstatusplus two software-tracked bits. - A tiny boot-time tool like
ep_wdt_bootcheckturns a raw bitmask into an actionable, human-readable reboot reason.
Conclusion
Boot status reporting is one of the most practically useful — and most overlooked — parts of the Linux watchdog subsystem. Once your board can reliably answer “why did we just reboot,” a huge category of field-debugging guesswork disappears. In the next lecture of this free linux device drivers course, we move from ioctls to the watchdog’s sysfs interface, which lets you inspect and manage a watchdog device without writing any C code at all.
Frequently Asked Questions
What is the difference between WDIOC_GETSTATUS and WDIOC_GETBOOTSTATUS?
GETSTATUS reports the watchdog’s current internal state (such as whether a keep-alive ping was recently received), while GETBOOTSTATUS reports the reason the system’s last reset occurred, as recorded by the watchdog hardware or driver.
Do all watchdog drivers support both ioctls?
No. Legacy miscellaneous-device drivers may support only one or the other, or return identical values for both. Drivers built on the generic watchdog framework get consistent, core-handled behavior for both ioctls.
Why does WDIOF_KEEPALIVEPING sometimes disappear between reads?
It is a one-shot flag: the core atomically clears the internal keep-alive bit as soon as it is read, so a second immediate read can legitimately show it unset even if the watchdog is healthy.
Can WDIOC_GETBOOTSTATUS fail?
Yes, if the underlying driver doesn’t populate a meaningful bootstatus value the ioctl can return 0, and on drivers that never wire up boot status reporting at all, the call can fail with -ENOTTY.
Is boot status persisted across multiple reboots?
Generally no — it typically reflects only the most recent reset, since the value is usually derived from a hardware reset-cause register that gets overwritten on the next reset.
Where should I call WDIOC_GETBOOTSTATUS in my init flow?
As early as possible, ideally in a dedicated boot-diagnostics step before application services start, so the reason for the previous reset is logged before anything else can interfere.
What kernel config option is needed for these ioctls?
The watchdog character-device interface itself, not a separate status-specific option; any driver registered through the generic watchdog framework automatically exposes both ioctls via the shared core.
Keep Learning the Watchdog Subsystem
Next up: managing watchdog devices entirely from user space through sysfs — no C code required.
Next Lecture Browse the Full Course