Linux CPUFreq Governors Hands-On Guide-Free Linux Device Drivers Training Online

Linux CPUFreq Governors Hands-On Guide

Linux CPUFreq Governors Hands-On Guide

The hands-on companion to our CPU idle and CPUFreq lecture — real sysfs commands, hotplug practice, and an original CPUFreq notifier kernel module, part of this free Linux kernel development course.

This lecture is the hands-on half of our CPU idle and CPUFreq pair. In the previous lecture we covered the theory: C-states, idle governors, Operating Performance Points, and every CPUFreq governor from performance to schedutil. Here we actually touch the sysfs interfaces on a running Linux system, hotplug a CPU, flip governors, and — because reading documentation only gets you so far as a driver engineer — write, build, and load a small original kernel module called ep_freq_notifier that hooks the CPUFreq transition notifier chain and logs every frequency change to the kernel log. Every command below is shown with the actual sysfs path and a realistic expected output, exactly as you would see it on a typical x86_64 or ARM embedded Linux target.

What You Will Learn

Reading cpuidle state stats Disabling a C-state Hotplugging a CPU offline/online Listing CPUFreq governors Switching the active governor Setting a fixed frequency Writing a cpufreq notifier module Reading devfreq from sysfs

Prerequisites

Read the CPU Idle and CPUFreq explanation lecture first Basic C and kernel module build knowledge Root/sudo access on a Linux target or VM Kernel headers matching your running kernel

Exploring CPU Idle States From User Space

Start by identifying which CPUIdle driver and governor are active. These two files live directly under the shared cpuidle directory, not per-CPU:

$ cat /sys/devices/system/cpu/cpuidle/current_driver
intel_idle

$ cat /sys/devices/system/cpu/cpuidle/current_governor_ro
menu

$ cat /sys/devices/system/cpu/cpuidle/available_governors
menu ladder teo

Now look at CPU 0’s individual idle states. Each stateN directory is one C-state, numbered shallowest to deepest:

$ ls /sys/devices/system/cpu/cpu0/cpuidle/
state0  state1  state2  state3

$ for s in /sys/devices/system/cpu/cpu0/cpuidle/state*; do
    printf "%-8s latency=%-6s residency=%-6s usage=%-10s %s\n" \
      "$(basename $s)" "$(cat $s/latency)" "$(cat $s/residency)" \
      "$(cat $s/usage)" "$(cat $s/name)"
  done
state0   latency=0      residency=0      usage=8213       POLL
state1   latency=2      residency=2      usage=193442     C1
state2   latency=10     residency=20     usage=88213      C1E
state3   latency=133    residency=300    usage=402117     C6

To stop the governor from ever selecting the deepest state on this particular CPU (useful when chasing a real-time latency spike), write 1 to its disable attribute:

$ echo 1 | sudo tee /sys/devices/system/cpu/cpu0/cpuidle/state3/disable
1

$ cat /sys/devices/system/cpu/cpu0/cpuidle/state3/disable
1

Remember this is per-CPU — repeat it for every CPU you want affected, and re-enable with echo 0 once you are done debugging.

Hands-On CPU Hotplugging And Idle Statistics

Hotplug is controlled by the per-CPU online attribute. Taking a CPU offline removes it from scheduling entirely — a much bigger step than simply letting the idle governor put it in a deep C-state:

$ cat /sys/devices/system/cpu/cpu3/online
1

$ nproc
4

$ echo 0 | sudo tee /sys/devices/system/cpu/cpu3/online
0

$ nproc
3

$ cat /sys/devices/system/cpu/cpu3/online
0

$ cat /sys/devices/system/cpu/cpu3/cpuidle/state0/name
cat: /sys/devices/system/cpu/cpu3/cpuidle/state0/name: No such file or directory
# cpuidle attributes disappear while the CPU is offline

$ echo 1 | sudo tee /sys/devices/system/cpu/cpu3/online
1

$ nproc
4

Notice that once cpu3 comes back online, its cpuidle directory is rebuilt from scratch and any per-state disable flags you had set earlier are reset to their driver defaults — a common gotcha when scripting hotplug-based power policies.

Reading And Changing CPUFreq Governors From sysfs

Listing Available Governors And Frequencies

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_driver
acpi-cpufreq

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_available_governors
performance powersave schedutil

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
schedutil

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_available_frequencies
2400000 2200000 2000000 1800000 1600000 1400000 1200000 1000000 800000

$ cat /sys/devices/system/cpu/cpu0/cpufreq/cpuinfo_cur_freq
1601342

$ cat /sys/devices/system/cpu/cpu0/cpufreq/affected_cpus
0 1 2 3

Switching Governors And Setting A Fixed Frequency

$ echo performance | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
performance

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
2400000

$ echo userspace | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
userspace

$ echo 1200000 | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_setspeed
1200000

$ cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_cur_freq
1200000

# restore the sane default before moving on
$ echo schedutil | sudo tee /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor
schedutil

Because scaling_governor applies to the whole policy, this same write affects every CPU listed in affected_cpus for cpu0’s policy — on this system that is all four cores.

Building ep_freq_notifier: A CPUFreq Transition Logger

Reading sysfs tells you the current state; a kernel-side notifier tells you about every transition as it happens, which is exactly what a driver needs when it must react to a frequency change (adjusting a bus timing register, for example). The CPUFreq core exposes a standard kernel notifier chain for this purpose. ep_freq_notifier is a small, original module that registers a callback on the CPUFREQ_TRANSITION_NOTIFIER chain and logs the old and new frequency for every CPU whenever the active governor changes it.

Module Source

// ep_freq_notifier.c — minimal CPUFreq transition logger
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/init.h>
#include <linux/cpufreq.h>
#include <linux/notifier.h>

static int ep_freq_transition_callback(struct notifier_block *nb,
                                        unsigned long val, void *data)
{
        struct cpufreq_freqs *freqs = data;

        /* Log only after the transition has completed */
        if (val != CPUFREQ_POSTCHANGE)
                return NOTIFY_DONE;

        pr_info("ep_freq_notifier: cpu%u transitioned %u kHz -> %u kHz\n",
                freqs->cpu, freqs->old, freqs->new);

        return NOTIFY_OK;
}

static struct notifier_block ep_freq_nb = {
        .notifier_call = ep_freq_transition_callback,
};

static int __init ep_freq_notifier_init(void)
{
        int ret;

        ret = cpufreq_register_notifier(&ep_freq_nb, CPUFREQ_TRANSITION_NOTIFIER);
        if (ret) {
                pr_err("ep_freq_notifier: failed to register notifier (%d)\n", ret);
                return ret;
        }

        pr_info("ep_freq_notifier: loaded, watching CPUFreq transitions\n");
        return 0;
}

static void __exit ep_freq_notifier_exit(void)
{
        cpufreq_unregister_notifier(&ep_freq_nb, CPUFREQ_TRANSITION_NOTIFIER);
        pr_info("ep_freq_notifier: unloaded\n");
}

module_init(ep_freq_notifier_init);
module_exit(ep_freq_notifier_exit);

MODULE_LICENSE("GPL");
MODULE_AUTHOR("EmbeddedPathashala");
MODULE_DESCRIPTION("ep_freq_notifier - minimal CPUFreq transition logger");

The callback ignores the CPUFREQ_PRECHANGE phase and only logs on CPUFREQ_POSTCHANGE, so each real transition produces exactly one log line rather than two. freqs->cpu, freqs->old, and freqs->new come straight from the struct cpufreq_freqs the CPUFreq core passes as the notifier’s data argument — this same pattern is what real drivers (memory controllers, bus bridges, thermal-aware peripherals) use to stay in sync with CPU frequency changes.

Makefile And Build Steps

# Makefile
obj-m += ep_freq_notifier.o

KDIR := /lib/modules/$(shell uname -r)/build
PWD  := $(shell pwd)

all:
	$(MAKE) -C $(KDIR) M=$(PWD) modules

clean:
	$(MAKE) -C $(KDIR) M=$(PWD) clean
$ make
make -C /lib/modules/6.8.0-generic/build M=/home/dev/ep_freq_notifier modules
  CC [M]  /home/dev/ep_freq_notifier/ep_freq_notifier.o
  MODPOST /home/dev/ep_freq_notifier/Module.symvers
  CC [M]  /home/dev/ep_freq_notifier/ep_freq_notifier.mod.o
  LD [M]  /home/dev/ep_freq_notifier/ep_freq_notifier.ko

Loading The Module And Reading dmesg

$ sudo insmod ep_freq_notifier.ko

$ dmesg | tail -n 1
[ 1210.004112] ep_freq_notifier: loaded, watching CPUFreq transitions

# Now generate some load so schedutil actually changes frequency
$ stress-ng --cpu 2 --timeout 10s &

$ dmesg | tail -n 5
[ 1211.512331] ep_freq_notifier: cpu2 transitioned 800000 kHz -> 1800000 kHz
[ 1211.998812] ep_freq_notifier: cpu2 transitioned 1800000 kHz -> 2400000 kHz
[ 1214.221033] ep_freq_notifier: cpu3 transitioned 800000 kHz -> 2000000 kHz
[ 1221.774213] ep_freq_notifier: cpu2 transitioned 2400000 kHz -> 1000000 kHz
[ 1221.981204] ep_freq_notifier: cpu3 transitioned 2000000 kHz -> 800000 kHz

$ sudo rmmod ep_freq_notifier

$ dmesg | tail -n 1
[ 1230.331199] ep_freq_notifier: unloaded

If your platform’s scaling driver bypasses the notifier layer entirely (this is the case with intel_pstate in its default active mode), you will see the module load cleanly but never log a transition — switch intel_pstate to passive mode, or test on a driver that goes through the standard governor path, such as acpi-cpufreq or a typical ARM cpufreq-dt driver.

A Quick Look At devfreq From The Command Line

The same governor/driver split applies to non-CPU devices through devfreq. On an SoC with a scalable GPU, you would see something like this:

$ ls /sys/class/devfreq/
18000000.gpu

$ cat /sys/class/devfreq/18000000.gpu/governor
simple_ondemand

$ cat /sys/class/devfreq/18000000.gpu/available_frequencies
200000000 300000000 400000000 525000000 650000000

$ cat /sys/class/devfreq/18000000.gpu/cur_freq
400000000

$ echo performance | sudo tee /sys/class/devfreq/18000000.gpu/governor
performance

Common Mistakes And Troubleshooting

  • Writing to scaling_governor with the wrong permissions. Most distributions restrict this file to root; use sudo tee instead of a plain redirect, since sudo echo x > file only elevates echo, not the shell’s own redirection.
  • Expecting notifier callbacks under intel_pstate active mode. intel_pstate’s default active mode bypasses the generic governor layer, so CPUFREQ_TRANSITION_NOTIFIER never fires there — this trips up almost everyone the first time they test a notifier module on a modern Intel laptop.
  • Forgetting hotplugged CPUs reset their idle state configuration. Any per-CPU disable flags set under cpuidle are lost when that CPU is taken offline and back online; reapply them after hotplug in any automated script.
  • Loading a module built against the wrong kernel headers. insmod will fail with “Invalid module format” if KDIR pointed at headers that do not match uname -r exactly.
  • Not restoring the governor after testing userspace/performance. Leaving a policy pinned to performance after a benchmark run quietly burns power on the next boot-persistent config reload; always restore schedutil (or the platform default) when done.

Best Practices For Hands-On CPU Idle And CPUFreq Work

  • Always read the current governor and frequency before changing anything, so you can restore the original state afterward.
  • Test cpufreq notifier code against a driver that actually runs the generic governor path (acpi-cpufreq, cpufreq-dt), not intel_pstate active mode.
  • Unregister every notifier you register — an unbalanced cpufreq_register_notifier/cpufreq_unregister_notifier pair leaves a stale callback pointer active after your module unloads and will oops the kernel on the next transition.
  • Script hotplug and idle-state changes together; assume cpuidle attributes are transient across an offline/online cycle.
  • Prefer rate_limit_us tuning over switching away from schedutil when you just need to change how aggressively frequency reacts to bursts.

Summary And Key Takeaways

You now have hands-on command-line fluency with both halves of CPU idle and CPUFreq: reading and disabling individual C-states, hotplugging a CPU and observing how its cpuidle tree disappears and reappears, and reading, switching, and pinning CPUFreq governors and frequencies through sysfs. You also built ep_freq_notifier, an original kernel module that hooks the CPUFreq transition notifier chain with cpufreq_register_notifier and logs every real frequency change through pr_info — the same mechanism production drivers use to stay synchronized with CPU frequency scaling. Combined with the architectural explanation lecture, you now have both the theory and the practical sysfs/kernel-module fluency to reason about, debug, and tune CPU idle and CPUFreq behavior on real embedded Linux hardware.

Frequently Asked Questions About CPU Idle And CPUFreq

Why didn’t my ep_freq_notifier module log anything on my laptop?

Most Intel laptops default to intel_pstate in active mode, which bypasses the generic CPUFreq governor layer and the standard transition notifier chain entirely. Test on acpi-cpufreq, cpufreq-dt (common on ARM boards), or switch intel_pstate to passive mode.

Why does echo governor > scaling_governor fail even with sudo?

The redirection happens in your unprivileged shell before sudo ever runs, so only the echo command is elevated, not the write. Use echo value | sudo tee /path/to/file instead, which pipes the value into a root-owned tee process.

Do I need to re-disable a cpuidle state after hotplugging that CPU?

Yes. When a CPU goes offline its cpuidle sysfs directory is torn down, and when it comes back online the directory is rebuilt from the driver’s defaults, so any previously written disable flags are lost and must be reapplied.

Is it safe to leave the governor set to performance after testing?

It works, but it defeats CPUFreq’s entire purpose by pinning every core at maximum frequency and voltage regardless of load. Always restore schedutil, or your platform’s intended default, once testing or benchmarking is complete.

What happens if I forget to call cpufreq_unregister_notifier on module exit?

The CPUFreq core keeps a pointer to your notifier_block on its chain. Once your module is unloaded, that memory is gone, and the next frequency transition invokes a callback pointer into freed module memory, which reliably crashes the kernel.

Can userspace directly request a frequency that isn’t in scaling_available_frequencies?

No. Writing to scaling_setspeed under the userspace governor is validated against the frequency table the scaling driver reported; requesting an unsupported value is rejected or rounded to the nearest supported OPP, depending on the driver.

Enjoyed This Hands-On CPUFreq Lecture?

This is part of EmbeddedPathashala’s free Linux device drivers course. Continue to the next lecture, or catch up on the CPU Idle and CPUFreq theory if you landed here first.

Go To Next Lecture Back To Course Index

Leave a Reply

Your email address will not be published. Required fields are marked *