Cortex-M4 Architecture Deep Dive
Understand every block inside the Cortex-M4 processor, why ARM beats AVR and MSP430 for modern embedded design, and meet the real chip you will program throughout this course.
Why Manufacturers Choose ARM Cortex-M
Before diving into the block diagram, it helps to understand why ARM Cortex-M has become the default choice for new embedded designs. The reasons go beyond marketing — they are concrete engineering trade-offs that directly affect the products you build.
The Thumb-2 Instruction Set — Why It Matters
One of the most misunderstood aspects of Cortex-M programming is the Thumb-2 instruction set. Let’s demystify it completely.
Maximum power, maximum code size
Used in Cortex-A
Smaller code, less power
Limited capability
Best of both worlds
Cortex-M ONLY uses this
A critical point: Cortex-M processors execute only Thumb-2 instructions. There is no 32-bit ARM mode available. This is a deliberate design choice — the hardware is simpler, cheaper, and more power-efficient as a result.
A benchmark compiled for Cortex-M in Thumb-2 mode typically produces code 25–30% smaller than the same code compiled for classic 32-bit ARM mode. Smaller code means less flash memory needed, which directly reduces chip cost. On a production run of 10,000 units, moving from 256 KB flash to 128 KB flash can save thousands of dollars.
/* C source — compiler picks the best encoding automatically */
uint32_t add_and_shift(uint32_t a, uint32_t b)
{
return (a + b) << 3;
}
/* ARM GCC generates Thumb-2 assembly (arm-none-eabi-gcc -O2 -mthumb): */
/* ADD R0, R0, R1 ; 16-bit encoding — adds two registers */
/* LSL R0, R0, #3 ; 16-bit encoding — left shift by 3 */
/* BX LR ; 16-bit encoding — return from function */
/* Total: 6 bytes for this entire function. */
/* The same function in classic 32-bit ARM would take 12 bytes. */
ARM Cortex-M vs AVR vs MSP430 — Honest Comparison
Two major alternatives to ARM Cortex-M exist in the embedded market: the AVR architecture from Microchip (formerly Atmel) and the MSP430 from Texas Instruments. Both are excellent for specific use cases. Understanding their limitations helps you see exactly what ARM Cortex-M solves.
- 8-bit or 32-bit data bus (AVR8 vs AVR32)
- 16-bit address space on AVR8 (64 KB limit)
- No hardware floating point
- Simple interrupt system (fixed priority)
- Harvard architecture: separate data/code buses complicate pointers
- Extremely well documented for beginners
- Arduino boards use AVR — massive hobbyist ecosystem
- Simple, predictable cycle timing
- Typical: ATmega328P (16 MHz, 2 KB SRAM, 32 KB flash)
- Industry-leading ultra-low power consumption
- Active mode: as low as 100 µA/MHz
- LPM4 standby: 0.1 µA
- 16-bit architecture only — limited compute
- No floating point unit
- Limited peripheral selection vs STM32
- Smaller ecosystem, fewer third-party libraries
- More expensive per DMIPS than ARM Cortex-M
- Best for: energy-harvesting sensors, always-on monitoring
- Full 32-bit — 4 GB address space
- Hardware multiply in single cycle
- Hardware divide (no software loop needed)
- Optional single-precision FPU (M4)
- NVIC: up to 240 interrupts, 256 priorities
- MPU for memory protection and RTOS support
- Unified von Neumann-style memory map (simple pointer model)
- Low power comparable to MSP430 in stop modes
- Massive ecosystem: FreeRTOS, Zephyr, HAL, countless libraries
- Same skills work across ST, NXP, Nordic, TI chips
AVR is still fine for simple hobbyist projects and rapid Arduino prototyping. MSP430 remains competitive in applications that literally run for 10+ years on a coin cell and don’t need significant computation. But for any serious embedded product — IoT devices, wearables, industrial control, audio processing — ARM Cortex-M is the professional choice.
Inside the Cortex-M4: Full Block Diagram
Now let’s open up the Cortex-M4 and understand every block inside it. This diagram matches the official ARM Cortex-M4 Technical Reference Manual (TRM). Every block shown here is present in chips like the STM32F407 and STM32F411.
Register bank: R0–R15
Thumb-2 ISA (ARMv7E-M)
Hardware multiply & divide
SysTick 24-bit timer
IEEE 754 compliant
32 FP registers (S0–S31)
FMAC instruction
Up to 240 IRQs
8–256 priority levels
Tail-chaining
Up to 8 regions
Privilege-level enforcement
Detects interrupts
while core is powered off
Records every instruction
executed for replay
8 hardware breakpoints
Flash address remapping
4 comparators
Cycle counter (CYCCNT)
Profiling counters
printf-style debug
over SWO pin — no UART needed
Allows debugger to read
any memory/register
Serialises trace data
to external capture device
or JTAG Debug Port
2-wire SWD or 4-wire JTAG
Every Block Explained — What It Does and Why It Exists
1. The Cortex-M4 CPU Core
The core is a 3-stage pipeline processor: it simultaneously fetches the next instruction from flash, decodes the current instruction, and executes the previous one. This overlap means most simple instructions complete in a single clock cycle at full speed.
The register bank contains 16 registers (R0–R15), all 32 bits wide. Most computations use R0–R12 as general-purpose scratch registers. R13 is the Stack Pointer, R14 is the Link Register (stores return addresses), and R15 is the Program Counter.
A hardware multiply-accumulate unit (MAC) performs a 32×32-bit multiply in a single clock cycle, and a hardware divide unit completes in 2–12 cycles. On 8-bit AVR, a 32-bit multiply requires dozens of instructions. This difference is enormous for signal processing and control loops.
2. Floating Point Unit (FPU)
The FPU is optional — the Cortex-M3 does not have one; the Cortex-M4 typically includes it. It implements the IEEE 754 single-precision standard (32-bit floats) and adds 32 dedicated floating-point registers (S0–S31, which can also be viewed as 16 double-word registers D0–D15 for loading/storing 64-bit values — but arithmetic is single-precision only).
Without a hardware FPU, floating-point operations are emulated in software — a 32-bit float add might take 50–100 cycles via a software library. With the hardware FPU, the same operation takes 1 cycle. For motor control, PID loops, and audio processing, this difference determines whether a design is feasible at a given clock speed.
/* The FPU is disabled by default after reset to save power.
Enable it by setting CP10 and CP11 bits in the CPACR register.
Must be done before any floating-point instruction executes. */
#define FPU_CPACR (*((volatile uint32_t *)0xE000ED88U))
void fpu_enable(void)
{
/* Set bits 20:23 to 0b1111 — full access to CP10 and CP11 */
FPU_CPACR |= (0xFU << 20);
/* Memory barriers to ensure the setting takes effect before
the first FP instruction. Required by the Cortex-M4 TRM. */
__asm volatile ("dsb");
__asm volatile ("isb");
}
int main(void)
{
fpu_enable();
float result = 3.14159f * 2.0f; /* Now runs in 1 cycle on the FPU */
return (int)result;
}
STM32CubeIDE and most ARM GCC toolchain configurations enable the FPU automatically when you compile with -mfpu=fpv4-sp-d16 -mfloat-abi=hard. The startup code included by the IDE calls the FPU enable sequence before main(). When programming bare-metal from scratch, you must enable it yourself, as shown above.
3. NVIC — Nested Vectored Interrupt Controller
The NVIC is arguably the most important peripheral inside the Cortex-M processor. It manages the entire interrupt system and is tightly coupled to the CPU core — so tightly that interrupt response takes only 12 clock cycles from interrupt assertion to the first instruction of the ISR executing.
Key features that make the NVIC far superior to legacy 8-bit interrupt controllers:
- Up to 240 external interrupt lines — chip vendors can implement 1 to 240 of these. STM32F411 uses 62 maskable channels plus the 16 Cortex-M4 internal exceptions.
- 16 priority levels (on STM32F411) — lower number = higher priority. A priority-0 interrupt can preempt a priority-15 ISR that is already running. This is called preemption.
- Tail-chaining — if two interrupts are pending, the NVIC services the second one immediately after the first without re-saving/restoring CPU state. This saves 12 cycles per chained interrupt.
- Automatic state save/restore — the CPU automatically pushes R0–R3, R12, LR, PC, and xPSR to the stack on interrupt entry. You do not write any save/restore code in your ISR.
- Late-arrival handling — if a higher-priority interrupt arrives while the CPU is still in the pre-emption save phase, the NVIC switches to serve the higher-priority one first.
/* NVIC register base address (from ARM Cortex-M4 TRM) */
#define NVIC_ISER0 (*((volatile uint32_t *)0xE000E100U))
#define NVIC_IPR_BASE ((volatile uint32_t *)0xE000E400U)
/* STM32F411: IRQ number for USART1 = 37 (from Reference Manual RM0383) */
#define USART1_IRQn 37
void nvic_enable_irq(uint32_t irq_num, uint8_t priority)
{
/* Step 1: Set the priority (8 priority bits, upper 4 are implemented) */
/* Priority register: each IRQ gets 8 bits, 4 IRQs per 32-bit register */
uint32_t reg_index = irq_num / 4;
uint32_t bit_shift = (irq_num % 4) * 8;
NVIC_IPR_BASE[reg_index] |= ((uint32_t)(priority << 4) << bit_shift);
/* Step 2: Enable the interrupt in ISER (Interrupt Set Enable Register) */
/* Each ISER register handles 32 IRQs; ISER0 covers IRQ 0–31, ISER1: 32–63 */
volatile uint32_t *iser = (volatile uint32_t *)0xE000E100U;
iser[irq_num / 32] = (1U << (irq_num % 32));
}
int main(void)
{
nvic_enable_irq(USART1_IRQn, 5); /* Enable USART1 IRQ at priority 5 */
/* Enable global interrupts (clear PRIMASK) */
__asm volatile ("cpsie i");
while (1) { /* application loop */ }
}
/* USART1 ISR — name must match the vector table entry */
void USART1_IRQHandler(void)
{
/* Read received byte and echo it */
volatile uint32_t *USART1_DR = (volatile uint32_t *)0x40011004U;
volatile uint32_t *USART1_SR = (volatile uint32_t *)0x40011000U;
if (*USART1_SR & (1U << 5)) { /* RXNE flag set? */
uint8_t byte = (uint8_t)(*USART1_DR & 0xFF);
*USART1_DR = byte; /* Echo back */
}
}
4. MPU — Memory Protection Unit
The MPU lets you divide the 4 GB address space into up to 8 protected regions. Each region has configurable access rules: read-only, read-write, execute-never, privileged-only, etc. If code tries to access a region in a way that violates the rules, the CPU triggers a MemManage fault — the fault handler can then reset the offending task, log an error, or halt the system.
This is essential for RTOS-based systems: the kernel runs in privileged mode and configures the MPU so that each user task can only access its own memory. One buggy task cannot corrupt another task’s data or the kernel’s data. The STM32F411 datasheet confirms the MPU covers up to 8 subregions per region, with sizes between 32 bytes and the full 4 GB.
Imagine a medical device running two FreeRTOS tasks: one reads sensor data, one drives a display. Without MPU, a bug in the display task (e.g., an out-of-bounds array write) could silently corrupt the sensor data buffer — a dangerous situation in a medical context. With the MPU enabled, any illegal access immediately triggers a fault, the RTOS catches it, logs the fault address, and restarts only the faulty task. This is how safety-certified firmware is built.
5. DWT — Data Watchpoint and Trace
The DWT has two main uses that embedded developers rely on heavily:
- Hardware watchpoints — set a data watchpoint on a memory address. The CPU halts automatically when that address is read or written. Invaluable for catching memory corruption bugs.
- Cycle counter (CYCCNT) — a 32-bit counter that increments every CPU clock cycle. Use it to measure exactly how long a function takes in cycles, without a logic analyser or oscilloscope. Accurate to within ±1 cycle.
#define DWT_CTRL (*((volatile uint32_t *)0xE0001000U))
#define DWT_CYCCNT (*((volatile uint32_t *)0xE0001004U))
#define CoreDebug_DEMCR (*((volatile uint32_t *)0xE000EDFC U))
void dwt_init(void)
{
CoreDebug_DEMCR |= (1U << 24); /* Enable DWT */
DWT_CYCCNT = 0; /* Reset counter */
DWT_CTRL |= (1U << 0); /* Start counting */
}
uint32_t dwt_get_cycles(void)
{
return DWT_CYCCNT;
}
/* Example: time a sorting function */
void measure_sort(int *arr, int len)
{
uint32_t start = dwt_get_cycles();
bubble_sort(arr, len); /* your function here */
uint32_t end = dwt_get_cycles();
uint32_t cycles_taken = end - start;
/* At 100 MHz: cycles_taken / 100 = microseconds */
}
6. ITM — Instrumentation Trace Macrocell
The ITM is a hardware printf debug channel. Instead of routing debug messages through a UART (which requires wiring and affects timing), you write to ITM stimulus registers. The messages flow out through the Serial Wire Output (SWO) pin — the same wire used for SWD debugging — and appear in your IDE’s debug console.
#include <stdint.h>
#define ITM_STIM0 (*((volatile uint32_t *)0xE0000000U))
#define ITM_TER (*((volatile uint32_t *)0xE0000E00U))
/* Send one character via ITM stimulus port 0 */
void itm_putchar(char c)
{
/* Wait until port 0 is ready (bit 0 of STIM0) */
while ((ITM_STIM0 & 1U) == 0);
/* Write the byte — this queues it to the SWO output */
*((volatile uint8_t *)0xE0000000U) = (uint8_t)c;
}
void itm_print(const char *s)
{
while (*s) itm_putchar(*s++);
}
int main(void)
{
/* Enable ITM stimulus port 0 */
ITM_TER |= (1U << 0);
itm_print("Hello from ITM!\n");
while (1) {}
}
7. ETM — Embedded Trace Macrocell
The ETM records every single instruction the CPU executes and streams the trace data out through dedicated trace pins (TRACED[3:0]). An external trace capture device (like a SEGGER J-Trace or a Lauterbach) stores this stream, allowing you to replay the exact execution history after a bug occurs — even in real-time with no CPU slowdown.
ETM is the most powerful debugging tool available. It can show you exactly which code path led to a crash, without needing to reproduce the bug or halt the CPU. It is used in avionics, automotive, and industrial applications where bugs must be analysed in production.
8. FPB — Flash Patch Breakpoint
The FPB provides 8 hardware breakpoints — address comparators that halt the CPU when the program counter reaches a specific address. Unlike software breakpoints (which replace instructions with a BKPT opcode in flash), hardware breakpoints work without modifying flash memory and can be set in read-only regions. The FPB also has a patching capability to remap flash addresses to SRAM, useful for fixing bugs in field-deployed firmware without full re-flashing.
9. WIC — Wake-up Interrupt Controller
When the Cortex-M4 enters deep sleep mode, the CPU core itself is powered off. The WIC is a tiny, ultra-low-power circuit that stays powered and watches for interrupt signals. When an external interrupt fires, the WIC powers on the CPU, and the NVIC processes the interrupt normally. This allows the dynamic efficiency features of STM32 chips — the CPU sleeps between events, yet responds in microseconds.
The Three AHB-Lite Bus Interfaces
The Cortex-M4 exposes three separate bus interfaces to the outside world. Each has a specific purpose that optimises performance:
Connects to Flash memory
Read-only, 32-bit
Used for fetching instructions
Connects to Flash memory
Read-only, 32-bit
Used for const data in flash
Connects to SRAM & peripherals
Read/write, 32-bit
All peripheral register accesses
512 KB
128 KB
Peripherals
(DMA, GPIO)
Peripherals
(USB OTG)
Low-speed
peripherals
High-speed
peripherals
From Core to Chip: Meet the STM32F411
Now let’s see how all the Cortex-M4 blocks above appear in a real production chip. The STM32F411xC/xE from STMicroelectronics is an excellent learning target — it is one of the most popular Cortex-M4 chips on development boards and in student labs worldwide.
| Parameter | STM32F411xC | STM32F411xE |
|---|---|---|
| CPU Core | ARM Cortex-M4 with FPU (single-precision) | |
| Maximum CPU Frequency | 100 MHz | |
| Performance | 125 DMIPS (Dhrystone 2.1 benchmark) | |
| Flash Memory | 256 KB | 512 KB |
| SRAM | 128 KB (zero wait state at CPU clock speed) | |
| Supply Voltage | 1.7 V to 3.6 V | |
| Temperature Range | −40 °C to +125 °C (junction) | |
| General Purpose Timers | 7 (six 16-bit + two 32-bit) | |
| Advanced-control Timer | 1 (TIM1 — PWM with dead-time for motor control) | |
| SysTick Timer | 24-bit countdown — used by RTOS scheduler | |
| I2C Interfaces | 3 (SMBus/PMBus compatible) | |
| USART | 3 (up to 12.5 Mbit/s) | |
| SPI / I2S | 5 SPI / 5 I2S (2 full-duplex I2S) | |
| USB | USB 2.0 OTG Full Speed (12 Mbit/s) with on-chip PHY | |
| SDIO | 1 (SD card / MMC / eMMC) | |
| ADC | 1 × 12-bit, up to 2.4 MSPS, 10 or 16 channels | |
| DMA | 2 controllers × 8 streams = 16 DMA streams | |
| GPIO Pins | 36 / 50 / 81 (package dependent) | 36 / 50 / 81 |
| Interrupt Channels | 62 maskable + 16 Cortex-M4 exceptions | |
| Debug Interface | SWD (2-wire) and JTAG (4-wire) | |
| ART Accelerator | Yes — eliminates flash wait states at 100 MHz | |
| CRC Unit | Hardware CRC-32 calculation | |
| Unique ID | 96-bit factory-programmed unique device ID | |
| Power — Stop mode | 42 µA typical at 25°C (fast wakeup) | |
| Power — Standby mode | 1.8 µA at 25°C (without RTC) | |
The ART Accelerator — Zero-Wait Flash at 100 MHz
Flash memory is physically slower than a 100 MHz CPU clock. Without acceleration, the CPU would need to insert wait states (idle cycles) every time it fetches an instruction from flash — significantly cutting effective throughput.
STMicroelectronics solves this with their proprietary ART Accelerator (Adaptive Real-Time memory Accelerator). It adds a prefetch queue and branch cache between the CPU and flash. The CPU’s ICode bus requests instructions, and the ART Accelerator serves them from its cache with zero wait states — as if the flash were as fast as SRAM.
Requests next
instruction every
10 ns @ 100 MHz
Prefetch queue:
reads ahead in flash
Branch cache:
stores hot code paths
Physical access time:
~30–60 ns
(needs 3–6 wait states
at 100 MHz without ART)
DMA — Direct Memory Access Controller
The STM32F411 has two independent DMA controllers (DMA1 and DMA2), each with 8 streams, giving 16 DMA streams total. A DMA stream can transfer data between any combination of memory and peripheral without CPU involvement.
Example: you configure DMA to receive 256 bytes from a USART and write them directly to a buffer in SRAM. The CPU is free to execute other code. When the DMA completes, it fires an interrupt. The CPU wakes, processes the buffer, and goes back to other work. This is fundamentally more efficient than polling or interrupt-per-byte approaches.
#include <stdint.h>
/* DMA2 Stream0 register base (from STM32F411 Reference Manual RM0383) */
#define DMA2_BASE 0x40026400U
#define DMA2_S0CR (*((volatile uint32_t *)(DMA2_BASE + 0x10U)))
#define DMA2_S0NDTR (*((volatile uint32_t *)(DMA2_BASE + 0x14U)))
#define DMA2_S0PAR (*((volatile uint32_t *)(DMA2_BASE + 0x18U)))
#define DMA2_S0M0AR (*((volatile uint32_t *)(DMA2_BASE + 0x1CU)))
/* RCC: Enable DMA2 clock */
#define RCC_AHB1ENR (*((volatile uint32_t *)0x40023830U))
#define RCC_DMA2EN (1U << 22)
static uint32_t src_buf[64] = { [0 ... 63] = 0xDEADBEEFU };
static uint32_t dst_buf[64];
void dma2_memcpy(void)
{
/* Enable DMA2 clock */
RCC_AHB1ENR |= RCC_DMA2EN;
/* Configure DMA2 Stream0 for memory-to-memory */
DMA2_S0CR = 0; /* Disable stream first */
DMA2_S0PAR = (uint32_t)src_buf; /* Source: src_buf */
DMA2_S0M0AR = (uint32_t)dst_buf; /* Destination: dst_buf */
DMA2_S0NDTR = 64; /* 64 items to transfer */
/* Direction: memory-to-memory, increment both, 32-bit words, channel 0 */
DMA2_S0CR = (1U << 14) | /* MEM2MEM bit: memory-to-memory mode */
(1U << 10) | /* MINC: memory pointer increment */
(1U << 9) | /* PINC: peripheral pointer increment */
(2U << 11) | /* MSIZE: 32-bit memory data width */
(2U << 13); /* PSIZE: 32-bit peripheral data width*/
DMA2_S0CR |= (1U << 0); /* EN: Enable the stream — transfer starts */
/* Wait for transfer complete (TC flag in DMA2_LISR, bit 5) */
volatile uint32_t *DMA2_LISR = (volatile uint32_t *)0x40026400U;
while (!(*DMA2_LISR & (1U << 5)));
}
Understanding NVIC Priority Groups
One of the most important — and most confused — topics in Cortex-M programming is interrupt priority grouping. Let’s make it crystal clear.
The STM32F411 implements 4 priority bits per interrupt (the upper 4 of the 8-bit priority field). These 4 bits are split into two fields by the PRIGROUP setting in the Application Interrupt and Reset Control Register (AIRCR):
| PRIGROUP | Preemption Priority Bits | Sub-Priority Bits | Preemption Levels | Sub-levels per Preemption |
|---|---|---|---|---|
| 0b011 (default) | 4 bits | 0 bits | 16 | 1 |
| 0b100 | 3 bits | 1 bit | 8 | 2 |
| 0b101 | 2 bits | 2 bits | 4 | 4 |
| 0b110 | 1 bit | 3 bits | 2 | 8 |
| 0b111 | 0 bits | 4 bits | 1 | 16 |
Preemption priority determines if one ISR can interrupt another that is currently running (higher preemption priority wins). Sub-priority only matters when two interrupts with the same preemption priority are pending simultaneously — the one with the lower sub-priority number fires first. Sub-priority never causes preemption.
On STM32 chips with 4 implemented priority bits, the valid priority values are 0, 16, 32, 48 … 240 (multiples of 16) when writing raw to the 8-bit IPR register. Many beginners write values 0–15 directly and wonder why priority 5 and priority 7 behave identically — the lower 4 bits are ignored. Always shift your priority value left by 4 bits, or use the STM32 HAL’s NVIC_SetPriority() which handles this for you.
Setting Up the Development Environment
Before writing firmware, you need a toolchain on your PC. This course uses STM32CubeIDE — a free, Eclipse-based IDE developed by STMicroelectronics. It integrates:
- Editor — syntax highlighting, code completion, navigation
- Compiler — ARM GCC (arm-none-eabi-gcc) bundled inside
- Linker — produces the .elf and .bin files loaded to the chip
- Flash programmer — programs the binary over ST-LINK (USB)
- Debugger — GDB + ST-LINK for breakpoints, step-through, register view
- SWV / ITM viewer — displays ITM printf output and DWT profiling
If you prefer working without an IDE, the complete open-source toolchain is:
# Install ARM GCC compiler
sudo apt install gcc-arm-none-eabi binutils-arm-none-eabi
# Install OpenOCD (flash programmer and GDB server)
sudo apt install openocd
# Install udev rules for ST-LINK
sudo apt install stlink-tools
# Verify installation
arm-none-eabi-gcc --version
openocd --version
# Compile (assuming main.c and startup.s are present)
arm-none-eabi-gcc -mcpu=cortex-m4 -mthumb -mfpu=fpv4-sp-d16 \
-mfloat-abi=hard -O2 -Wl,-T,linker.ld \
startup.s main.c -o firmware.elf
# Convert to binary
arm-none-eabi-objcopy -O binary firmware.elf firmware.bin
# Flash via ST-LINK using OpenOCD
openocd -f interface/stlink.cfg \
-f target/stm32f4x.cfg \
-c "program firmware.elf verify reset exit"
The Hardware Target: STM32F407 Discovery Board
This course uses the STM32F407G-DISC1 (STM32 Discovery) board as the primary hardware target. It is also compatible with the STM32F411 Nucleo boards since both use Cortex-M4 cores with the same programmer’s model.
| Board Feature | STM32F407G-DISC1 Details | Why It Matters for Learning |
|---|---|---|
| Microcontroller | STM32F407VGT6 (Cortex-M4, 168 MHz, 1 MB flash, 192 KB SRAM) | A step up from F411 — same M4 core, more flash/SRAM for larger projects |
| Built-in Programmer | ST-LINK/V2 (mini USB connector, SWD interface) | No external programmer needed — plug USB and flash immediately |
| User LEDs | 4 LEDs: LD3 (orange, PD13), LD4 (green, PD12), LD5 (red, PD14), LD6 (blue, PD15) | 4 GPIOs to toggle for LED experiments without extra wiring |
| User Button | B1 button connected to PA0 | Interrupt / polling input for button experiments |
| MEMS accelerometer | LIS302DL or LIS3DSH connected via SPI | Real sensor to practice SPI communication |
| USB OTG | Full speed + high speed (with external PHY connector) | USB device/host experiments without extra hardware |
| Audio DAC | CS43L22 connected via I2S + I2C | Audio output experiments — play tones via I2S/I2C |
| Power Supply | USB-powered (5V → 3V regulator on board) | Single USB cable powers everything |
A First Look at the STM32 Clock Tree
One of the most important things to understand about STM32 chips — and one that trips up almost every beginner — is the clock tree. Peripherals on an STM32 are clocked by different bus frequencies, and each peripheral must have its clock enabled before you can use it.
Internal 16 MHz RC
Default after reset
±1% accuracy
External crystal
4–26 MHz
Better accuracy
Multiplies input
up to 100 MHz
Configured via RCC
100 MHz
System clock
feeds CPU, AHB
100 MHz
GPIO, DMA, Flash
100 MHz
USART1, SPI1, ADC
50 MHz
USART2, I2C, SPI2/3
General-purpose timers
/* Registers from STM32F411 Reference Manual (RM0383) */
#define RCC_CR (*((volatile uint32_t *)0x40023800U))
#define RCC_PLLCFGR (*((volatile uint32_t *)0x40023804U))
#define RCC_CFGR (*((volatile uint32_t *)0x40023808U))
#define FLASH_ACR (*((volatile uint32_t *)0x40023C00U))
void system_clock_100MHz(void)
{
/* Step 1: Enable ART Accelerator and set 3 flash wait states for 100 MHz */
FLASH_ACR = (1U << 9) | /* PRFTEN: prefetch enable */
(1U << 10)| /* ICEN: instruction cache */
(1U << 11)| /* DCEN: data cache */
3U; /* LATENCY: 3 wait states */
/* Step 2: Configure PLL: HSI (16 MHz) as source
PLLM=8 : divides 16 MHz to 2 MHz (VCO input)
PLLN=200: multiplies 2 MHz to 400 MHz (VCO output)
PLLP=4 : divides 400 MHz to 100 MHz (SYSCLK)
PLLQ=8 : divides 400 MHz to 50 MHz (USB — needs 48 MHz, not used here) */
RCC_PLLCFGR = (8U << 0) | /* PLLM = 8 */
(200U << 6) | /* PLLN = 200 */
(1U << 16) | /* PLLP = 4 (bits 17:16 = 01 means /4) */
(8U << 24) | /* PLLQ = 8 */
(0U << 22); /* PLLSRC = HSI */
/* Step 3: Enable PLL and wait for it to lock */
RCC_CR |= (1U << 24); /* PLLON = 1 */
while (!(RCC_CR & (1U << 25))); /* Wait for PLLRDY */
/* Step 4: Set AHB prescaler = 1, APB1 = /2 (50 MHz), APB2 = /1 (100 MHz) */
RCC_CFGR = (4U << 10) | /* PPRE1: APB1 /2 = 50 MHz */
(0U << 13); /* PPRE2: APB2 /1 = 100 MHz */
/* Step 5: Switch SYSCLK to PLL */
RCC_CFGR |= (2U << 0); /* SW = PLL */
while (((RCC_CFGR >> 2) & 3) != 2); /* Wait for SWS = PLL */
}
Frequently Asked Questions
Lecture 3 dives into the register bank in detail — R0 through R15, the two stack pointers (MSP and PSP), the Link Register, the Program Counter, and the Program Status Register (xPSR). Understanding these is the foundation for writing interrupt handlers and RTOS code.

2 Comments