Architecture & Why ARM Wins-Embedded C Training Institute in Hyderabad

Cortex-M4 Architecture Deep Dive

Lecture 02 — Architecture & Why ARM Wins

Cortex-M4 Architecture Deep Dive

Understand every block inside the Cortex-M4 processor, why ARM beats AVR and MSP430 for modern embedded design, and meet the real chip you will program throughout this course.

Beginner → Intermediate Estimated read: 30 min Reference: STM32F411 Datasheet (DocID026289)

Why Manufacturers Choose ARM Cortex-M

Before diving into the block diagram, it helps to understand why ARM Cortex-M has become the default choice for new embedded designs. The reasons go beyond marketing — they are concrete engineering trade-offs that directly affect the products you build.

💰
32-bit at 8-bit Price
Cortex-M0 and M3 chips cost the same or less than legacy 8-bit microcontrollers while delivering far more processing power. You get a full 32-bit address space, hardware multiply, and deterministic interrupt latency for pennies.
⚡
High Performance
The STM32F411 Cortex-M4 hits 125 DMIPS at 100 MHz using the ART Accelerator. That is roughly 50x the raw throughput of a classic 16 MHz 8-bit AVR running the same DMIPS benchmark — at comparable cost.
🔋
Ultra Low Power
The STM32F411 in Stop mode draws just 42 µA at 25°C. In Standby mode that drops to 1.8 µA. This makes ARM Cortex-M chips competitive with dedicated ultra-low-power parts for battery-powered designs.
💾
Code Density (Thumb-2)
Thumb-2 instructions are a mix of 16-bit compact encodings and 32-bit powerful encodings. The result is code that is 25–30% smaller than pure 32-bit ARM code, reducing flash requirements and cost.
🔴
Powerful Interrupt System
The NVIC handles up to 240 external interrupt sources (chip-vendor configurable) with 8–256 priority levels. It automatically saves CPU state on interrupt entry and restores it on exit — in hardware, with zero overhead code.
🔨
Customisable Silicon
Chip vendors can include or omit the FPU, DSP extensions, MPU, ETM trace, and cache. One ARM architecture license covers everything from a simple sensor node (M0, no FPU) to a high-performance controller (M7, double-precision FPU, cache).
🏭
RTOS Friendly
The Cortex-M includes hardware support for RTOS: two stack pointers (MSP and PSP), privileged/unprivileged execution modes, and the MPU for task memory isolation. FreeRTOS and Zephyr are designed around these features.
📚
Excellent Documentation
ARM publishes the full Architecture Reference Manual, Technical Reference Manual, and Programming Manual as free downloads. Every chip vendor publishes a detailed Reference Manual. No other embedded architecture has this level of public documentation.

The Thumb-2 Instruction Set — Why It Matters

One of the most misunderstood aspects of Cortex-M programming is the Thumb-2 instruction set. Let’s demystify it completely.

Evolution of ARM Instruction Sets
ARM (32-bit) All instructions 32 bits wide
Maximum power, maximum code size
Used in Cortex-A
→
Thumb (16-bit) All instructions 16 bits wide
Smaller code, less power
Limited capability
→
Thumb-2 (Mixed) 16-bit AND 32-bit instructions
Best of both worlds
Cortex-M ONLY uses this
Thumb-2 eliminates the mode-switch penalty of classic ARM/Thumb interworking. The CPU fetches 16-bit or 32-bit words from flash seamlessly in a single stream.

A critical point: Cortex-M processors execute only Thumb-2 instructions. There is no 32-bit ARM mode available. This is a deliberate design choice — the hardware is simpler, cheaper, and more power-efficient as a result.

Practical Code Size Impact

A benchmark compiled for Cortex-M in Thumb-2 mode typically produces code 25–30% smaller than the same code compiled for classic 32-bit ARM mode. Smaller code means less flash memory needed, which directly reduces chip cost. On a production run of 10,000 units, moving from 256 KB flash to 128 KB flash can save thousands of dollars.

Example: 16-bit vs 32-bit Thumb-2 instructions in the same C function
/* C source — compiler picks the best encoding automatically */
uint32_t add_and_shift(uint32_t a, uint32_t b)
{
    return (a + b) << 3;
}

/* ARM GCC generates Thumb-2 assembly (arm-none-eabi-gcc -O2 -mthumb): */
/* ADD  R0, R0, R1     ; 16-bit encoding — adds two registers            */
/* LSL  R0, R0, #3    ; 16-bit encoding — left shift by 3                */
/* BX   LR            ; 16-bit encoding — return from function           */

/* Total: 6 bytes for this entire function.                               */
/* The same function in classic 32-bit ARM would take 12 bytes.           */

ARM Cortex-M vs AVR vs MSP430 — Honest Comparison

Two major alternatives to ARM Cortex-M exist in the embedded market: the AVR architecture from Microchip (formerly Atmel) and the MSP430 from Texas Instruments. Both are excellent for specific use cases. Understanding their limitations helps you see exactly what ARM Cortex-M solves.

AVR (Microchip/Atmel)
  • 8-bit or 32-bit data bus (AVR8 vs AVR32)
  • 16-bit address space on AVR8 (64 KB limit)
  • No hardware floating point
  • Simple interrupt system (fixed priority)
  • Harvard architecture: separate data/code buses complicate pointers
  • Extremely well documented for beginners
  • Arduino boards use AVR — massive hobbyist ecosystem
  • Simple, predictable cycle timing
  • Typical: ATmega328P (16 MHz, 2 KB SRAM, 32 KB flash)
MSP430 (Texas Instruments)
  • Industry-leading ultra-low power consumption
  • Active mode: as low as 100 µA/MHz
  • LPM4 standby: 0.1 µA
  • 16-bit architecture only — limited compute
  • No floating point unit
  • Limited peripheral selection vs STM32
  • Smaller ecosystem, fewer third-party libraries
  • More expensive per DMIPS than ARM Cortex-M
  • Best for: energy-harvesting sensors, always-on monitoring
ARM Cortex-M3 / M4
  • Full 32-bit — 4 GB address space
  • Hardware multiply in single cycle
  • Hardware divide (no software loop needed)
  • Optional single-precision FPU (M4)
  • NVIC: up to 240 interrupts, 256 priorities
  • MPU for memory protection and RTOS support
  • Unified von Neumann-style memory map (simple pointer model)
  • Low power comparable to MSP430 in stop modes
  • Massive ecosystem: FreeRTOS, Zephyr, HAL, countless libraries
  • Same skills work across ST, NXP, Nordic, TI chips
When AVR or MSP430 Still Makes Sense

AVR is still fine for simple hobbyist projects and rapid Arduino prototyping. MSP430 remains competitive in applications that literally run for 10+ years on a coin cell and don’t need significant computation. But for any serious embedded product — IoT devices, wearables, industrial control, audio processing — ARM Cortex-M is the professional choice.

Inside the Cortex-M4: Full Block Diagram

Now let’s open up the Cortex-M4 and understand every block inside it. This diagram matches the official ARM Cortex-M4 Technical Reference Manual (TRM). Every block shown here is present in chips like the STM32F407 and STM32F411.

CPU Core blocks
Debug & Trace blocks
Bus interfaces
Security & Protection
Interrupt & Power
ARM Cortex-M4 Processor Internal Architecture
Cortex-M4 Core 3-stage pipeline (Fetch / Decode / Execute)
Register bank: R0–R15
Thumb-2 ISA (ARMv7E-M)
Hardware multiply & divide
SysTick 24-bit timer
FPU Single-precision
IEEE 754 compliant
32 FP registers (S0–S31)
FMAC instruction
NVIC Nested Vectored Interrupt Controller
Up to 240 IRQs
8–256 priority levels
Tail-chaining
MPU Memory Protection Unit
Up to 8 regions
Privilege-level enforcement
WIC Wake-up Interrupt Controller
Detects interrupts
while core is powered off
ETM Embedded Trace Macrocell
Records every instruction
executed for replay
FPB Flash Patch Breakpoint
8 hardware breakpoints
Flash address remapping
AHB Lite Bus Matrix (32-bit) — ICode / DCode / System buses
DWT Data Watchpoint & Trace
4 comparators
Cycle counter (CYCCNT)
Profiling counters
ITM Instrumentation Trace Macrocell
printf-style debug
over SWO pin — no UART needed
AHB-AP AHB Access Port
Allows debugger to read
any memory/register
TPIU Trace Port Interface Unit
Serialises trace data
to external capture device
SW-DP / SWJ-DP Serial Wire Debug Port
or JTAG Debug Port
2-wire SWD or 4-wire JTAG
CoreSight ROM Table — lets the debugger auto-discover all debug blocks
Everything inside the dashed boundary is the ARM Cortex-M4 processor IP licensed by ARM. The chip vendor (e.g., STMicroelectronics) connects flash, SRAM, and peripherals to the AHB/APB buses outside this boundary.

Every Block Explained — What It Does and Why It Exists

1. The Cortex-M4 CPU Core

The core is a 3-stage pipeline processor: it simultaneously fetches the next instruction from flash, decodes the current instruction, and executes the previous one. This overlap means most simple instructions complete in a single clock cycle at full speed.

The register bank contains 16 registers (R0–R15), all 32 bits wide. Most computations use R0–R12 as general-purpose scratch registers. R13 is the Stack Pointer, R14 is the Link Register (stores return addresses), and R15 is the Program Counter.

A hardware multiply-accumulate unit (MAC) performs a 32×32-bit multiply in a single clock cycle, and a hardware divide unit completes in 2–12 cycles. On 8-bit AVR, a 32-bit multiply requires dozens of instructions. This difference is enormous for signal processing and control loops.

STM32F411 Datasheet (DocID026289, Section 3.1): “The ARM Cortex-M4 with FPU 32-bit RISC processor features exceptional code-efficiency, delivering the high-performance expected from an ARM core in the memory size usually associated with 8- and 16-bit devices.”

2. Floating Point Unit (FPU)

The FPU is optional — the Cortex-M3 does not have one; the Cortex-M4 typically includes it. It implements the IEEE 754 single-precision standard (32-bit floats) and adds 32 dedicated floating-point registers (S0–S31, which can also be viewed as 16 double-word registers D0–D15 for loading/storing 64-bit values — but arithmetic is single-precision only).

Without a hardware FPU, floating-point operations are emulated in software — a 32-bit float add might take 50–100 cycles via a software library. With the hardware FPU, the same operation takes 1 cycle. For motor control, PID loops, and audio processing, this difference determines whether a design is feasible at a given clock speed.

Enabling the FPU before use (bare-metal C, Cortex-M4)
/* The FPU is disabled by default after reset to save power.
   Enable it by setting CP10 and CP11 bits in the CPACR register.
   Must be done before any floating-point instruction executes. */

#define FPU_CPACR   (*((volatile uint32_t *)0xE000ED88U))

void fpu_enable(void)
{
    /* Set bits 20:23 to 0b1111 — full access to CP10 and CP11 */
    FPU_CPACR |= (0xFU << 20);

    /* Memory barriers to ensure the setting takes effect before
       the first FP instruction. Required by the Cortex-M4 TRM. */
    __asm volatile ("dsb");
    __asm volatile ("isb");
}

int main(void)
{
    fpu_enable();

    float result = 3.14159f * 2.0f;  /* Now runs in 1 cycle on the FPU */
    return (int)result;
}
Key Point

STM32CubeIDE and most ARM GCC toolchain configurations enable the FPU automatically when you compile with -mfpu=fpv4-sp-d16 -mfloat-abi=hard. The startup code included by the IDE calls the FPU enable sequence before main(). When programming bare-metal from scratch, you must enable it yourself, as shown above.

3. NVIC — Nested Vectored Interrupt Controller

The NVIC is arguably the most important peripheral inside the Cortex-M processor. It manages the entire interrupt system and is tightly coupled to the CPU core — so tightly that interrupt response takes only 12 clock cycles from interrupt assertion to the first instruction of the ISR executing.

Key features that make the NVIC far superior to legacy 8-bit interrupt controllers:

  • Up to 240 external interrupt lines — chip vendors can implement 1 to 240 of these. STM32F411 uses 62 maskable channels plus the 16 Cortex-M4 internal exceptions.
  • 16 priority levels (on STM32F411) — lower number = higher priority. A priority-0 interrupt can preempt a priority-15 ISR that is already running. This is called preemption.
  • Tail-chaining — if two interrupts are pending, the NVIC services the second one immediately after the first without re-saving/restoring CPU state. This saves 12 cycles per chained interrupt.
  • Automatic state save/restore — the CPU automatically pushes R0–R3, R12, LR, PC, and xPSR to the stack on interrupt entry. You do not write any save/restore code in your ISR.
  • Late-arrival handling — if a higher-priority interrupt arrives while the CPU is still in the pre-emption save phase, the NVIC switches to serve the higher-priority one first.
STM32F411 Datasheet (DocID026289, Section 3.10): “The devices embed a nested vectored interrupt controller able to manage 16 priority levels, and handle up to 62 maskable interrupt channels plus the 16 interrupt lines of the Cortex-M4 with FPU.”
Enabling an interrupt in the NVIC — bare-metal C
/* NVIC register base address (from ARM Cortex-M4 TRM) */
#define NVIC_ISER0   (*((volatile uint32_t *)0xE000E100U))
#define NVIC_IPR_BASE ((volatile uint32_t *)0xE000E400U)

/* STM32F411: IRQ number for USART1 = 37 (from Reference Manual RM0383) */
#define USART1_IRQn  37

void nvic_enable_irq(uint32_t irq_num, uint8_t priority)
{
    /* Step 1: Set the priority (8 priority bits, upper 4 are implemented) */
    /* Priority register: each IRQ gets 8 bits, 4 IRQs per 32-bit register */
    uint32_t reg_index = irq_num / 4;
    uint32_t bit_shift  = (irq_num % 4) * 8;
    NVIC_IPR_BASE[reg_index] |= ((uint32_t)(priority << 4) << bit_shift);

    /* Step 2: Enable the interrupt in ISER (Interrupt Set Enable Register) */
    /* Each ISER register handles 32 IRQs; ISER0 covers IRQ 0–31, ISER1: 32–63 */
    volatile uint32_t *iser = (volatile uint32_t *)0xE000E100U;
    iser[irq_num / 32] = (1U << (irq_num % 32));
}

int main(void)
{
    nvic_enable_irq(USART1_IRQn, 5);  /* Enable USART1 IRQ at priority 5 */

    /* Enable global interrupts (clear PRIMASK) */
    __asm volatile ("cpsie i");

    while (1) { /* application loop */ }
}

/* USART1 ISR — name must match the vector table entry */
void USART1_IRQHandler(void)
{
    /* Read received byte and echo it */
    volatile uint32_t *USART1_DR  = (volatile uint32_t *)0x40011004U;
    volatile uint32_t *USART1_SR  = (volatile uint32_t *)0x40011000U;

    if (*USART1_SR & (1U << 5)) {      /* RXNE flag set? */
        uint8_t byte = (uint8_t)(*USART1_DR & 0xFF);
        *USART1_DR = byte;              /* Echo back */
    }
}

4. MPU — Memory Protection Unit

The MPU lets you divide the 4 GB address space into up to 8 protected regions. Each region has configurable access rules: read-only, read-write, execute-never, privileged-only, etc. If code tries to access a region in a way that violates the rules, the CPU triggers a MemManage fault — the fault handler can then reset the offending task, log an error, or halt the system.

This is essential for RTOS-based systems: the kernel runs in privileged mode and configures the MPU so that each user task can only access its own memory. One buggy task cannot corrupt another task’s data or the kernel’s data. The STM32F411 datasheet confirms the MPU covers up to 8 subregions per region, with sizes between 32 bytes and the full 4 GB.

MPU Real-World Use Case

Imagine a medical device running two FreeRTOS tasks: one reads sensor data, one drives a display. Without MPU, a bug in the display task (e.g., an out-of-bounds array write) could silently corrupt the sensor data buffer — a dangerous situation in a medical context. With the MPU enabled, any illegal access immediately triggers a fault, the RTOS catches it, logs the fault address, and restarts only the faulty task. This is how safety-certified firmware is built.

5. DWT — Data Watchpoint and Trace

The DWT has two main uses that embedded developers rely on heavily:

  • Hardware watchpoints — set a data watchpoint on a memory address. The CPU halts automatically when that address is read or written. Invaluable for catching memory corruption bugs.
  • Cycle counter (CYCCNT) — a 32-bit counter that increments every CPU clock cycle. Use it to measure exactly how long a function takes in cycles, without a logic analyser or oscilloscope. Accurate to within ±1 cycle.
Using DWT cycle counter for function timing — bare-metal C
#define DWT_CTRL    (*((volatile uint32_t *)0xE0001000U))
#define DWT_CYCCNT  (*((volatile uint32_t *)0xE0001004U))
#define CoreDebug_DEMCR (*((volatile uint32_t *)0xE000EDFC U))

void dwt_init(void)
{
    CoreDebug_DEMCR |= (1U << 24);  /* Enable DWT */
    DWT_CYCCNT = 0;                  /* Reset counter */
    DWT_CTRL   |= (1U << 0);        /* Start counting */
}

uint32_t dwt_get_cycles(void)
{
    return DWT_CYCCNT;
}

/* Example: time a sorting function */
void measure_sort(int *arr, int len)
{
    uint32_t start = dwt_get_cycles();
    bubble_sort(arr, len);           /* your function here */
    uint32_t end   = dwt_get_cycles();

    uint32_t cycles_taken = end - start;
    /* At 100 MHz: cycles_taken / 100 = microseconds */
}

6. ITM — Instrumentation Trace Macrocell

The ITM is a hardware printf debug channel. Instead of routing debug messages through a UART (which requires wiring and affects timing), you write to ITM stimulus registers. The messages flow out through the Serial Wire Output (SWO) pin — the same wire used for SWD debugging — and appear in your IDE’s debug console.

ITM printf — sends characters over SWO without UART
#include <stdint.h>

#define ITM_STIM0  (*((volatile uint32_t *)0xE0000000U))
#define ITM_TER    (*((volatile uint32_t *)0xE0000E00U))

/* Send one character via ITM stimulus port 0 */
void itm_putchar(char c)
{
    /* Wait until port 0 is ready (bit 0 of STIM0) */
    while ((ITM_STIM0 & 1U) == 0);

    /* Write the byte — this queues it to the SWO output */
    *((volatile uint8_t *)0xE0000000U) = (uint8_t)c;
}

void itm_print(const char *s)
{
    while (*s) itm_putchar(*s++);
}

int main(void)
{
    /* Enable ITM stimulus port 0 */
    ITM_TER |= (1U << 0);

    itm_print("Hello from ITM!\n");

    while (1) {}
}

7. ETM — Embedded Trace Macrocell

The ETM records every single instruction the CPU executes and streams the trace data out through dedicated trace pins (TRACED[3:0]). An external trace capture device (like a SEGGER J-Trace or a Lauterbach) stores this stream, allowing you to replay the exact execution history after a bug occurs — even in real-time with no CPU slowdown.

ETM is the most powerful debugging tool available. It can show you exactly which code path led to a crash, without needing to reproduce the bug or halt the CPU. It is used in avionics, automotive, and industrial applications where bugs must be analysed in production.

8. FPB — Flash Patch Breakpoint

The FPB provides 8 hardware breakpoints — address comparators that halt the CPU when the program counter reaches a specific address. Unlike software breakpoints (which replace instructions with a BKPT opcode in flash), hardware breakpoints work without modifying flash memory and can be set in read-only regions. The FPB also has a patching capability to remap flash addresses to SRAM, useful for fixing bugs in field-deployed firmware without full re-flashing.

9. WIC — Wake-up Interrupt Controller

When the Cortex-M4 enters deep sleep mode, the CPU core itself is powered off. The WIC is a tiny, ultra-low-power circuit that stays powered and watches for interrupt signals. When an external interrupt fires, the WIC powers on the CPU, and the NVIC processes the interrupt normally. This allows the dynamic efficiency features of STM32 chips — the CPU sleeps between events, yet responds in microseconds.

The Three AHB-Lite Bus Interfaces

The Cortex-M4 exposes three separate bus interfaces to the outside world. Each has a specific purpose that optimises performance:

Cortex-M4 External Bus Connections to Chip Memory & Peripherals
ICode Bus (I-Bus) Instruction fetch
Connects to Flash memory
Read-only, 32-bit
Used for fetching instructions
DCode Bus (D-Bus) Data accesses to Flash
Connects to Flash memory
Read-only, 32-bit
Used for const data in flash
System Bus (S-Bus) Data + instruction (SRAM)
Connects to SRAM & peripherals
Read/write, 32-bit
All peripheral register accesses
Multi-AHB Bus Matrix (chip vendor logic — outside the Cortex-M4 boundary)
Flash
512 KB
SRAM
128 KB
AHB1
Peripherals
(DMA, GPIO)
AHB2
Peripherals
(USB OTG)
APB1
Low-speed
peripherals
APB2
High-speed
peripherals
Separating instruction fetch (ICode) from data access (DCode / S-Bus) means the CPU can fetch the next instruction while simultaneously reading data from flash — a key performance optimisation.

From Core to Chip: Meet the STM32F411

Now let’s see how all the Cortex-M4 blocks above appear in a real production chip. The STM32F411xC/xE from STMicroelectronics is an excellent learning target — it is one of the most popular Cortex-M4 chips on development boards and in student labs worldwide.

Source: STM32F411xC/xE Datasheet, DocID026289 Rev 6 (December 2016) — the official document used for all specifications in this section.
ParameterSTM32F411xCSTM32F411xE
CPU CoreARM Cortex-M4 with FPU (single-precision)
Maximum CPU Frequency100 MHz
Performance125 DMIPS (Dhrystone 2.1 benchmark)
Flash Memory256 KB512 KB
SRAM128 KB (zero wait state at CPU clock speed)
Supply Voltage1.7 V to 3.6 V
Temperature Range−40 °C to +125 °C (junction)
General Purpose Timers7 (six 16-bit + two 32-bit)
Advanced-control Timer1 (TIM1 — PWM with dead-time for motor control)
SysTick Timer24-bit countdown — used by RTOS scheduler
I2C Interfaces3 (SMBus/PMBus compatible)
USART3 (up to 12.5 Mbit/s)
SPI / I2S5 SPI / 5 I2S (2 full-duplex I2S)
USBUSB 2.0 OTG Full Speed (12 Mbit/s) with on-chip PHY
SDIO1 (SD card / MMC / eMMC)
ADC1 × 12-bit, up to 2.4 MSPS, 10 or 16 channels
DMA2 controllers × 8 streams = 16 DMA streams
GPIO Pins36 / 50 / 81 (package dependent)36 / 50 / 81
Interrupt Channels62 maskable + 16 Cortex-M4 exceptions
Debug InterfaceSWD (2-wire) and JTAG (4-wire)
ART AcceleratorYes — eliminates flash wait states at 100 MHz
CRC UnitHardware CRC-32 calculation
Unique ID96-bit factory-programmed unique device ID
Power — Stop mode42 µA typical at 25°C (fast wakeup)
Power — Standby mode1.8 µA at 25°C (without RTC)

The ART Accelerator — Zero-Wait Flash at 100 MHz

Flash memory is physically slower than a 100 MHz CPU clock. Without acceleration, the CPU would need to insert wait states (idle cycles) every time it fetches an instruction from flash — significantly cutting effective throughput.

STMicroelectronics solves this with their proprietary ART Accelerator (Adaptive Real-Time memory Accelerator). It adds a prefetch queue and branch cache between the CPU and flash. The CPU’s ICode bus requests instructions, and the ART Accelerator serves them from its cache with zero wait states — as if the flash were as fast as SRAM.

STM32F411 Datasheet (Section 3.2): “Based on CoreMark benchmark, the performance achieved thanks to the ART accelerator is equivalent to 0 wait state program execution from Flash memory at a CPU frequency up to 100 MHz.”
ART Accelerator — How Flash Wait States Are Eliminated
CPU Core
Requests next
instruction every
10 ns @ 100 MHz
→
ART Accelerator
Prefetch queue:
reads ahead in flash
Branch cache:
stores hot code paths
→
Flash Memory
Physical access time:
~30–60 ns
(needs 3–6 wait states
at 100 MHz without ART)
The ART Accelerator hides flash latency by speculating which instructions the CPU will need next and pre-loading them. Cache hits (typical for sequential code and loops) cost 0 extra cycles.

DMA — Direct Memory Access Controller

The STM32F411 has two independent DMA controllers (DMA1 and DMA2), each with 8 streams, giving 16 DMA streams total. A DMA stream can transfer data between any combination of memory and peripheral without CPU involvement.

Example: you configure DMA to receive 256 bytes from a USART and write them directly to a buffer in SRAM. The CPU is free to execute other code. When the DMA completes, it fires an interrupt. The CPU wakes, processes the buffer, and goes back to other work. This is fundamentally more efficient than polling or interrupt-per-byte approaches.

DMA memory-to-memory transfer concept (bare-metal C)
#include <stdint.h>

/* DMA2 Stream0 register base (from STM32F411 Reference Manual RM0383) */
#define DMA2_BASE     0x40026400U
#define DMA2_S0CR     (*((volatile uint32_t *)(DMA2_BASE + 0x10U)))
#define DMA2_S0NDTR   (*((volatile uint32_t *)(DMA2_BASE + 0x14U)))
#define DMA2_S0PAR    (*((volatile uint32_t *)(DMA2_BASE + 0x18U)))
#define DMA2_S0M0AR   (*((volatile uint32_t *)(DMA2_BASE + 0x1CU)))

/* RCC: Enable DMA2 clock */
#define RCC_AHB1ENR   (*((volatile uint32_t *)0x40023830U))
#define RCC_DMA2EN    (1U << 22)

static uint32_t src_buf[64] = { [0 ... 63] = 0xDEADBEEFU };
static uint32_t dst_buf[64];

void dma2_memcpy(void)
{
    /* Enable DMA2 clock */
    RCC_AHB1ENR |= RCC_DMA2EN;

    /* Configure DMA2 Stream0 for memory-to-memory */
    DMA2_S0CR   = 0;                         /* Disable stream first      */
    DMA2_S0PAR  = (uint32_t)src_buf;         /* Source: src_buf           */
    DMA2_S0M0AR = (uint32_t)dst_buf;         /* Destination: dst_buf      */
    DMA2_S0NDTR = 64;                        /* 64 items to transfer      */

    /* Direction: memory-to-memory, increment both, 32-bit words, channel 0 */
    DMA2_S0CR = (1U << 14) |  /* MEM2MEM bit: memory-to-memory mode */
                (1U << 10) |  /* MINC: memory pointer increment     */
                (1U << 9)  |  /* PINC: peripheral pointer increment */
                (2U << 11) |  /* MSIZE: 32-bit memory data width    */
                (2U << 13);   /* PSIZE: 32-bit peripheral data width*/

    DMA2_S0CR |= (1U << 0);   /* EN: Enable the stream — transfer starts */

    /* Wait for transfer complete (TC flag in DMA2_LISR, bit 5) */
    volatile uint32_t *DMA2_LISR = (volatile uint32_t *)0x40026400U;
    while (!(*DMA2_LISR & (1U << 5)));
}

Understanding NVIC Priority Groups

One of the most important — and most confused — topics in Cortex-M programming is interrupt priority grouping. Let’s make it crystal clear.

The STM32F411 implements 4 priority bits per interrupt (the upper 4 of the 8-bit priority field). These 4 bits are split into two fields by the PRIGROUP setting in the Application Interrupt and Reset Control Register (AIRCR):

PRIGROUPPreemption Priority BitsSub-Priority BitsPreemption LevelsSub-levels per Preemption
0b011 (default)4 bits0 bits161
0b1003 bits1 bit82
0b1012 bits2 bits44
0b1101 bit3 bits28
0b1110 bits4 bits116

Preemption priority determines if one ISR can interrupt another that is currently running (higher preemption priority wins). Sub-priority only matters when two interrupts with the same preemption priority are pending simultaneously — the one with the lower sub-priority number fires first. Sub-priority never causes preemption.

The Most Common Beginner Bug with NVIC

On STM32 chips with 4 implemented priority bits, the valid priority values are 0, 16, 32, 48 … 240 (multiples of 16) when writing raw to the 8-bit IPR register. Many beginners write values 0–15 directly and wonder why priority 5 and priority 7 behave identically — the lower 4 bits are ignored. Always shift your priority value left by 4 bits, or use the STM32 HAL’s NVIC_SetPriority() which handles this for you.

Setting Up the Development Environment

Before writing firmware, you need a toolchain on your PC. This course uses STM32CubeIDE — a free, Eclipse-based IDE developed by STMicroelectronics. It integrates:

  • Editor — syntax highlighting, code completion, navigation
  • Compiler — ARM GCC (arm-none-eabi-gcc) bundled inside
  • Linker — produces the .elf and .bin files loaded to the chip
  • Flash programmer — programs the binary over ST-LINK (USB)
  • Debugger — GDB + ST-LINK for breakpoints, step-through, register view
  • SWV / ITM viewer — displays ITM printf output and DWT profiling
1
Download STM32CubeIDE
Go to st.com, search for STM32CubeIDE. Download the installer for your OS (Windows / macOS / Linux). It is free with a free STMicro account. The installer bundles everything: IDE, ARM GCC, ST-LINK drivers.
2
Install ST-LINK USB Drivers (Windows)
On Windows, the ST-LINK virtual COM port and mass storage drivers are installed automatically with CubeIDE. On Linux, add your user to the plugdev group and install the udev rules provided by STMicro. On macOS, no drivers are needed.
3
Create a New Project
File → New → STM32 Project → select your board (e.g., NUCLEO-F411RE or STM32F407G-DISC1). CubeIDE generates the startup file, linker script, and system clock initialisation automatically.
4
Build and Flash
Press the hammer icon (Build) to compile. Press the bug icon (Debug) to compile, flash, and halt at main() with the debugger attached. The ST-LINK interface on the Nucleo/Discovery board handles the USB-to-SWD programming automatically — no separate programmer needed.
5
Verify: Blink the LED
The classic “Hello World” for embedded: toggle the LED in a loop. If your LED blinks, your entire toolchain — compiler, linker, flash programmer, and hardware — is confirmed working. Don’t skip this step before attempting more complex code.
Alternative Open-Source Toolchain (Linux / macOS)

If you prefer working without an IDE, the complete open-source toolchain is:

Linux install (Ubuntu/Debian)
# Install ARM GCC compiler
sudo apt install gcc-arm-none-eabi binutils-arm-none-eabi

# Install OpenOCD (flash programmer and GDB server)
sudo apt install openocd

# Install udev rules for ST-LINK
sudo apt install stlink-tools

# Verify installation
arm-none-eabi-gcc --version
openocd --version
Compile and flash a project from command line
# Compile (assuming main.c and startup.s are present)
arm-none-eabi-gcc -mcpu=cortex-m4 -mthumb -mfpu=fpv4-sp-d16 \
    -mfloat-abi=hard -O2 -Wl,-T,linker.ld \
    startup.s main.c -o firmware.elf

# Convert to binary
arm-none-eabi-objcopy -O binary firmware.elf firmware.bin

# Flash via ST-LINK using OpenOCD
openocd -f interface/stlink.cfg \
        -f target/stm32f4x.cfg \
        -c "program firmware.elf verify reset exit"

The Hardware Target: STM32F407 Discovery Board

This course uses the STM32F407G-DISC1 (STM32 Discovery) board as the primary hardware target. It is also compatible with the STM32F411 Nucleo boards since both use Cortex-M4 cores with the same programmer’s model.

Board FeatureSTM32F407G-DISC1 DetailsWhy It Matters for Learning
Microcontroller STM32F407VGT6 (Cortex-M4, 168 MHz, 1 MB flash, 192 KB SRAM) A step up from F411 — same M4 core, more flash/SRAM for larger projects
Built-in Programmer ST-LINK/V2 (mini USB connector, SWD interface) No external programmer needed — plug USB and flash immediately
User LEDs 4 LEDs: LD3 (orange, PD13), LD4 (green, PD12), LD5 (red, PD14), LD6 (blue, PD15) 4 GPIOs to toggle for LED experiments without extra wiring
User Button B1 button connected to PA0 Interrupt / polling input for button experiments
MEMS accelerometer LIS302DL or LIS3DSH connected via SPI Real sensor to practice SPI communication
USB OTG Full speed + high speed (with external PHY connector) USB device/host experiments without extra hardware
Audio DAC CS43L22 connected via I2S + I2C Audio output experiments — play tones via I2S/I2C
Power Supply USB-powered (5V → 3V regulator on board) Single USB cable powers everything
Suggested images: STM32F407 Discovery board top-down photo showing LED positions and button • STM32CubeIDE debug session screenshot showing register view and breakpoint • SWD probe connection diagram (2-wire: SWDIO + SWCLK) • DWT cycle counter graph overlaid on C code in profiler view

A First Look at the STM32 Clock Tree

One of the most important things to understand about STM32 chips — and one that trips up almost every beginner — is the clock tree. Peripherals on an STM32 are clocked by different bus frequencies, and each peripheral must have its clock enabled before you can use it.

STM32F411 Datasheet (Section 3.12): “On reset the 16 MHz internal RC oscillator is selected as the default CPU clock. The application can then select as system clock either the RC oscillator or an external 4–26 MHz clock source. This clock source is input to a PLL thus allowing to increase the frequency up to 100 MHz.”
STM32F411 Simplified Clock Tree
HSI
Internal 16 MHz RC
Default after reset
±1% accuracy
HSE
External crystal
4–26 MHz
Better accuracy
→
Main PLL
Multiplies input
up to 100 MHz
Configured via RCC
→
SYSCLK
100 MHz
System clock
feeds CPU, AHB
→
AHB Bus
100 MHz
GPIO, DMA, Flash
APB2 Bus
100 MHz
USART1, SPI1, ADC
APB1 Bus
50 MHz
USART2, I2C, SPI2/3
General-purpose timers
APB1 maximum is 50 MHz on STM32F411. APB2 and AHB maximum is 100 MHz. Timer clocks can be 2× their APB clock if the APB prescaler ≠ 1.
Configuring STM32F411 to run at 100 MHz using PLL from HSI (bare-metal C)
/* Registers from STM32F411 Reference Manual (RM0383) */
#define RCC_CR         (*((volatile uint32_t *)0x40023800U))
#define RCC_PLLCFGR    (*((volatile uint32_t *)0x40023804U))
#define RCC_CFGR       (*((volatile uint32_t *)0x40023808U))
#define FLASH_ACR      (*((volatile uint32_t *)0x40023C00U))

void system_clock_100MHz(void)
{
    /* Step 1: Enable ART Accelerator and set 3 flash wait states for 100 MHz */
    FLASH_ACR = (1U << 9) |   /* PRFTEN: prefetch enable   */
                (1U << 10)|   /* ICEN:   instruction cache */
                (1U << 11)|   /* DCEN:   data cache        */
                3U;           /* LATENCY: 3 wait states    */

    /* Step 2: Configure PLL: HSI (16 MHz) as source
       PLLM=8 : divides 16 MHz to 2 MHz (VCO input)
       PLLN=200: multiplies 2 MHz to 400 MHz (VCO output)
       PLLP=4 : divides 400 MHz to 100 MHz (SYSCLK)
       PLLQ=8 : divides 400 MHz to 50 MHz (USB — needs 48 MHz, not used here) */
    RCC_PLLCFGR = (8U << 0)   |  /* PLLM = 8  */
                  (200U << 6) |  /* PLLN = 200 */
                  (1U << 16)  |  /* PLLP = 4 (bits 17:16 = 01 means /4) */
                  (8U << 24)  |  /* PLLQ = 8  */
                  (0U << 22);    /* PLLSRC = HSI */

    /* Step 3: Enable PLL and wait for it to lock */
    RCC_CR |= (1U << 24);              /* PLLON = 1 */
    while (!(RCC_CR & (1U << 25)));   /* Wait for PLLRDY */

    /* Step 4: Set AHB prescaler = 1, APB1 = /2 (50 MHz), APB2 = /1 (100 MHz) */
    RCC_CFGR = (4U << 10) |  /* PPRE1: APB1 /2  = 50 MHz  */
               (0U << 13);   /* PPRE2: APB2 /1  = 100 MHz */

    /* Step 5: Switch SYSCLK to PLL */
    RCC_CFGR |= (2U << 0);            /* SW = PLL  */
    while (((RCC_CFGR >> 2) & 3) != 2); /* Wait for SWS = PLL */
}

Frequently Asked Questions

What is the difference between the STM32F407 and STM32F411?
Both use the same ARM Cortex-M4 core with FPU and are fully software-compatible. The STM32F407 runs at up to 168 MHz with 1 MB flash and 192 KB SRAM, plus more advanced peripherals (Ethernet, camera interface, DAC, two I2S). The STM32F411 runs at up to 100 MHz with 512 KB flash and 128 KB SRAM, targeting a lower-power, lower-cost profile. For this course, the programmer’s model — registers, NVIC, SysTick, bus structure — is identical between them.
Do I need to enable the clock for every peripheral before using it?
Yes, on all STM32 chips. After reset, the clock to most peripherals is gated off to save power. Before accessing any peripheral registers, you must set the corresponding bit in the RCC (Reset and Clock Control) AHB1ENR, AHB2ENR, APB1ENR, or APB2ENR register. Forgetting this is the single most common beginner bug — your code seems correct but the peripheral does nothing.
What is SWD and how is it different from JTAG?
SWD (Serial Wire Debug) is a 2-wire ARM-standard debug interface: SWDIO (data, bidirectional) and SWCLK (clock). JTAG is an older standard using 4–5 wires: TMS, TCK, TDI, TDO, and optional TRST. SWD is preferred on Cortex-M because it saves GPIO pins and is just as capable for programming and stepping through code. The ST-LINK on Nucleo/Discovery boards uses SWD by default. Only if you need full instruction trace (ETM) do you need additional trace pins beyond SWD.
Can I use STM32CubeIDE on Linux?
Yes. STM32CubeIDE has native Linux x86_64 and ARM64 installers. After installation, add the udev rule for ST-LINK (sudo cp /opt/st/stm32cubeide_*/plugins/com.st.stm32cube.ide.mcu.externaltools.stlink-gdb-server*/tools/udev/rules.d/*.rules /etc/udev/rules.d/ then sudo udevadm control –reload) and add your user to the plugdev group. Then connect the board and flash normally.
What does the 96-bit Unique ID in STM32 chips do?
Every STM32 chip has a 96-bit (12-byte) factory-programmed unique identifier at address 0x1FFF7A10 on STM32F4 devices. This ID is guaranteed to be unique across all STM32 chips ever manufactured. It is used for: generating unique Bluetooth or Ethernet MAC addresses from a single firmware image, licensing software to a specific device, unique device serial numbers for asset tracking, and cryptographic key derivation.
What is the difference between AHB and APB buses?
AHB (Advanced High-performance Bus) is the fast bus directly connected to the CPU — memory, DMA, and high-speed peripherals (GPIO, USB) sit on AHB. It runs at the full system clock speed (up to 100 MHz on STM32F411). APB (Advanced Peripheral Bus) is a lower-speed bus for peripherals that don’t need full clock speed — USART, I2C, SPI, timers. APB1 is the slow bus (max 50 MHz on STM32F411) and APB2 is the fast bus (max 100 MHz). Putting slower peripherals on APB reduces power consumption and simplifies bus arbitration.
Why does the STM32F411 have two DMA controllers?
Having two independent DMA controllers with 8 streams each allows multiple simultaneous transfers. For example, DMA1 can be transferring audio samples from an I2S peripheral to SRAM while DMA2 simultaneously transfers data from SRAM to a SPI display — all without any CPU involvement and without either transfer blocking the other. The multi-AHB bus matrix ensures both DMA controllers can access different bus slaves simultaneously without conflicts.
Next: The Cortex-M Programmer’s Model

Lecture 3 dives into the register bank in detail — R0 through R15, the two stack pointers (MSP and PSP), the Link Register, the Program Counter, and the Program Status Register (xPSR). Understanding these is the foundation for writing interrupt handlers and RTOS code.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *