Cortex-M Processor Peripherals
The ARM Cortex-M3/M4 processor integrates a rich set of peripherals directly on-chip alongside the CPU core: NVIC, SCB, SysTick, MPU, FPU, and Debug. All are memory-mapped in the Private Peripheral Bus (PPB) region and accessible through CMSIS headers. This lecture covers each peripheral’s role, key registers, and firmware usage.
Introduction
EmbeddedPathashala — ARM Cortex-M Course Home All Lectures STM32 Projects
Processor-Internal Peripheral Overview
Unlike the vendor peripherals (USART, SPI, GPIO) that ST adds to the STM32 chip, the following peripherals are part of the ARM Cortex-M processor core itself. They are identical across every Cortex-M3 and Cortex-M4 implementation — whether the chip is from ST, NXP, TI, or any other ARM licensee. Their addresses, register layouts, and behaviour are all defined by ARM and cannot be changed by the chip vendor.
Thumb-2
3-stage pipeline
Interrupt Controller
Up to 240 IRQs
Single-precision
IEEE 754
extensions
Unit
8 regions
Block
Fault & config
timer
RTOS tick
ETM, TPIU
CoreSight
| Peripheral | Base Address | Present on | Key function |
|---|---|---|---|
| ITM | 0xE0000000 | M3, M4 | Instrumentation Trace Macrocell — printf-style debug over SWO |
| DWT | 0xE0001000 | M3, M4 | Data Watchpoint & Trace — cycle counter, watchpoints |
| FPB | 0xE0002000 | M3, M4 | Flash Patch & Breakpoint — hardware breakpoints (up to 8) |
| SCS (NVIC) | 0xE000E000 | M0, M3, M4 | System Control Space — NVIC, SysTick, SCB, MPU registers |
| TPIU | 0xE0040000 | M3, M4 | Trace Port Interface — serialises trace data to pins |
| ETM | 0xE0041000 | M3, M4 (optional) | Embedded Trace Macrocell — instruction trace stream |
NVIC — Nested Vectored Interrupt Controller
The NVIC is the heart of the Cortex-M exception system. It accepts interrupt requests from up to 240 peripheral sources, prioritises them, and signals the CPU core to take the appropriate exception. The “nested” part means the NVIC can preempt a lower-priority ISR with a higher-priority one automatically.
Key NVIC capabilities:
• Up to 240 external IRQs (implementation-defined; STM32F411 uses ~85)
• 256 programmable priority levels (8-bit priority fields; STM32F4 implements the top 4 bits = 16 levels)
• Tail-chaining: back-to-back ISR execution without full un-stack/re-stack overhead
• Late arrival: higher-priority IRQ arriving during stacking preempts immediately
• Interrupt enable/disable per IRQ, plus global masking via PRIMASK / BASEPRI
NVIC registers (inside SCS at 0xE000E000)
| Register | Offset | Purpose |
|---|---|---|
| ISER[0..7] | 0xE000E100 | Interrupt Set-Enable Registers — write 1 to enable an IRQ |
| ICER[0..7] | 0xE000E180 | Interrupt Clear-Enable Registers — write 1 to disable an IRQ |
| ISPR[0..7] | 0xE000E200 | Interrupt Set-Pending Registers — software-trigger an IRQ |
| ICPR[0..7] | 0xE000E280 | Interrupt Clear-Pending Registers — cancel a pending IRQ |
| IABR[0..7] | 0xE000E300 | Interrupt Active Bit Registers — read which IRQs are active |
| IPR[0..59] | 0xE000E400 | Interrupt Priority Registers — 8 bits per IRQ, 4 IRQs per 32-bit word |
| STIR | 0xE000EF00 | Software Trigger Interrupt Register — write IRQ number to trigger |
NVIC usage with CMSIS
#include "stm32f4xx.h" /* includes CMSIS core_cm4.h */
/* Enable USART2 IRQ at priority 5 (STM32F4: 4-bit priority → value 5<<4 = 0x50) */
void nvic_setup_usart2(void)
{
NVIC_SetPriority(USART2_IRQn, 5); /* priority level 5 */
NVIC_EnableIRQ(USART2_IRQn); /* set bit in ISER */
}
/* Disable USART2 IRQ */
void nvic_disable_usart2(void)
{
NVIC_DisableIRQ(USART2_IRQn); /* set bit in ICER */
}
/* Software-trigger an IRQ (useful for testing or deferred work) */
void trigger_usart2_isr_software(void)
{
NVIC_SetPendingIRQ(USART2_IRQn); /* write to ISPR */
}
/* Direct register access (no CMSIS) — enable IRQ 38 (USART2 on STM32F4) */
#define NVIC_ISER1 (*(volatile uint32_t *)0xE000E104U)
NVIC_ISER1 = (1U << (38 - 32)); /* IRQ38 is bit 6 of ISER[1] */
Priority grouping
The 4 implemented priority bits on STM32F4 are split into preemption priority (determines nesting) and sub-priority (breaks ties between same-preemption-level IRQs). This split is configured in the SCB’s AIRCR register via the PRIGROUP field. CMSIS provides NVIC_SetPriorityGrouping() to set it.
/* Example: 3 bits preemption, 1 bit sub-priority (group 4) */
NVIC_SetPriorityGrouping(4); /* PRIGROUP = 4 → [7:5] preempt, [4] sub */
/* IRQ with preempt=3, sub=0 */
uint32_t pri = NVIC_EncodePriority(4, 3, 0);
NVIC_SetPriority(TIM2_IRQn, pri);
SCB — System Control Block
The System Control Block (SCB) contains the configuration and status registers for the processor core — fault status, sleep control, exception prioritisation for system exceptions (HardFault, MemManage, BusFault, UsageFault, SVC, PendSV, SysTick), and the vector table offset register.
| Register | Address | Key function |
|---|---|---|
| CPUID | 0xE000ED00 | Read-only — implementer (ARM=0x41), part number, revision |
| ICSR | 0xE000ED04 | Interrupt Control & State — pending NMI/SysTick/PendSV, active exception number |
| VTOR | 0xE000ED08 | Vector Table Offset — relocate the vector table to SRAM or a different Flash address |
| AIRCR | 0xE000ED0C | Application Interrupt & Reset Control — PRIGROUP, SYSRESETREQ (software reset), VECTCLRACTIVE |
| SCR | 0xE000ED10 | System Control — SLEEPDEEP, SLEEPONEXIT, SEVONPEND |
| CCR | 0xE000ED14 | Configuration & Control — STKALIGN (8-byte exception stack), DIV_0_TRP, UNALIGN_TRP |
| SHPR1–3 | 0xE000ED18 | System Handler Priority — priorities for MemManage, BusFault, UsageFault, SVC, PendSV, SysTick |
| SHCSR | 0xE000ED24 | System Handler Control & State — enable MemManage/BusFault/UsageFault; pending/active bits |
| CFSR | 0xE000ED28 | Configurable Fault Status — UFSR (UsageFault), BFSR (BusFault), MMFSR (MemManage) combined |
| HFSR | 0xE000ED2C | HardFault Status — DEBUGEVT, FORCED (escalated from configurable fault), VECTTBL |
| MMFAR | 0xE000ED34 | MemManage Fault Address — address that caused the fault (valid when MMFSR.MMARVALID=1) |
| BFAR | 0xE000ED38 | BusFault Address — address of faulting memory access (valid when BFSR.BFARVALID=1) |
Common SCB operations
#include "core_cm4.h" /* CMSIS */
/* 1. Software system reset */
void system_reset(void)
{
__DSB(); /* ensure all memory writes are complete */
SCB->AIRCR = (0x5FAUL << SCB_AIRCR_VECTKEY_Pos) |
SCB_AIRCR_SYSRESETREQ_Msk;
while (1); /* wait for reset */
}
/* 2. Relocate vector table to SRAM (after copying it there) */
void relocate_vtable_to_sram(void)
{
/* Copy vector table from Flash to SRAM */
memcpy((void *)0x20000000, (void *)0x08000000, 0x200);
__DSB();
SCB->VTOR = 0x20000000; /* now uses SRAM vector table */
__DSB();
}
/* 3. Enable configurable faults */
void enable_configurable_faults(void)
{
SCB->SHCSR |= SCB_SHCSR_USGFAULTENA_Msk | /* UsageFault */
SCB_SHCSR_BUSFAULTENA_Msk | /* BusFault */
SCB_SHCSR_MEMFAULTENA_Msk; /* MemManage */
}
/* 4. Read CPUID to identify the processor */
void print_cpuid(void)
{
uint32_t cpuid = SCB->CPUID;
uint8_t impl = (cpuid >> 24) & 0xFF; /* 0x41 = ARM */
uint8_t part = (cpuid >> 4) & 0xFFF; /* 0xC24 = Cortex-M4 */
uint8_t rev = cpuid & 0xF;
/* STM32F411: impl=0x41, part=0xC24, rev=1 */
}
SysTick — System Timer
SysTick is a 24-bit countdown timer built into every Cortex-M processor. It fires the SysTick exception (exception number 15, the highest-numbered system exception) every time the counter counts from the reload value down to zero. Its primary use is as the RTOS scheduling tick — a free-running periodic interrupt that triggers the task switcher.
| Register | Address | Purpose |
|---|---|---|
| SYST_CSR | 0xE000E010 | Control & Status: ENABLE (start), TICKINT (enable IRQ), CLKSOURCE (AHB or AHB/8), COUNTFLAG (read-to-clear overflow) |
| SYST_RVR | 0xE000E014 | Reload Value Register — 24-bit value loaded when counter reaches 0 |
| SYST_CVR | 0xE000E018 | Current Value Register — write any value to clear to 0 |
| SYST_CALIB | 0xE000E01C | Calibration — TENMS field (value for 10 ms tick at default frequency) |
Configuring SysTick for a 1 ms tick (STM32F411 at 100 MHz)
/* STM32F411 core clock = 100 MHz = 100,000,000 Hz
For 1 ms tick: reload = (100,000,000 / 1000) - 1 = 99999 */
void systick_init_1ms(void)
{
SysTick->LOAD = 99999U; /* reload value for 1 ms */
SysTick->VAL = 0U; /* clear current value */
SysTick->CTRL = SysTick_CTRL_CLKSOURCE_Msk | /* AHB clock (not /8) */
SysTick_CTRL_TICKINT_Msk | /* enable SysTick IRQ */
SysTick_CTRL_ENABLE_Msk; /* start counter */
}
/* SysTick exception handler — called every 1 ms */
volatile uint32_t g_tick_ms = 0;
void SysTick_Handler(void)
{
g_tick_ms++;
}
/* Blocking delay using SysTick */
void delay_ms(uint32_t ms)
{
uint32_t start = g_tick_ms;
while ((g_tick_ms - start) < ms);
}
/* Non-blocking elapsed time */
uint32_t millis(void)
{
return g_tick_ms;
}
SysTick priority: SysTick shares its priority register (SHPR3[31:24]) with PendSV. In FreeRTOS, SysTick and PendSV are set to the lowest priority (0xFF) so they never preempt application ISRs. Configure priority before enabling the timer.
MPU — Memory Protection Unit
The MPU is an optional (but universally present on Cortex-M3/M4) hardware unit that enforces access permissions on memory regions. It generates a MemManage fault when code violates its rules, making it a critical safety tool for RTOS task isolation and stack overflow detection.
MPU capabilities on Cortex-M3/M4:
• Up to 8 independently configurable memory regions
• Each region: base address, size (32 bytes to 4 GB), access permissions, attributes (cacheable, bufferable, shareable)
• Per-region sub-region disable (8 equal sub-regions per region, each independently maskable)
• Separate privilege and unprivileged access permissions (read/write/execute per mode)
• Background region: optional default for privileged code when MPU is enabled
| Register | Address | Purpose |
|---|---|---|
| MPU_TYPE | 0xE000ED90 | Read-only — DREGION (number of regions, 8 on Cortex-M4), SEPARATE (0 = unified) |
| MPU_CTRL | 0xE000ED94 | ENABLE, HFNMIENA (enable during NMI/HardFault), PRIVDEFENA (background region for priv code) |
| MPU_RNR | 0xE000ED98 | Region Number Register — select which region (0–7) the next write targets |
| MPU_RBAR | 0xE000ED9C | Region Base Address — base address (aligned to region size) + VALID + REGION shortcut |
| MPU_RASR | 0xE000EDA0 | Region Attribute & Size — size, AP[2:0] (access permissions), XN, TEX, S, C, B, SRD (sub-region disable) |
MPU stack guard example
/* Configure region 7 as a no-access guard at the bottom of the stack
to detect stack overflow — causes MemManageFault before corruption */
void mpu_set_stack_guard(uint32_t stack_bottom)
{
/* Disable MPU before configuring */
MPU->CTRL = 0;
/* Region 7: 32-byte no-access region at stack_bottom */
MPU->RNR = 7U;
MPU->RBAR = (stack_bottom & 0xFFFFFFE0U); /* 32-byte aligned */
MPU->RASR = (0b100 << 1) | /* SIZE = 4 → 2^(4+1) = 32 bytes */
(0b000 << 24) | /* AP = 000 → no access (priv or unpriv) */
(1U << 28) | /* XN = 1 → no execute */
(1U << 0); /* ENABLE this region */
/* Re-enable MPU with background region for privileged code */
__DSB();
MPU->CTRL = MPU_CTRL_ENABLE_Msk | MPU_CTRL_PRIVDEFENA_Msk;
__ISB();
}
/* Now any PUSH that crosses into the guard zone fires MemManage_Handler
rather than silently corrupting data below the stack. */
FPU — Floating-Point Unit (Cortex-M4F)
The Cortex-M4F includes a single-precision IEEE 754 floating-point unit. It adds 32 single-precision registers (S0–S31, also accessible as 16 double-precision D0–D15) and the FPSCR (FP Status and Control Register). The FPU must be explicitly enabled in firmware before any FP instruction executes — it is disabled after reset to save power.
Enabling the FPU
/* FPU access is controlled through the Coprocessor Access Control Register (CPACR)
in the SCB — address 0xE000ED88
CP10 and CP11 must both be set to 0b11 (full access) */
void fpu_enable(void)
{
/* Set CP10 and CP11 full access */
SCB->CPACR |= (0xFU << 20); /* bits 23:20 → CP11 and CP10 = 11b each */
__DSB();
__ISB(); /* flush pipeline so next FP instruction sees the change */
}
/* GCC / STM32CubeIDE automatically calls SystemInit() which calls fpu_enable()
if compiled with -mfpu=fpv4-sp-d16 -mfloat-abi=hard */
FPU and exception stacking
When the FPU is active and an exception fires, the hardware uses lazy stacking — it reserves space for the FP context (S0–S15, FPSCR = 17 words = 68 bytes) on the stack but does not actually write the FP registers unless the handler uses an FP instruction. This reduces interrupt latency for ISRs that don’t use floating point.
/* Total exception frame with FPU lazy stacking: */
/* 8 integer regs (32 bytes) + 1 reserved word (4 bytes) + 17 FP words (68 bytes) = 104 bytes */
/* To check if lazy stacking allocated space: */
uint32_t lspact = (FPU->FPCCR >> 0) & 1U; /* LSPACT bit in FPCCR (0xE000EF34) */
if (lspact) {
/* FP context is allocated but not yet saved — first FP instruction in ISR
triggers the actual save automatically */
}
| FPU Register | Address | Purpose |
|---|---|---|
| FPCCR | 0xE000EF34 | FP Context Control — ASPEN (auto save), LSPEN (lazy save), LSPACT (context reserved) |
| FPCAR | 0xE000EF38 | FP Context Address — address of the lazy-allocated FP frame on the stack |
| FPDSCR | 0xE000EF3C | FP Default Status Control — default rounding mode and flush-to-zero for new threads |
| MVFR0–1 | 0xE000EF40 | Media & FP Feature Registers — identify FPU capabilities (read-only) |
Debug Infrastructure
The Cortex-M debug system (ARM CoreSight) provides non-invasive hardware tracing, breakpoints, watchpoints, and register inspection — all accessible through standard debug ports (SWD or JTAG) without modifying the firmware.
ITM
Instrumentation Trace Macrocell: Provides a low-overhead printf-style output channel via the SWO (Single-Wire Output) pin. You write a byte to ITM->PORT[0].u8 and it appears in the debugger’s SWO viewer. Typically 50–100× faster than UART-based logging with no timing disruption.
DWT
Data Watchpoint & Trace: Contains a 32-bit free-running cycle counter (DWT_CYCCNT) that counts every processor clock cycle. Useful for profiling — read before and after a code section and subtract. Also provides up to 4 hardware data watchpoints (trigger on read/write to a specific address) without code modification.
FPB
Flash Patch & Breakpoint: Provides up to 8 hardware instruction breakpoints and 2 literal patch points. Hardware breakpoints halt the processor without any software overhead, unlike software breakpoints (BKPT instruction) that consume a word in Flash.
Using DWT_CYCCNT for microsecond profiling
/* Enable DWT cycle counter — must enable DWT first via CoreDebug */
void dwt_enable(void)
{
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk; /* enable DWT */
DWT->CYCCNT = 0; /* reset counter */
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk; /* start counting */
}
/* Measure execution time of a function in microseconds */
uint32_t measure_us(void (*fn)(void))
{
uint32_t start = DWT->CYCCNT;
fn();
uint32_t end = DWT->CYCCNT;
/* Divide by cycles per microsecond (= CPU freq in MHz) */
return (end - start) / 100U; /* 100 MHz → divide by 100 */
}
/* ITM printf via SWO */
int itm_putchar(int c)
{
while (ITM->PORT[0].u32 == 0); /* wait for port ready */
ITM->PORT[0].u8 = (uint8_t)c;
return c;
}
System Exception Numbers and Priorities
Cortex-M defines a set of system exceptions with fixed exception numbers (1–15). They are processor-internal, unlike external IRQs (numbers 16+). Their priority registers are in SHPR1–3 of the SCB, not in NVIC_IPR.
| Exception # | Name | Priority | Purpose |
|---|---|---|---|
| 1 | Reset | –3 (fixed, highest) | Power-on or software reset entry point |
| 2 | NMI | –2 (fixed) | Non-maskable interrupt — cannot be disabled or preempted except by Reset |
| 3 | HardFault | –1 (fixed) | Catch-all fault — triggered when configurable faults are disabled or escalate |
| 4 | MemManage | Programmable | MPU access violation or XN region execute attempt |
| 5 | BusFault | Programmable | Precise or imprecise memory access error (bad address, AHB error) |
| 6 | UsageFault | Programmable | Undefined instruction, divide by zero (if enabled), unaligned access, invalid state |
| 11 | SVCall | Programmable | Supervisor Call — RTOS system call / privilege escalation entry point |
| 12 | Debug Monitor | Programmable | Software debug breakpoints when not in halt-mode debug |
| 14 | PendSV | Programmable | Pendable service — RTOS context switch trigger (set to lowest priority) |
| 15 | SysTick | Programmable | System timer overflow — RTOS scheduling tick |
| 16+ | IRQ0…IRQ239 | Programmable (NVIC) | External peripheral interrupts (USART, SPI, GPIO, etc.) |
Frequently Asked Questions
Are these peripherals accessible from user (unprivileged) code?
No. The Private Peripheral Bus (PPB) region (0xE0000000–0xE00FFFFF) is always privileged-access-only, regardless of the MPU configuration. Any unprivileged read or write to a PPB address triggers a MemManage or HardFault. RTOS tasks running unprivileged must use SVC calls to access NVIC, SCB, or SysTick through the kernel.
Can SysTick be used as a general-purpose timer, not just for RTOS?
Yes. SysTick is simply a 24-bit countdown timer with a periodic interrupt. In bare-metal code without an RTOS, it is commonly used for delay_ms() and millis() implementations. It is also free on chips that have no spare hardware timers, since it is always present in the Cortex-M core.
Does the MPU completely prevent all memory corruption?
No — the MPU only controls access permissions for defined regions. Code that does illegal pointer arithmetic within an allowed region, or integer overflows that stay within a valid region, are not caught. The MPU is most effective for catching stack overflows (with a guard region), isolating RTOS tasks from each other, and preventing user code from touching kernel memory.
What is the difference between PRIMASK and BASEPRI?
PRIMASK is a 1-bit register — setting it to 1 masks all interrupts below NMI and HardFault (effectively disabling all IRQs). BASEPRI is an 8-bit register — it masks all exceptions with priority numerically greater than or equal to the BASEPRI value, allowing high-priority ISRs to still fire. FreeRTOS uses BASEPRI to implement critical sections without disabling high-priority ISRs.
Is the FPU available on all Cortex-M4 chips?
Cortex-M4F (with FPU) is the standard configuration and is used by STM32F4, STM32F3, STM32L4, and most modern Cortex-M4 implementations. A minority of Cortex-M4 cores were produced without the FPU (Cortex-M4 without the “F” suffix). Always check the device’s part number — STM32F411 is a Cortex-M4F and includes the FPU.
NVIC SCB System Control Block SysTick Timer MPU Memory Protection Cortex-M4 FPU Debug ITM DWT System Exceptions CMSIS STM32F411 Free Embedded Course
© EmbeddedPathashala — Free Embedded Systems Course — ARM Cortex-M3/M4 Series

2 Comments