Cortex-M Register Model and Inline Assembly
Master the complete Cortex-M register bank, understand how function calls use PC and LR, learn Thread vs Handler modes and access levels, then write real ARM GCC inline assembly to read and write CPU registers directly from C.
What is a Programmer’s Model?
A programmer’s model is the view of the processor that your code sees. It defines: which registers exist, how many bits they are, what each one does, what execution modes the processor supports, and what access rules apply. You don’t need to know how transistors implement a register — but you must know the programmer’s model cold, because every piece of firmware you write operates within these rules.
The Cortex-M3 and Cortex-M4 share the same ARMv7-M programmer’s model. Code compiled for one runs unmodified on the other (unless M4-specific FPU or DSP instructions are used). This is by design — ARM defines the programmer’s model, and chip vendors implement it faithfully.
Cortex-M registers fall into two fundamentally different categories that determine how you access them:
R0–R15, PSR, PRIMASK, FAULTMASK, BASEPRI, CONTROL. These live inside the CPU core and have no memory address. You cannot read them with a pointer dereference. You must use assembly instructions (or intrinsic functions).
NVIC, MPU, SCB, SysTick, DWT, ITM, and all MCU peripherals (GPIO, USART, I2C, TIM…). These sit at fixed addresses in the 4 GB address space. You read/write them with volatile uint32_t * pointers in plain C.
The Complete Cortex-M Register Bank
The Cortex-M has 16 programmer-visible integer registers, each 32 bits wide, plus several special-purpose registers. Here is every register, what it does, and how the compiler uses it:
How Function Calls Work: PC and LR
Understanding the Program Counter (PC) and Link Register (LR) is fundamental — not just for writing firmware, but for understanding interrupts, context switches, and debugging stack traces. Let’s trace exactly what happens at the hardware level when one function calls another.
1. LR = PC + 4 (address of fun1_instruction_B)
2. PC = address of fun2 (CPU jumps there)
POP {PC} loads the saved LR value back into PC.
CPU resumes execution at fun1_instruction_B.
What Happens with Nested Calls
When fun1 calls fun2 which calls fun3, the LR gets overwritten by the second BL. The compiler automatically generates a PUSH at the start of fun2 to save LR to the stack before it is overwritten. On return, it POPs the saved value back into PC. Leaf functions (those that call nothing) never need to save LR, saving stack space and cycles.
/* C source */
void fun3(void) { /* leaf — no calls inside */ }
void fun2(void) {
fun3(); /* fun2 calls fun3 — LR must be saved */
}
void fun1(void) {
fun2(); /* fun1 calls fun2 — LR must be saved */
}
/* Generated ARM Thumb-2 assembly (arm-none-eabi-gcc -O1 -mthumb) */
fun3:
BX LR ; leaf: just return, no push needed
fun2:
PUSH {LR} ; save return address — fun2 calls fun3
BL fun3 ; LR = addr after BL; PC jumps to fun3
POP {PC} ; restore saved LR into PC → returns to fun1
fun1:
PUSH {LR} ; save return address — fun1 calls fun2
BL fun2 ; LR = addr after BL; PC jumps to fun2
POP {PC} ; restore saved LR into PC → returns to caller
Every nested function call uses stack space (to save LR, and any callee-saved registers R4–R11). Deeply recursive functions with no base case, or ISRs that call deeply-nested functions, can exhaust the stack. On Cortex-M, a stack overflow silently overwrites heap or BSS data — undefined behaviour follows. Production firmware always includes a stack watermark check or MPU guard region at the bottom of the stack.
The Program Status Register (xPSR)
The PSR is actually three overlapping registers accessed as a unified 32-bit word called xPSR. Each sub-register monitors different aspects of the processor state:
| Register | Full Name | Bits | What It Tracks |
|---|---|---|---|
| APSR | Application Program Status Register | 31:27 | N, Z, C, V, Q arithmetic flags — updated by data-processing instructions |
| IPSR | Interrupt Program Status Register | 8:0 | Current exception number — 0 in Thread mode, non-zero in Handler mode |
| EPSR | Execution Program Status Register | 26:25, 15:10, 24 | T (Thumb) bit must always be 1 on Cortex-M; IT block state bits |
The Arithmetic Flags in Practice
/* C code */
int a = 5, b = -5;
int result = a + b; /* result = 0 */
/* Generated assembly (conceptually): */
MOV R0, #5 ; R0 = 5
MOV R1, #-5 ; R1 = 0xFFFFFFFB
ADDS R2, R0, R1 ; R2 = 0, updates flags:
; N=0 (result not negative)
; Z=1 (result is zero) ← Z flag set
; C=1 (unsigned carry out)
; V=0 (no signed overflow)
/* Conditional branch uses these flags: */
BEQ label ; branch if Z=1 (i.e., if result == 0)
Operating Modes: Thread Mode and Handler Mode
The Cortex-M has two operating modes. Understanding which mode you are in determines which stack pointer is used and whether certain instructions are allowed.
- This is the normal mode — where your application code runs
- The CPU enters Thread mode after reset
- The CPU returns to Thread mode when an exception (interrupt) handler completes
- Can use either MSP or PSP (selected by CONTROL register bit 1)
- Can be either Privileged or Unprivileged (selected by CONTROL register bit 0)
- IPSR = 0 in Thread mode (no exception is active)
- All application tasks in FreeRTOS run in Thread mode
- This is the exception mode — entered when any interrupt or exception fires
- CPU automatically enters Handler mode on any exception
- Interrupt service routines (ISRs) run in Handler mode
- Always uses MSP (Main Stack Pointer) — never PSP
- Always Privileged — Handler mode is always fully privileged
- IPSR ≠ 0 (shows which exception number is active)
- CPU automatically returns to Thread mode at end of ISR using EXC_RETURN
(Privileged) MSP active
Normal application code
(Always Privileged) MSP always used
ISR runs here
(Unprivileged) PSP used (with RTOS)
User tasks run here
Access Levels: Privileged vs Unprivileged
The access level is an independent dimension from the operating mode. It controls what the currently-running code is allowed to do:
- Full access to all instructions and registers
- Can read/write PRIMASK, FAULTMASK, BASEPRI
- Can modify the CONTROL register
- Can access all MPU-protected regions
- Can use MSR and MRS instructions on special registers
- Handler mode is always privileged
- Thread mode is privileged after reset (CONTROL[0] = 0)
- Cannot write PRIMASK, FAULTMASK, BASEPRI, CONTROL
- Cannot use MSR on most special registers
- MPU can prevent access to privileged memory regions
- Cannot modify system control registers (SCB, NVIC directly)
- To regain privilege: must raise a software exception (SVC)
- Used by user tasks in RTOS (FreeRTOS MPU port)
- Entered by setting CONTROL[0] = 1 in Thread mode
The Two Stack Pointers: MSP and PSP
R13 appears as a single register but is actually banked — there are two physical 32-bit registers behind R13, and only one is visible at a time:
| Register | Full Name | Used By | Selected By | Typical Use |
|---|---|---|---|---|
| MSP | Main Stack Pointer | Handler mode (always) + Thread mode (if CONTROL[1]=0) | CONTROL[1] = 0 | Kernel stack, ISR stack, bare-metal applications |
| PSP | Process Stack Pointer | Thread mode only (when CONTROL[1]=1) | CONTROL[1] = 1 | User-task stacks in RTOS (each task has its own PSP) |
In a bare-metal application, you only need MSP. In an RTOS like FreeRTOS, the kernel configures each task to use PSP. When a context switch occurs, the scheduler saves the PSP value for the outgoing task and loads the PSP value for the incoming task. The MSP is reserved for the RTOS kernel and interrupt handlers, ensuring that ISR stack usage never interferes with user task stacks.
application + ISRs
FreeRTOS kernel uses MSP
The CONTROL Register — All 3 Bits Explained
| Bit | Name | Reset Value | Meaning When 0 | Meaning When 1 |
|---|---|---|---|---|
| Bit 0 | nPRIV | 0 | Thread mode is Privileged | Thread mode is Unprivileged |
| Bit 1 | SPSEL | 0 | Thread mode uses MSP | Thread mode uses PSP |
| Bit 2 | FPCA | 0 | FP context not active (Cortex-M4 only) | FP context is active — FPU state saved on exception |
#include <stdint.h>
/* Read the CONTROL register into a C variable */
uint32_t control_reg;
__asm volatile ("MRS %0, CONTROL"
: "=r"(control_reg) /* output: %0 = any register */
: /* no inputs */
);
/* Set bit 1 (SPSEL): Thread mode now uses PSP */
control_reg |= (1U << 1);
__asm volatile ("MSR CONTROL, %0"
: /* no outputs */
: "r"(control_reg) /* input: %0 = any register */
);
/* Required: ISB after MSR CONTROL to ensure pipeline sees the change */
__asm volatile ("ISB");
/* Set PSP to a specific stack address before switching to PSP */
uint32_t task_stack_top = 0x20004000U;
__asm volatile ("MSR PSP, %0"
:
: "r"(task_stack_top)
);
Non-Memory-Mapped vs Memory-Mapped Registers
This distinction directly determines how you write firmware. Getting it wrong (trying to use a pointer to access R0, or using inline assembly to access GPIOA_ODR) produces wrong code or won’t compile.
Or GCC intrinsics: __get_PRIMASK() etc.
+ All MCU peripherals: GPIO, RCC, USART, SPI, I2C, TIM, ADC, USB, CAN…
ARM GCC Inline Assembly — Complete Guide
Inline assembly lets you embed ARM Thumb-2 instructions directly inside a C function. You need this to access non-memory-mapped registers (CONTROL, PRIMASK, MSP, PSP, xPSR) and for precise hardware-level operations the C compiler cannot express.
Basic Syntax Form
The ARM GCC inline assembly uses the extended AT&T inline assembly format. The full four-section form is:
: output operand list /* Section 2: outputs */
: input operand list /* Section 3: inputs */
: “clobber list” /* Section 4: side-effects */
);
Prevents GCC from eliminating or reordering this asm block. Always use volatile for register reads/writes — the compiler has no way to know if removing the instruction would change observable hardware state.
Tells GCC: the assembly writes a result into this C variable. GCC allocates a register for it. “=” means write-only.
Tells GCC: load this C variable into a register before the asm runs. Use %0, %1… in the code string to reference them.
Tells GCC what the assembly modifies beyond the declared outputs. “cc” = flags modified; “memory” = memory barrier; named registers = those registers are overwritten.
Operand Constraints — Quick Reference
Constraints tell GCC what kind of register or storage location to use for each operand:
Practical Example 1 — Simple Register Operations
#include <stdint.h>
void asm_basics_demo(void)
{
uint32_t data = 0x1234ABCDU;
uint32_t result;
/* Move C variable 'data' to any ARM register (%0 = input operand 0).
The "r" constraint lets GCC pick any general-purpose register.
GCC will emit: MOV Rx, #0x1234ABCD (or LDR if needed) before this asm. */
__asm volatile (
"MOV %0, %1\n\t" /* move data into result's register */
: "=r"(result) /* output: result <- register %0 */
: "r"(data) /* input: data -> register %1 */
);
/* result now == 0x1234ABCD */
/* Perform ADD in assembly: result = data + 10 */
__asm volatile (
"ADD %0, %1, #10\n\t"
: "=r"(result)
: "r"(data)
);
/* result now == 0x1234ABD7 */
}
Practical Example 2 — Memory Load/Store (LDR, STR)
/* Assembly: LDR R0,[R1] — load 32-bit word from address in R1 into R0
LDR R1,[R2] — load from address in R2 into R1
ADD R1, R0 — R1 = R1 + R0
STR R1,[R3] — store R1 to address in R3 */
uint32_t a = 10, b = 20, c = 0;
/* Individual asm statements — one instruction each */
__asm volatile ("LDR R0, [%0]" : : "r"(&a) : "R0"); /* R0 = a */
__asm volatile ("LDR R1, [%0]" : : "r"(&b) : "R1"); /* R1 = b */
__asm volatile ("ADD R1, R0" : : : "R1"); /* R1 = R1+R0 */
__asm volatile ("STR R1, [%0]" : : "r"(&c) : "memory"); /* c = R1 */
/* Better: all four instructions in one asm block with named operands */
__asm volatile (
"LDR R0, [%[src_a]]\n\t" /* named input operand [src_a] = &a */
"LDR R1, [%[src_b]]\n\t"
"ADD R1, R0\n\t"
"STR R1, [%[dst_c]]\n\t"
: /* no output operands */
: [src_a] "r" (&a), /* named input: &a */
[src_b] "r" (&b),
[dst_c] "r" (&c)
: "R0", "R1", "memory" /* clobbers: R0, R1 overwritten */
);
/* c is now 30 (10 + 20) */
Practical Example 3 — Reading Special Registers
#include <stdint.h>
/* MRS (Move from Special Register to Register) reads special registers.
MSR (Move from Register to Special Register) writes them.
Both MRS and MSR are privileged instructions — only work in privileged mode. */
static inline uint32_t get_primask(void)
{
uint32_t val;
__asm volatile ("MRS %0, PRIMASK" : "=r"(val));
return val; /* bit 0 = 1 means interrupts disabled */
}
static inline void set_primask(uint32_t val)
{
__asm volatile ("MSR PRIMASK, %0" : : "r"(val) : "memory");
}
static inline uint32_t get_control(void)
{
uint32_t val;
__asm volatile ("MRS %0, CONTROL" : "=r"(val));
return val;
}
static inline uint32_t get_msp(void)
{
uint32_t val;
__asm volatile ("MRS %0, MSP" : "=r"(val));
return val;
}
static inline uint32_t get_psp(void)
{
uint32_t val;
__asm volatile ("MRS %0, PSP" : "=r"(val));
return val;
}
void demo_special_regs(void)
{
uint32_t ctrl = get_control();
printf("CONTROL = 0x%08lX\n", ctrl);
printf(" nPRIV (bit0) = %lu %s\n", ctrl & 1,
(ctrl & 1) ? "Unprivileged" : "Privileged");
printf(" SPSEL (bit1) = %lu %s\n", (ctrl >> 1) & 1,
(ctrl >> 1) & 1 ? "PSP active" : "MSP active");
printf("MSP = 0x%08lX\n", get_msp());
printf("PSP = 0x%08lX\n", get_psp());
printf("PRIMASK = %lu (0=enabled, 1=disabled)\n", get_primask());
}
CONTROL = 0x00000000 nPRIV (bit0) = 0 Privileged SPSEL (bit1) = 0 MSP active MSP = 0x20020000 PSP = 0x00000000 PRIMASK = 0 (0=enabled, 1=disabled)
After reset, the processor is in Thread mode, Privileged, using MSP. PSP = 0 because it has not been initialised. PRIMASK = 0 means all configured interrupts are enabled.
Practical Example 4 — Critical Section (Disable/Enable Interrupts)
/* Disable all interrupts — enter critical section */
static inline void disable_interrupts(void)
{
__asm volatile ("CPSID I" : : : "memory");
/* CPSID I: Change Processor State, Interrupt Disable
Sets PRIMASK = 1. All IRQs except NMI and HardFault are masked.
"memory" clobber: prevents compiler moving memory ops across this barrier. */
}
/* Enable all interrupts — exit critical section */
static inline void enable_interrupts(void)
{
__asm volatile ("CPSIE I" : : : "memory");
/* CPSIE I: Change Processor State, Interrupt Enable — clears PRIMASK */
}
/* Safer pattern: save and restore PRIMASK state (handles nesting) */
static inline uint32_t enter_critical(void)
{
uint32_t saved;
__asm volatile (
"MRS %0, PRIMASK\n\t" /* save current interrupt state */
"CPSID I\n\t" /* disable interrupts */
: "=r"(saved) : : "memory"
);
return saved;
}
static inline void exit_critical(uint32_t saved)
{
__asm volatile (
"MSR PRIMASK, %0\n\t" /* restore saved state */
: : "r"(saved) : "memory"
);
}
/* Usage in application code */
static volatile uint32_t shared_counter = 0;
void increment_counter(void)
{
uint32_t irq_state = enter_critical();
shared_counter++; /* atomic — no ISR can interrupt this */
exit_critical(irq_state);
}
ARM Calling Convention (AAPCS) — What the Compiler Assumes
When you write inline assembly that touches registers, you must respect the ARM Procedure Call Standard (AAPCS). Violating it corrupts the register state seen by surrounding C code in unpredictable ways.
| Register(s) | Role | Who Saves It? | Your Inline Asm Rule |
|---|---|---|---|
| R0–R3 | Arguments / return value / scratch | Caller-saved — callee may destroy | Declare in clobber list if you modify them, or use output/input operands |
| R12 (IP) | Intra-procedure scratch | Caller-saved — callee may destroy | Declare in clobber list if you modify it |
| R4–R11 | Callee-saved (preserved across calls) | Callee-saved — if modified, must restore | If your asm modifies R4–R11, PUSH them first and POP them after |
| R13 (SP) | Stack pointer — must be 4-byte aligned | Must always be valid | Never destroy SP unless you restore it exactly. Must be 8-byte aligned at function calls. |
| R14 (LR) | Link register / scratch in leaf functions | Caller-saved — overwritten by BL | If your asm calls BL, LR is overwritten. Save LR first if needed. |
| APSR flags | Condition flags | Not preserved across calls | Add “cc” to clobber list if your asm sets N,Z,C,V flags |
Ready-to-Use Inline Assembly Patterns
#include <stdint.h>
/* ── No-Operation delay ── */
static inline void nop(void)
{
__asm volatile ("NOP");
}
/* ── Data Memory Barrier — ensures all memory accesses complete ── */
static inline void dmb(void)
{
__asm volatile ("DMB" : : : "memory");
}
/* ── Data Synchronisation Barrier — stronger than DMB ── */
static inline void dsb(void)
{
__asm volatile ("DSB" : : : "memory");
}
/* ── Instruction Synchronisation Barrier — flush pipeline ── */
static inline void isb(void)
{
__asm volatile ("ISB" : : : "memory");
}
/* ── Wait For Interrupt — enter sleep until next interrupt ── */
static inline void wfi(void)
{
__asm volatile ("WFI");
}
/* ── Wait For Event — used with SEV for multi-core sync ── */
static inline void wfe(void)
{
__asm volatile ("WFE");
}
/* ── Software Breakpoint — halts the debugger ── */
static inline void bkpt(void)
{
__asm volatile ("BKPT #0");
}
/* ── Count Leading Zeros (for fast log2, bit-weight) ── */
static inline uint32_t clz(uint32_t val)
{
uint32_t result;
__asm volatile ("CLZ %0, %1" : "=r"(result) : "r"(val));
return result; /* 0 if val has bit31 set; 32 if val == 0 */
}
/* ── Bit Reverse (for FFT butterfly operations, CRC) ── */
static inline uint32_t rbit(uint32_t val)
{
uint32_t result;
__asm volatile ("RBIT %0, %1" : "=r"(result) : "r"(val));
return result;
}
/* ── Byte-Reverse (for endian conversion) ── */
static inline uint32_t rev(uint32_t val)
{
uint32_t result;
__asm volatile ("REV %0, %1" : "=r"(result) : "r"(val));
return result; /* 0x12345678 → 0x78563412 */
}
Frequently Asked Questions
Lecture 5 covers the ARM Cortex-M 4 GB address space in detail — how it is partitioned into code, SRAM, peripheral, and system regions, how memory-mapped registers are addressed, and how to read a memory map from a chip datasheet and Reference Manual.

2 Comments