Cortex-M Register Model and Inline Assembly-Embedded C Programming Training Online

Cortex-M Register Model and Inline Assembly

Lecture 04 — Programmer’s Model

Cortex-M Register Model and Inline Assembly

Master the complete Cortex-M register bank, understand how function calls use PC and LR, learn Thread vs Handler modes and access levels, then write real ARM GCC inline assembly to read and write CPU registers directly from C.

Intermediate Estimated read: 35 min Target: All Cortex-M3 / M4 chips

What is a Programmer’s Model?

A programmer’s model is the view of the processor that your code sees. It defines: which registers exist, how many bits they are, what each one does, what execution modes the processor supports, and what access rules apply. You don’t need to know how transistors implement a register — but you must know the programmer’s model cold, because every piece of firmware you write operates within these rules.

The Cortex-M3 and Cortex-M4 share the same ARMv7-M programmer’s model. Code compiled for one runs unmodified on the other (unless M4-specific FPU or DSP instructions are used). This is by design — ARM defines the programmer’s model, and chip vendors implement it faithfully.

The Two Categories of Registers

Cortex-M registers fall into two fundamentally different categories that determine how you access them:

Non-memory-mapped registers
R0–R15, PSR, PRIMASK, FAULTMASK, BASEPRI, CONTROL. These live inside the CPU core and have no memory address. You cannot read them with a pointer dereference. You must use assembly instructions (or intrinsic functions).
Memory-mapped registers
NVIC, MPU, SCB, SysTick, DWT, ITM, and all MCU peripherals (GPIO, USART, I2C, TIM…). These sit at fixed addresses in the 4 GB address space. You read/write them with volatile uint32_t * pointers in plain C.

The Complete Cortex-M Register Bank

The Cortex-M has 16 programmer-visible integer registers, each 32 bits wide, plus several special-purpose registers. Here is every register, what it does, and how the compiler uses it:

Cortex-M3 / M4 Register Bank (ARMv7-M)
General Purpose Registers
R0
Argument 1 / return value / scratch
R1
Argument 2 / return value (high) / scratch
R2
Argument 3 / scratch
R3
Argument 4 / scratch
R4
Callee-saved — caller expects unchanged
R5
Callee-saved
R6
Callee-saved
R7
Callee-saved
R8
Callee-saved (high register)
R9
Callee-saved / platform register
R10
Callee-saved
R11
Callee-saved / frame pointer (optional)
R12
Intra-procedure scratch (IP) — caller-saved
Special Purpose Registers
R13 / SP
Stack Pointer — banked as MSP & PSP
R14 / LR
Link Register — stores return address on call
R15 / PC
Program Counter — address of next instruction
Program Status Registers
APSR
Application PSR — N,Z,C,V,Q flags
IPSR
Interrupt PSR — current exception number
EPSR
Execution PSR — Thumb state (T bit), IT bits
Exception Mask & Control
PRIMASK
1 bit — disables all exceptions except NMI & HardFault
FAULTMASK
1 bit — disables all exceptions including HardFault
BASEPRI
8 bits — masks interrupts at or below a priority level
CONTROL
3 bits — selects stack, privilege level, FPU context
R0–R3 and R12 are caller-saved (the called function may freely destroy them). R4–R11 are callee-saved (the called function must preserve them — push on entry, pop on exit). This rule is the ARM Procedure Call Standard (AAPCS).

How Function Calls Work: PC and LR

Understanding the Program Counter (PC) and Link Register (LR) is fundamental — not just for writing firmware, but for understanding interrupts, context switches, and debugging stack traces. Let’s trace exactly what happens at the hardware level when one function calls another.

Function Call Mechanism — PC and LR Step by Step
Caller: fun1()
fun1_instruction_A; ← PC points here
BL fun2 ← branch & link
fun1_instruction_B; ← LR stored here
fun1_instruction_C;
What BL does (Branch with Link):
1. LR = PC + 4 (address of fun1_instruction_B)
2. PC = address of fun2 (CPU jumps there)
PC jumps to fun2
→
PC = LR on return
←
Callee: fun2()
PUSH {R4, LR} ← save LR on stack
fun2_instruction_1;
fun2_instruction_2;
POP {R4, PC} ← PC = saved LR
What return does:
POP {PC} loads the saved LR value back into PC.
CPU resumes execution at fun1_instruction_B.
The hardware-level function call costs exactly 2 instructions: BL (or BLX) to call, POP {PC} to return. The Link Register (R14) is the key — it temporarily holds the return address between the call and the return.

What Happens with Nested Calls

When fun1 calls fun2 which calls fun3, the LR gets overwritten by the second BL. The compiler automatically generates a PUSH at the start of fun2 to save LR to the stack before it is overwritten. On return, it POPs the saved value back into PC. Leaf functions (those that call nothing) never need to save LR, saving stack space and cycles.

Disassembly showing LR save/restore in nested calls
/* C source */
void fun3(void) { /* leaf — no calls inside */ }

void fun2(void) {
    fun3();  /* fun2 calls fun3 — LR must be saved */
}

void fun1(void) {
    fun2();  /* fun1 calls fun2 — LR must be saved */
}

/* Generated ARM Thumb-2 assembly (arm-none-eabi-gcc -O1 -mthumb) */

fun3:
    BX   LR              ; leaf: just return, no push needed

fun2:
    PUSH {LR}            ; save return address — fun2 calls fun3
    BL   fun3            ; LR = addr after BL; PC jumps to fun3
    POP  {PC}            ; restore saved LR into PC → returns to fun1

fun1:
    PUSH {LR}            ; save return address — fun1 calls fun2
    BL   fun2            ; LR = addr after BL; PC jumps to fun2
    POP  {PC}            ; restore saved LR into PC → returns to caller
Why Stack Overflow Crashes Are Predictable

Every nested function call uses stack space (to save LR, and any callee-saved registers R4–R11). Deeply recursive functions with no base case, or ISRs that call deeply-nested functions, can exhaust the stack. On Cortex-M, a stack overflow silently overwrites heap or BSS data — undefined behaviour follows. Production firmware always includes a stack watermark check or MPU guard region at the bottom of the stack.

The Program Status Register (xPSR)

The PSR is actually three overlapping registers accessed as a unified 32-bit word called xPSR. Each sub-register monitors different aspects of the processor state:

RegisterFull NameBitsWhat It Tracks
APSR Application Program Status Register 31:27 N, Z, C, V, Q arithmetic flags — updated by data-processing instructions
IPSR Interrupt Program Status Register 8:0 Current exception number — 0 in Thread mode, non-zero in Handler mode
EPSR Execution Program Status Register 26:25, 15:10, 24 T (Thumb) bit must always be 1 on Cortex-M; IT block state bits
xPSR 32-bit Layout — All Sub-registers Combined
31
N
30
Z
29
C
28
V
27
Q
26:25
IT[1:0]
24
T
23:20
Reserved
19:16
GE[3:0]
15:10
IT[7:2]
9:8
Res
8:0
ISR_NUMBER
N — Negative flag (result was negative) Z — Zero flag (result was zero) C — Carry flag (unsigned overflow) V — Overflow flag (signed overflow) Q — Saturation flag (DSP operations) T — Thumb state — must always be 1 on Cortex-M GE — SIMD greater-equal flags (M4 DSP) ISR_NUMBER — current exception number
The T bit is always 1 on Cortex-M. If it ever becomes 0 (e.g., by branching to an even address), the CPU raises a HardFault. This is why function pointers in Thumb code always have their LSB set to 1.

The Arithmetic Flags in Practice

How flags work — ADDS updates N, Z, C, V automatically
/* C code */
int a = 5, b = -5;
int result = a + b;  /* result = 0 */

/* Generated assembly (conceptually): */
MOV  R0, #5        ; R0 = 5
MOV  R1, #-5       ; R1 = 0xFFFFFFFB
ADDS R2, R0, R1    ; R2 = 0, updates flags:
                   ;   N=0 (result not negative)
                   ;   Z=1 (result is zero)  ← Z flag set
                   ;   C=1 (unsigned carry out)
                   ;   V=0 (no signed overflow)

/* Conditional branch uses these flags: */
BEQ  label         ; branch if Z=1 (i.e., if result == 0)

Operating Modes: Thread Mode and Handler Mode

The Cortex-M has two operating modes. Understanding which mode you are in determines which stack pointer is used and whether certain instructions are allowed.

Thread Mode
  • This is the normal mode — where your application code runs
  • The CPU enters Thread mode after reset
  • The CPU returns to Thread mode when an exception (interrupt) handler completes
  • Can use either MSP or PSP (selected by CONTROL register bit 1)
  • Can be either Privileged or Unprivileged (selected by CONTROL register bit 0)
  • IPSR = 0 in Thread mode (no exception is active)
  • All application tasks in FreeRTOS run in Thread mode
Handler Mode
  • This is the exception mode — entered when any interrupt or exception fires
  • CPU automatically enters Handler mode on any exception
  • Interrupt service routines (ISRs) run in Handler mode
  • Always uses MSP (Main Stack Pointer) — never PSP
  • Always Privileged — Handler mode is always fully privileged
  • IPSR ≠ 0 (shows which exception number is active)
  • CPU automatically returns to Thread mode at end of ISR using EXC_RETURN
Mode Transition Diagram
Reset CPU starts here
→
Thread Mode
(Privileged) MSP active
Normal application code
Exception occurs ↓
↑ EXC_RETURN
Handler Mode
(Always Privileged) MSP always used
ISR runs here
CONTROL[0] = 1 ↓
↑ Cannot go back easily
Thread Mode
(Unprivileged) PSP used (with RTOS)
User tasks run here
Once CONTROL[0] is set to 1 (unprivileged), Thread mode code cannot set it back to 0 — only Handler mode (privileged) can do that. This is exactly the security boundary used by RTOS kernels.

Access Levels: Privileged vs Unprivileged

The access level is an independent dimension from the operating mode. It controls what the currently-running code is allowed to do:

Privileged Level
  • Full access to all instructions and registers
  • Can read/write PRIMASK, FAULTMASK, BASEPRI
  • Can modify the CONTROL register
  • Can access all MPU-protected regions
  • Can use MSR and MRS instructions on special registers
  • Handler mode is always privileged
  • Thread mode is privileged after reset (CONTROL[0] = 0)
Unprivileged Level
  • Cannot write PRIMASK, FAULTMASK, BASEPRI, CONTROL
  • Cannot use MSR on most special registers
  • MPU can prevent access to privileged memory regions
  • Cannot modify system control registers (SCB, NVIC directly)
  • To regain privilege: must raise a software exception (SVC)
  • Used by user tasks in RTOS (FreeRTOS MPU port)
  • Entered by setting CONTROL[0] = 1 in Thread mode

The Two Stack Pointers: MSP and PSP

R13 appears as a single register but is actually banked — there are two physical 32-bit registers behind R13, and only one is visible at a time:

RegisterFull NameUsed BySelected ByTypical Use
MSP Main Stack Pointer Handler mode (always) + Thread mode (if CONTROL[1]=0) CONTROL[1] = 0 Kernel stack, ISR stack, bare-metal applications
PSP Process Stack Pointer Thread mode only (when CONTROL[1]=1) CONTROL[1] = 1 User-task stacks in RTOS (each task has its own PSP)

In a bare-metal application, you only need MSP. In an RTOS like FreeRTOS, the kernel configures each task to use PSP. When a context switch occurs, the scheduler saves the PSP value for the outgoing task and loads the PSP value for the incoming task. The MSP is reserved for the RTOS kernel and interrupt handlers, ensuring that ISR stack usage never interferes with user task stacks.

MSP vs PSP in a Bare-Metal vs RTOS System
Bare-Metal (no RTOS)
MSP → main stack
Used for everything:
application + ISRs
PSP not used
CONTROL[1] = 0
RTOS System (e.g. FreeRTOS)
MSP → kernel + ISR stack
All interrupt handlers use MSP
FreeRTOS kernel uses MSP
PSP → each task’s own stack
Task A: PSP = 0x20004000
Task B: PSP = 0x20003000
On context switch: scheduler saves PSP for Task A, loads PSP for Task B
The PSP/MSP split means a runaway user task that overflows its stack cannot corrupt the kernel’s MSP stack. This is a fundamental safety guarantee in RTOS-based systems.

The CONTROL Register — All 3 Bits Explained

BitNameReset ValueMeaning When 0Meaning When 1
Bit 0 nPRIV 0 Thread mode is Privileged Thread mode is Unprivileged
Bit 1 SPSEL 0 Thread mode uses MSP Thread mode uses PSP
Bit 2 FPCA 0 FP context not active (Cortex-M4 only) FP context is active — FPU state saved on exception
Reading and writing the CONTROL register via inline assembly
#include <stdint.h>

/* Read the CONTROL register into a C variable */
uint32_t control_reg;
__asm volatile ("MRS %0, CONTROL"
               : "=r"(control_reg)   /* output: %0 = any register */
               :                       /* no inputs */
               );

/* Set bit 1 (SPSEL): Thread mode now uses PSP */
control_reg |= (1U << 1);
__asm volatile ("MSR CONTROL, %0"
               :                       /* no outputs */
               : "r"(control_reg)    /* input: %0 = any register */
               );
/* Required: ISB after MSR CONTROL to ensure pipeline sees the change */
__asm volatile ("ISB");

/* Set PSP to a specific stack address before switching to PSP */
uint32_t task_stack_top = 0x20004000U;
__asm volatile ("MSR PSP, %0"
               :
               : "r"(task_stack_top)
               );

Non-Memory-Mapped vs Memory-Mapped Registers

This distinction directly determines how you write firmware. Getting it wrong (trying to use a pointer to access R0, or using inline assembly to access GPIOA_ODR) produces wrong code or won’t compile.

Register Access Method Summary
Non-Memory-Mapped Registers
Who they are:
R0–R15, APSR, IPSR, EPSR, PRIMASK, FAULTMASK, BASEPRI, CONTROL, MSP, PSP
Key property:
No memory address — they are physical flip-flops inside the CPU datapath
How to access:
Only via assembly: MOV, MRS, MSR, PUSH, POP
Or GCC intrinsics: __get_PRIMASK() etc.
Memory-Mapped Registers
Who they are:
NVIC, MPU, SCB, SysTick, DWT, ITM (0xE0000000–0xE00FFFFF)
+ All MCU peripherals: GPIO, RCC, USART, SPI, I2C, TIM, ADC, USB, CAN…
Key property:
Each has a fixed address in the 4 GB memory map (documented in Reference Manual)
How to access:
Plain C with volatile pointer: *((volatile uint32_t*)0x40020014U) |= (1<<5);
The golden rule: if it’s inside the CPU core, use assembly or intrinsics. If it has an address in the memory map, use a C pointer.

ARM GCC Inline Assembly — Complete Guide

Inline assembly lets you embed ARM Thumb-2 instructions directly inside a C function. You need this to access non-memory-mapped registers (CONTROL, PRIMASK, MSP, PSP, xPSR) and for precise hardware-level operations the C compiler cannot express.

Basic Syntax Form

The ARM GCC inline assembly uses the extended AT&T inline assembly format. The full four-section form is:

Inline Assembly Anatomy
__asm volatile ( “assembly code string” /* Section 1: instruction(s) */
    : output operand list /* Section 2: outputs */
    : input operand list /* Section 3: inputs */
    : “clobber list” /* Section 4: side-effects */
);
volatile
Prevents GCC from eliminating or reordering this asm block. Always use volatile for register reads/writes — the compiler has no way to know if removing the instruction would change observable hardware state.
Output operands — format: “=r”(var)
Tells GCC: the assembly writes a result into this C variable. GCC allocates a register for it. “=” means write-only.
Input operands — format: “r”(var)
Tells GCC: load this C variable into a register before the asm runs. Use %0, %1… in the code string to reference them.
Clobber list — e.g.: “r0″,”memory”,”cc”
Tells GCC what the assembly modifies beyond the declared outputs. “cc” = flags modified; “memory” = memory barrier; named registers = those registers are overwritten.

Operand Constraints — Quick Reference

Constraints tell GCC what kind of register or storage location to use for each operand:

“r”
Any general-purpose register (R0–R12). Most common constraint.
“l”
Low register only (R0–R7). Required for some 16-bit Thumb instructions.
“h”
High register only (R8–R15). Rarely needed.
“I”
Immediate constant 0–255. For ADD, SUB with immediate operand.
“i”
Any immediate integer constant.
“m”
Memory location. Used when the asm accesses memory directly.
“=r”
Write-only output into any register. The = prefix means output.
“+r”
Read-write register. The asm both reads and writes this register.
“&r”
Early-clobber output — GCC must not use same reg as any input.
“cc”
Clobber: asm modifies condition flags (APSR N,Z,C,V).
“memory”
Clobber: full memory barrier — GCC flushes all cached values.

Practical Example 1 — Simple Register Operations

Moving C variable data into R0 and back — the basics
#include <stdint.h>

void asm_basics_demo(void)
{
    uint32_t data = 0x1234ABCDU;
    uint32_t result;

    /* Move C variable 'data' to any ARM register (%0 = input operand 0).
       The "r" constraint lets GCC pick any general-purpose register.
       GCC will emit: MOV Rx, #0x1234ABCD (or LDR if needed) before this asm. */
    __asm volatile (
        "MOV %0, %1\n\t"     /* move data into result's register */
        : "=r"(result)        /* output: result <- register %0   */
        : "r"(data)           /* input:  data  -> register %1    */
    );
    /* result now == 0x1234ABCD */

    /* Perform ADD in assembly: result = data + 10 */
    __asm volatile (
        "ADD %0, %1, #10\n\t"
        : "=r"(result)
        : "r"(data)
    );
    /* result now == 0x1234ABD7 */
}

Practical Example 2 — Memory Load/Store (LDR, STR)

The add-via-memory example from the slides, fully annotated
/*  Assembly:    LDR R0,[R1]    — load 32-bit word from address in R1 into R0
                LDR R1,[R2]    — load from address in R2 into R1
                ADD R1, R0     — R1 = R1 + R0
                STR R1,[R3]    — store R1 to address in R3              */

uint32_t a = 10, b = 20, c = 0;

/* Individual asm statements — one instruction each */
__asm volatile ("LDR R0, [%0]" : : "r"(&a) : "R0");  /* R0 = a */
__asm volatile ("LDR R1, [%0]" : : "r"(&b) : "R1");  /* R1 = b */
__asm volatile ("ADD R1, R0"   : : : "R1");          /* R1 = R1+R0 */
__asm volatile ("STR R1, [%0]" : : "r"(&c) : "memory"); /* c = R1 */

/* Better: all four instructions in one asm block with named operands */
__asm volatile (
    "LDR R0, [%[src_a]]\n\t"   /* named input operand [src_a] = &a */
    "LDR R1, [%[src_b]]\n\t"
    "ADD R1, R0\n\t"
    "STR R1, [%[dst_c]]\n\t"
    :                              /* no output operands            */
    : [src_a] "r" (&a),           /* named input: &a               */
      [src_b] "r" (&b),
      [dst_c] "r" (&c)
    : "R0", "R1", "memory"        /* clobbers: R0, R1 overwritten  */
);
/* c is now 30 (10 + 20) */

Practical Example 3 — Reading Special Registers

Reading PRIMASK, BASEPRI, CONTROL, MSP, PSP using MRS
#include <stdint.h>

/* MRS (Move from Special Register to Register) reads special registers.
   MSR (Move from Register to Special Register) writes them.
   Both MRS and MSR are privileged instructions — only work in privileged mode. */

static inline uint32_t get_primask(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, PRIMASK" : "=r"(val));
    return val;  /* bit 0 = 1 means interrupts disabled */
}

static inline void set_primask(uint32_t val)
{
    __asm volatile ("MSR PRIMASK, %0" : : "r"(val) : "memory");
}

static inline uint32_t get_control(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, CONTROL" : "=r"(val));
    return val;
}

static inline uint32_t get_msp(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, MSP" : "=r"(val));
    return val;
}

static inline uint32_t get_psp(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, PSP" : "=r"(val));
    return val;
}

void demo_special_regs(void)
{
    uint32_t ctrl = get_control();
    printf("CONTROL = 0x%08lX\n", ctrl);
    printf("  nPRIV (bit0) = %lu  %s\n", ctrl & 1,
           (ctrl & 1) ? "Unprivileged" : "Privileged");
    printf("  SPSEL (bit1) = %lu  %s\n", (ctrl >> 1) & 1,
           (ctrl >> 1) & 1 ? "PSP active" : "MSP active");

    printf("MSP = 0x%08lX\n", get_msp());
    printf("PSP = 0x%08lX\n", get_psp());
    printf("PRIMASK = %lu  (0=enabled, 1=disabled)\n", get_primask());
}
Expected ITM Output (running on STM32F411 after reset)
CONTROL = 0x00000000
  nPRIV (bit0) = 0  Privileged
  SPSEL (bit1) = 0  MSP active
MSP = 0x20020000
PSP = 0x00000000
PRIMASK = 0  (0=enabled, 1=disabled)

After reset, the processor is in Thread mode, Privileged, using MSP. PSP = 0 because it has not been initialised. PRIMASK = 0 means all configured interrupts are enabled.

Practical Example 4 — Critical Section (Disable/Enable Interrupts)

Portable critical section using PRIMASK via inline assembly
/* Disable all interrupts — enter critical section */
static inline void disable_interrupts(void)
{
    __asm volatile ("CPSID I" : : : "memory");
    /* CPSID I: Change Processor State, Interrupt Disable
       Sets PRIMASK = 1. All IRQs except NMI and HardFault are masked.
       "memory" clobber: prevents compiler moving memory ops across this barrier. */
}

/* Enable all interrupts — exit critical section */
static inline void enable_interrupts(void)
{
    __asm volatile ("CPSIE I" : : : "memory");
    /* CPSIE I: Change Processor State, Interrupt Enable — clears PRIMASK */
}

/* Safer pattern: save and restore PRIMASK state (handles nesting) */
static inline uint32_t enter_critical(void)
{
    uint32_t saved;
    __asm volatile (
        "MRS   %0, PRIMASK\n\t"   /* save current interrupt state */
        "CPSID I\n\t"             /* disable interrupts            */
        : "=r"(saved) : : "memory"
    );
    return saved;
}

static inline void exit_critical(uint32_t saved)
{
    __asm volatile (
        "MSR PRIMASK, %0\n\t"     /* restore saved state          */
        : : "r"(saved) : "memory"
    );
}

/* Usage in application code */
static volatile uint32_t shared_counter = 0;

void increment_counter(void)
{
    uint32_t irq_state = enter_critical();
    shared_counter++;   /* atomic — no ISR can interrupt this */
    exit_critical(irq_state);
}

ARM Calling Convention (AAPCS) — What the Compiler Assumes

When you write inline assembly that touches registers, you must respect the ARM Procedure Call Standard (AAPCS). Violating it corrupts the register state seen by surrounding C code in unpredictable ways.

Register(s)RoleWho Saves It?Your Inline Asm Rule
R0–R3 Arguments / return value / scratch Caller-saved — callee may destroy Declare in clobber list if you modify them, or use output/input operands
R12 (IP) Intra-procedure scratch Caller-saved — callee may destroy Declare in clobber list if you modify it
R4–R11 Callee-saved (preserved across calls) Callee-saved — if modified, must restore If your asm modifies R4–R11, PUSH them first and POP them after
R13 (SP) Stack pointer — must be 4-byte aligned Must always be valid Never destroy SP unless you restore it exactly. Must be 8-byte aligned at function calls.
R14 (LR) Link register / scratch in leaf functions Caller-saved — overwritten by BL If your asm calls BL, LR is overwritten. Save LR first if needed.
APSR flags Condition flags Not preserved across calls Add “cc” to clobber list if your asm sets N,Z,C,V flags

Ready-to-Use Inline Assembly Patterns

Collection of commonly-needed inline assembly fragments for Cortex-M
#include <stdint.h>

/* ── No-Operation delay ── */
static inline void nop(void)
{
    __asm volatile ("NOP");
}

/* ── Data Memory Barrier — ensures all memory accesses complete ── */
static inline void dmb(void)
{
    __asm volatile ("DMB" : : : "memory");
}

/* ── Data Synchronisation Barrier — stronger than DMB ── */
static inline void dsb(void)
{
    __asm volatile ("DSB" : : : "memory");
}

/* ── Instruction Synchronisation Barrier — flush pipeline ── */
static inline void isb(void)
{
    __asm volatile ("ISB" : : : "memory");
}

/* ── Wait For Interrupt — enter sleep until next interrupt ── */
static inline void wfi(void)
{
    __asm volatile ("WFI");
}

/* ── Wait For Event — used with SEV for multi-core sync ── */
static inline void wfe(void)
{
    __asm volatile ("WFE");
}

/* ── Software Breakpoint — halts the debugger ── */
static inline void bkpt(void)
{
    __asm volatile ("BKPT #0");
}

/* ── Count Leading Zeros (for fast log2, bit-weight) ── */
static inline uint32_t clz(uint32_t val)
{
    uint32_t result;
    __asm volatile ("CLZ %0, %1" : "=r"(result) : "r"(val));
    return result;  /* 0 if val has bit31 set; 32 if val == 0 */
}

/* ── Bit Reverse (for FFT butterfly operations, CRC) ── */
static inline uint32_t rbit(uint32_t val)
{
    uint32_t result;
    __asm volatile ("RBIT %0, %1" : "=r"(result) : "r"(val));
    return result;
}

/* ── Byte-Reverse (for endian conversion) ── */
static inline uint32_t rev(uint32_t val)
{
    uint32_t result;
    __asm volatile ("REV %0, %1" : "=r"(result) : "r"(val));
    return result;  /* 0x12345678 → 0x78563412 */
}

Frequently Asked Questions

What happens if you forget the clobber list in inline assembly?
If your inline assembly modifies a register (say R4) but you don’t declare it in the clobber list, GCC assumes R4 still holds whatever C value it had before the asm block. The compiler may then use R4 for subsequent C operations, producing wrong results — silently, with no warning. This is one of the most common bugs in handwritten inline assembly. Always declare every register your asm writes to that isn’t an explicit output operand.
Why does Cortex-M only have one mode (no Real mode / Protected mode like x86)?
ARM Cortex-M’s design philosophy is simplicity and determinism. Instead of complex mode hierarchies, it has just two modes (Thread / Handler) and two access levels (Privileged / Unprivileged), controlled by a single 3-bit register (CONTROL). This is sufficient for embedded RTOS use cases and is much easier to reason about than x86’s ring hierarchy. The MPU enforces memory isolation that x86 handles with the MMU — but the Cortex-M MPU is simpler and has fixed-latency access checking, which matters for real-time systems.
Why is it dangerous to write to LR directly in C code?
LR holds the return address. Overwriting it changes where the current function returns to. This is occasionally useful (e.g., for implementing tailored call dispatchers or profiling hooks), but must be done with extreme care in assembly. In C code, the compiler controls LR and may have its own value stored there. Using inline assembly to write LR without declaring it as an output/clobber will corrupt the return sequence silently.
What is EXC_RETURN and how does the CPU know how to return from an interrupt?
When an interrupt fires, the CPU saves the return state to the stack and loads a magic value into LR — this is EXC_RETURN (e.g. 0xFFFFFFF9). When the ISR returns (via POP {PC} or BX LR), the CPU detects this special EXC_RETURN value and instead of branching to it as an address, it performs an exception return: restores APSR, PC, and other registers from the stack and returns to the correct mode. The specific EXC_RETURN value encodes whether to return to Thread/Handler mode, and whether to use MSP or PSP for the unstacking.
When should I use __asm volatile vs __asm (without volatile)?
Use __asm volatile for any asm that has hardware side-effects or that the compiler might otherwise eliminate — which is almost always. Use plain __asm (without volatile) only for pure arithmetic operations where you are absolutely certain the compiler can freely move, merge, or eliminate the instruction if the output is unused. In practice, always use volatile to be safe.
Can I use GCC built-in intrinsics instead of writing inline assembly?
Yes, and for most operations you should. The CMSIS library (distributed with every STM32 SDK) provides inline functions for all common Cortex-M special register accesses: __get_PRIMASK(), __set_PRIMASK(), __get_MSP(), __set_PSP(), __enable_irq(), __disable_irq(), __WFI(), __DMB(), __DSB(), __ISB(), and many more. These are implemented using the same inline assembly patterns shown in this lecture, but are tested and portable. In production code, prefer CMSIS intrinsics over hand-written asm.
Next: Cortex-M Memory Map

Lecture 5 covers the ARM Cortex-M 4 GB address space in detail — how it is partitioned into code, SRAM, peripheral, and system regions, how memory-mapped registers are addressed, and how to read a memory map from a chip datasheet and Reference Manual.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *