Stack Memory and the
Full Descending Model
Every function call, local variable, and interrupt relies on the stack. This lecture explains what the stack is used for, how ARM Cortex-M’s Full Descending model works, what happens to the stack during exceptions, and how to size and protect it correctly in your STM32 projects.
- What is the Stack?
- Three Uses of Stack Memory on Cortex-M
- SRAM Layout: Global Data, Heap, and Stack
- Four Stack Operation Models
- Full Ascending Stack (FA)
- Full Descending Stack (FD) — ARM Standard
- PUSH and POP Mechanics on Cortex-M
- Stack Frame During a Function Call
- Exception Entry Stack Frame
- MSP and PSP — Two Stack Pointers
- Sizing the Stack Correctly
- Stack Overflow Detection
- FAQ
1. What is the Stack?
The stack is a region of SRAM managed automatically by the processor using a dedicated register called the Stack Pointer (SP). It operates on a Last-In-First-Out (LIFO) principle: the last item pushed onto the stack is the first item popped off.
Unlike global variables (which live at fixed addresses) or heap memory (which you
allocate with malloc), stack memory is allocated and freed
automatically as functions are called and return. You never manage it manually —
the compiler generates the PUSH and POP instructions for you.
Last In, First Out
Think of a stack of plates. You always add to the top and remove from the top. On Cortex-M, the Stack Pointer always points to the topmost item. PUSH adds an item; POP removes it.
Compiler-managed
You write void foo(int x) { int y = x + 1; ... } and the compiler
automatically generates instructions to allocate space for y on the
stack when foo is called and release it when foo returns.
2. Three Uses of Stack Memory on Cortex-M
The stack is not just for local variables. On ARM Cortex-M it serves three distinct roles, and understanding all three is essential for sizing it correctly.
Temporary register storage
AAPCS (ARM Procedure Call Standard) requires callee-saved registers (R4–R11, LR) to be pushed at function entry and popped at return. When a function needs these registers internally, it saves the caller’s values on the stack first.
Local variable storage
Local variables declared inside a function live on the stack for the duration of that function call. When the function returns, the stack space is automatically reclaimed — no memory leak possible.
Exception context save
When an interrupt fires, the hardware automatically pushes a set of CPU registers (R0–R3, R12, LR, PC, xPSR) onto the stack before executing the ISR. On return the hardware pops them back. This is called the exception frame.
/* Compiler-generated stack usage for a typical function */
void example(int a, int b)
{
int result; /* → stack slot allocated */
int temp[4]; /* → 16 bytes on stack */
/* callee-saved registers used internally: */
/* PUSH {R4, R5, LR} at entry (12 bytes) */
result = a + b;
temp[0] = result;
/* POP {R4, R5, PC} at return */
}
/* Stack usage for this one function: 12 (regs) + 4 (result) + 16 (temp) = 32 bytes */
3. SRAM Layout: Global Data, Heap, and Stack
The 128 KB SRAM of STM32F411 (0x20000000–0x2001FFFF) is divided into three logical sections by the linker script. The boundaries are determined at link time based on actual usage.
STM32F411 SRAM Layout (128 KB)
| Section | Contents | Grows | Notes |
|---|---|---|---|
| Global data (.data + .bss) |
Initialised globals, uninitialised globals, static variables | Fixed at link time | Copied / zeroed by Reset_Handler. Size known at compile time. |
| Heap | Dynamically allocated memory (malloc, new) |
Upward (↑) toward stack | Optional; avoid in deeply embedded systems. Risk: heap + stack collision. |
| Stack | Local variables, saved registers, exception frames | Downward (↓) toward heap | Starts at top of SRAM (_estack). Stack overflow if it reaches heap. |
/* Linker script symbols that define the SRAM layout */
MEMORY {
RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 128K
FLASH (rx): ORIGIN = 0x08000000, LENGTH = 512K
}
SECTIONS {
.data : { ... } > RAM /* initialised globals */
.bss : { ... } > RAM /* zero-initialised globals */
/* Heap starts here — grows upward */
._user_heap_stack :
{
. = ALIGN(8);
PROVIDE(end = .); /* heap start */
. = . + 0x400; /* 1 KB heap */
. = . + 0x400; /* 1 KB stack minimum */
. = ALIGN(8);
} > RAM
/* Stack top = top of RAM */
_estack = ORIGIN(RAM) + LENGTH(RAM); /* 0x20020000 */
}
4. Four Stack Operation Models
Different processor architectures have defined different conventions for how the stack pointer behaves during PUSH and POP. ARM documents four models, of which Cortex-M uses one exclusively.
| Model | SP points to… | PUSH direction | Used by |
|---|---|---|---|
| Full Descending (FD) | Last written (full) slot | Decrements SP first, then writes | ARM Cortex-M (all variants) |
| Full Ascending (FA) | Last written (full) slot | Writes first, then increments SP | Some older ARM architectures |
| Empty Descending (ED) | Next free (empty) slot | Writes to SP, then decrements | Rare, some DSP architectures |
| Empty Ascending (EA) | Next free (empty) slot | Writes to SP, then increments | Rare, historical |
The two key dimensions are:
- Full vs Empty — does SP point to the last item written (Full) or the next free slot (Empty)?
- Ascending vs Descending — does the stack grow toward higher addresses (Ascending) or lower addresses (Descending)?
5. Full Ascending Stack (FA) — For Comparison
The slides show a Full Ascending diagram as a contrast. In FA, the stack grows toward higher addresses. SP always points to the last written location (full slot). PUSH writes the data first, then increments SP.
Full Ascending Stack — PUSH sequence (addresses increasing upward)
(write then SP++)
(write then SP++)
(SP grows UP ↑)
FA: SP always points to the last written slot. Stack grows toward higher addresses. NOT used by ARM Cortex-M.
6. Full Descending Stack (FD) — ARM Standard
In the Full Descending model, the stack grows toward lower addresses. SP always points to the last written item (the “top” of the stack — which is physically at the lowest address used so far). On a PUSH, SP is decremented first, then the data is written to the new SP address.
Full Descending Stack — PUSH sequence (used by ARM Cortex-M)
SP at _estack
(SP– then write)
(SP grows DOWN ↓)
(SP at lowest addr)
FD: SP always points to the last written item at the lowest address in use. Stack grows toward lower addresses. ARM Cortex-M uses this model exclusively.
7. PUSH and POP Mechanics on Cortex-M
The Cortex-M instruction set provides PUSH and POP instructions
that operate on a register list. They always use the Full Descending convention.
/* PUSH {reg-list} — pushes registers in descending register order
Each register takes 4 bytes on the stack.
SP is decremented by 4 × (number of registers) first. */
PUSH {R4, R5, R6, LR}
; Equivalent to:
; SP = SP - 16 ; make room for 4 registers
; [SP+12] = R4 ; highest register at highest address
; [SP+8] = R5
; [SP+4] = R6
; [SP+0] = LR ; lowest register at lowest address (SP points here)
/* POP {reg-list} — pops in ascending register order, SP increments */
POP {R4, R5, R6, PC}
; Equivalent to:
; R4 = [SP+12] ; restore callee-saved registers
; R5 = [SP+8]
; R6 = [SP+4]
; PC = [SP+0] ; return: PC = saved LR value
; SP = SP + 16 ; release the stack frame
POP {R4–R7, PC} rather than
a separate branch instruction. The saved LR (return address) is popped directly into PC,
causing the processor to jump back to the call site. This saves one instruction and
allows the compiler to combine register restore with function return.
SP alignment — always 8 bytes at function call boundary
/* AAPCS requires SP to be 8-byte aligned at any public function call.
The compiler inserts padding pushes if necessary. */
void caller(void)
{
/* SP is 8-byte aligned here (guaranteed by ABI) */
callee(1, 2, 3, 4);
/* After call returns, SP restored to same 8-byte aligned value */
}
/* If callee only saves 1 register (4 bytes), compiler adds a dummy push:
PUSH {R4, R7} ; 8 bytes — maintains alignment
instead of:
PUSH {R4} ; 4 bytes — would break alignment */
8. Stack Frame During a Function Call
Every function call creates a stack frame — a contiguous region on the stack that holds all the data belonging to that invocation of the function. Understanding stack frames lets you read disassembly, debug crashes, and reason about stack depth.
Stack Frame Layout — foo() calls bar()
foo calls
bar()
When bar() returns: its frame is discarded (SP moves up). foo’s frame is intact and execution resumes.
/* Compiled code for a simple function (arm-none-eabi-gcc -O1) */
int sum(int a, int b, int c)
{
int total = a + b + c; /* local variable */
return total;
}
/* Disassembly (total fits in R0 — no stack needed!): */
; ADD R0, R0, R1 ; R0 = a + b
; ADD R0, R0, R2 ; R0 = (a+b) + c
; BX LR ; return total in R0
/* With a larger function needing stack allocation: */
void complex(void)
{
uint8_t buffer[64]; /* 64 bytes on stack */
int i; /* 4 bytes on stack */
/* ... */
}
/* Disassembly entry: */
; PUSH {R4, LR} ; save callee-saved + return address (8 bytes)
; SUB SP, SP, #68 ; allocate buffer[64] + i + padding = 68 bytes
; ...
; ADD SP, SP, #68 ; release locals
; POP {R4, PC} ; restore + return
9. Exception Entry Stack Frame
When any exception or interrupt fires on Cortex-M, the processor performs an automatic exception entry stacking before the ISR begins. Eight registers are pushed onto the stack by hardware — no compiler code needed.
Hardware-Pushed Exception Frame (Cortex-M4, no FPU)
8 registers × 4 bytes = 32 bytes consumed by every exception entry. On exception return, hardware pops these 8 registers automatically.
Why R0–R3, R12, LR, PC, xPSR?
These are the registers the interrupted code might have been using at the exact moment of interruption. By saving them automatically, the ISR can use R0–R3 and R12 freely without corrupting the interrupted code’s state. R4–R11 are callee-saved — the ISR is responsible for saving those if it uses them.
With FPU: 26 registers!
If the Cortex-M4 FPU is enabled and the interrupted code was using floating-point, the hardware pushes an additional 18 FPU registers (S0–S15, FPSCR, reserved word) making the exception frame 26 words (104 bytes). This is controlled by the FPCA bit in the CONTROL register.
/* The EXC_RETURN value placed in LR by hardware encodes the return path.
Common values for Cortex-M4: */
// 0xFFFFFFF9 — Return to Thread mode using MSP (no FPU state)
// 0xFFFFFFFD — Return to Thread mode using PSP (no FPU state)
// 0xFFFFFFE9 — Return to Thread mode using MSP (with FPU state)
// 0xFFFFFFED — Return to Thread mode using PSP (with FPU state)
void SysTick_Handler(void)
{
/* On entry: hardware has already pushed R0-R3,R12,LR,PC,xPSR */
/* SP has moved down by 32 bytes */
system_tick++; /* safe to use R0-R3 freely here */
/* On exit: BX LR (EXC_RETURN) triggers hardware unstack */
}
10. MSP and PSP — Two Stack Pointers
Cortex-M provides two stack pointers to support RTOS designs where kernel and user tasks need separate stacks:
Main Stack Pointer
- Always used in Handler mode (ISRs, fault handlers)
- Used in Thread mode when CONTROL[1] (SPSEL) = 0
- Default after reset — startup code uses MSP
- Points to the OS/kernel stack
- Initial value loaded from vector table word 0
Process Stack Pointer
- Used in Thread mode when CONTROL[1] (SPSEL) = 1
- Never used in Handler mode
- Each RTOS task gets its own PSP value
- Allows the OS to isolate task stacks from the kernel stack
- Must be set explicitly by software before switching to PSP
/* Reading and writing MSP / PSP via MRS/MSR */
static inline uint32_t get_msp(void)
{
uint32_t val;
__asm volatile ("MRS %0, MSP" : "=r"(val));
return val;
}
static inline void set_psp(uint32_t val)
{
__asm volatile ("MSR PSP, %0" : : "r"(val) : "memory");
}
/* Switch Thread mode to use PSP (typical RTOS task switch init) */
static inline void switch_to_psp(void)
{
uint32_t ctrl;
__asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));
ctrl |= (1U << 1); /* set SPSEL = 1 → use PSP */
__asm volatile ("MSR CONTROL, %0\n\t" "ISB\n\t"
: : "r"(ctrl) : "memory");
}
/* FreeRTOS pattern: each task has its own stack array */
static uint32_t task1_stack[256]; /* 1 KB stack for task1 */
static uint32_t task2_stack[256]; /* 1 KB stack for task2 */
void start_task1(void)
{
/* PSP = top of task1_stack */
set_psp((uint32_t)(task1_stack + 256));
switch_to_psp();
/* Now task1 uses PSP; ISRs and kernel use MSP */
}
11. Sizing the Stack Correctly
Choosing the right stack size is one of the most important and underappreciated tasks in embedded system design. Too small and you overflow; too large and you waste precious SRAM.
What consumes stack space?
| Source | Bytes consumed | Notes |
|---|---|---|
| Exception frame (no FPU) | 32 bytes | Per interrupt nesting level |
| Exception frame (with FPU) | 104 bytes | If FPCA is set |
| Callee-saved registers (R4–R11, LR) | 36 bytes max | Per function that uses them |
| Local variables | Varies | Check compiler output with -fstack-usage |
| Function argument passing (>4 args) | 4 bytes per extra arg | First 4 args in R0–R3 |
| Recursive calls | N × frame size | Avoid recursion in embedded! |
Practical sizing strategy
/* Step 1: Let GCC report per-function stack usage */
/* Compile with: arm-none-eabi-gcc -fstack-usage ... */
/* Produces .su files: foo.c:23:void foo() 48 static */
/* Step 2: Find worst-case call chain manually or with tools like:
- cppcheck --check-level=exhaustive (limited)
- pc-lint, Polyspace (commercial)
- arm-none-eabi-nm + manual analysis */
/* Step 3: Add safety margin */
#define MAX_CALL_DEPTH_BYTES 512 /* from analysis */
#define MAX_IRQ_NESTING 4 /* 4 levels of nested IRQs */
#define IRQ_FRAME_BYTES 104 /* worst case with FPU */
#define SAFETY_MARGIN 256 /* 25% headroom */
#define MIN_STACK_SIZE (MAX_CALL_DEPTH_BYTES + \
MAX_IRQ_NESTING * IRQ_FRAME_BYTES + \
SAFETY_MARGIN)
/* = 512 + 416 + 256 = 1184 bytes → round up to 1536 or 2048 bytes */
/* In linker script: */
/* _Min_Stack_Size = 0x800; /* 2 KB minimum stack */
_Min_Stack_Size = 0x400 (1 KB) by default. This is often
too small for real applications that use printf (which calls _write, which may recurse
through the C library with significant stack usage). For projects using USB, FatFS, or
printf, consider 4–8 KB minimum stack.
12. Stack Overflow Detection
By default, Cortex-M has no automatic stack overflow detection. When the stack pointer goes below the bottom of the stack region, it silently overwrites .bss or .data — corrupting global variables and producing impossible-to-trace bugs. Here are four techniques to catch overflows.
Stack watermarking (canary pattern)
Fill the entire stack with a known pattern at startup. After running for a while, inspect how far the pattern has been overwritten to find the high-water mark.
/* Fill stack with 0xDEADBEEF pattern */
extern uint32_t _estack, _Min_Stack_Size;
void stack_paint(void) {
uint32_t *p = &_estack - (_Min_Stack_Size/4);
while (p < &_estack) *p++ = 0xDEADBEEFU;
}
/* Check: scan from bottom, count consecutive 0xDEADBEEF */
MPU stack guard region
Configure the MPU to make the last 32 bytes of the stack region read-only with no execute. Any access triggers a MemManage fault, catching overflow before it corrupts data.
/* MPU region 7: guard page below stack */
MPU_RBAR = (STACK_BOTTOM & ~0x1F) | (7 << 1) | 1;
MPU_RASR = (0b111 << 28) /* no access */
| (0b00100 << 1) /* 32 bytes */
| 1; /* enable */
PSPLIM / MSPLIM (Cortex-M33/M55)
Cortex-M33 and later add dedicated stack limit registers. The processor raises a fault when SP goes below the configured limit. Cortex-M4 does NOT have these — use MPU instead.
FreeRTOS stack overflow hooks
FreeRTOS checks the watermark at each context switch when
configCHECK_FOR_STACK_OVERFLOW is set. Call
vApplicationStackOverflowHook() to catch the overflow in your code.
/* Complete stack watermark check utility for bare-metal STM32 */
#include <stdint.h>
#define STACK_CANARY 0xDEADBEEFU
/* Called from Reset_Handler BEFORE main() */
void stack_paint(void)
{
extern uint32_t _estack;
extern uint32_t _Min_Stack_Size;
uint32_t stack_bottom = (uint32_t)&_estack - (uint32_t)&_Min_Stack_Size;
uint32_t *p = (uint32_t *)stack_bottom;
/* Paint from bottom up to current SP */
while ((uint32_t)p < __get_MSP()) {
*p++ = STACK_CANARY;
}
}
/* Call periodically or from HardFault handler */
uint32_t stack_get_unused_bytes(void)
{
extern uint32_t _estack;
extern uint32_t _Min_Stack_Size;
uint32_t stack_bottom = (uint32_t)&_estack - (uint32_t)&_Min_Stack_Size;
uint32_t *p = (uint32_t *)stack_bottom;
uint32_t unused = 0;
while (*p == STACK_CANARY) {
unused += 4;
p++;
}
return unused;
}
void check_stack_health(void)
{
uint32_t free_bytes = stack_get_unused_bytes();
if (free_bytes < 64) {
/* DANGER: less than 64 bytes of stack remaining! */
Error_Handler();
}
}
13. FAQ
Next: Exceptions and Interrupts — NVIC Deep Dive
The stack is where exception frames land. The next lecture covers the NVIC in depth — priority levels, grouping, enabling and pending interrupts, and writing ISRs that interact correctly with the stack and the exception model.

2 Comments