Stack Memory and theFull Descending Model-Embedded C Training Institute in Hyderabad

Cortex-M Stack Memory Full Descending | STM32 Embedded

Lecture 10 — ARM Cortex-M Programming

Stack Memory and the
Full Descending Model

Every function call, local variable, and interrupt relies on the stack. This lecture explains what the stack is used for, how ARM Cortex-M’s Full Descending model works, what happens to the stack during exceptions, and how to size and protect it correctly in your STM32 projects.

Covers source pages 81–90  |  STM32F411  |  ARM Cortex-M TRM

1. What is the Stack?

The stack is a region of SRAM managed automatically by the processor using a dedicated register called the Stack Pointer (SP). It operates on a Last-In-First-Out (LIFO) principle: the last item pushed onto the stack is the first item popped off.

Unlike global variables (which live at fixed addresses) or heap memory (which you allocate with malloc), stack memory is allocated and freed automatically as functions are called and return. You never manage it manually — the compiler generates the PUSH and POP instructions for you.

LIFO

Last In, First Out

Think of a stack of plates. You always add to the top and remove from the top. On Cortex-M, the Stack Pointer always points to the topmost item. PUSH adds an item; POP removes it.

AUTOMATIC

Compiler-managed

You write void foo(int x) { int y = x + 1; ... } and the compiler automatically generates instructions to allocate space for y on the stack when foo is called and release it when foo returns.

2. Three Uses of Stack Memory on Cortex-M

The stack is not just for local variables. On ARM Cortex-M it serves three distinct roles, and understanding all three is essential for sizing it correctly.

USE 1

Temporary register storage

AAPCS (ARM Procedure Call Standard) requires callee-saved registers (R4–R11, LR) to be pushed at function entry and popped at return. When a function needs these registers internally, it saves the caller’s values on the stack first.

USE 2

Local variable storage

Local variables declared inside a function live on the stack for the duration of that function call. When the function returns, the stack space is automatically reclaimed — no memory leak possible.

USE 3

Exception context save

When an interrupt fires, the hardware automatically pushes a set of CPU registers (R0–R3, R12, LR, PC, xPSR) onto the stack before executing the ISR. On return the hardware pops them back. This is called the exception frame.

/* Compiler-generated stack usage for a typical function */
void example(int a, int b)
{
    int result;          /* → stack slot allocated */
    int temp[4];         /* → 16 bytes on stack    */

    /* callee-saved registers used internally: */
    /* PUSH {R4, R5, LR} at entry  (12 bytes)  */

    result = a + b;
    temp[0] = result;

    /* POP {R4, R5, PC} at return  */
}

/* Stack usage for this one function: 12 (regs) + 4 (result) + 16 (temp) = 32 bytes */

3. SRAM Layout: Global Data, Heap, and Stack

The 128 KB SRAM of STM32F411 (0x20000000–0x2001FFFF) is divided into three logical sections by the linker script. The boundaries are determined at link time based on actual usage.

STM32F411 SRAM Layout (128 KB)

0x20000000 0x2001FFFF
Global Data
(.data + .bss)
Heap
(malloc)
Stack
(grows ↓)
RAM_START → _edata/_ebss _end → brk _estack (top) ↓ growing down
SectionContentsGrowsNotes
Global data
(.data + .bss)
Initialised globals, uninitialised globals, static variables Fixed at link time Copied / zeroed by Reset_Handler. Size known at compile time.
Heap Dynamically allocated memory (malloc, new) Upward (↑) toward stack Optional; avoid in deeply embedded systems. Risk: heap + stack collision.
Stack Local variables, saved registers, exception frames Downward (↓) toward heap Starts at top of SRAM (_estack). Stack overflow if it reaches heap.
Heap-stack collision — the silent killer
If you allocate too much on the heap and the stack grows down far enough, they collide. There is no hardware protection by default (without MPU configuration). The stack will silently overwrite heap data and vice versa, producing bizarre, non-reproducible bugs. This is why deeply embedded systems often avoid dynamic allocation entirely.
/* Linker script symbols that define the SRAM layout */
MEMORY {
    RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 128K
    FLASH (rx): ORIGIN = 0x08000000, LENGTH = 512K
}

SECTIONS {
    .data : { ... } > RAM   /* initialised globals */
    .bss  : { ... } > RAM   /* zero-initialised globals */

    /* Heap starts here — grows upward */
    ._user_heap_stack :
    {
        . = ALIGN(8);
        PROVIDE(end = .);    /* heap start */
        . = . + 0x400;       /* 1 KB heap  */
        . = . + 0x400;       /* 1 KB stack minimum */
        . = ALIGN(8);
    } > RAM

    /* Stack top = top of RAM */
    _estack = ORIGIN(RAM) + LENGTH(RAM);   /* 0x20020000 */
}

4. Four Stack Operation Models

Different processor architectures have defined different conventions for how the stack pointer behaves during PUSH and POP. ARM documents four models, of which Cortex-M uses one exclusively.

Model SP points to… PUSH direction Used by
Full Descending (FD) Last written (full) slot Decrements SP first, then writes ARM Cortex-M (all variants)
Full Ascending (FA) Last written (full) slot Writes first, then increments SP Some older ARM architectures
Empty Descending (ED) Next free (empty) slot Writes to SP, then decrements Rare, some DSP architectures
Empty Ascending (EA) Next free (empty) slot Writes to SP, then increments Rare, historical

The two key dimensions are:

  • Full vs Empty — does SP point to the last item written (Full) or the next free slot (Empty)?
  • Ascending vs Descending — does the stack grow toward higher addresses (Ascending) or lower addresses (Descending)?
ARM Cortex-M always uses Full Descending
This is hard-wired into the architecture — it is not configurable. All Cortex-M chips (M0, M0+, M3, M4, M7, M33…) use Full Descending. The PUSH instruction automatically decrements SP by 4 then writes, and POP reads then increments by 4.

5. Full Ascending Stack (FA) — For Comparison

The slides show a Full Ascending diagram as a contrast. In FA, the stack grows toward higher addresses. SP always points to the last written location (full slot). PUSH writes the data first, then increments SP.

Full Ascending Stack — PUSH sequence (addresses increasing upward)

free
free
free
← SP: 0xF0 (empty)
prev data
Before any PUSH
→
free
free
← SP: 0xF1 (full)
0xF1 ← just pushed
prev data
PUSH 0xF1
(write then SP++)
→
free
← SP: 0xF2 (full)
0xF2
0xF1
prev data
PUSH 0xF2
(write then SP++)
→
← SP: 0xF3 (full)
0xF3
0xF2
0xF1
prev data
PUSH 0xF3
(SP grows UP ↑)

FA: SP always points to the last written slot. Stack grows toward higher addresses. NOT used by ARM Cortex-M.

6. Full Descending Stack (FD) — ARM Standard

In the Full Descending model, the stack grows toward lower addresses. SP always points to the last written item (the “top” of the stack — which is physically at the lowest address used so far). On a PUSH, SP is decremented first, then the data is written to the new SP address.

Full Descending Stack — PUSH sequence (used by ARM Cortex-M)

prev data (high addr)
← SP (initial top)
free ↓
free ↓
free ↓
Initial state
SP at _estack
→
prev data
0xF1
← SP (SP-4 then write)
free ↓
free ↓
PUSH 0xF1
(SP– then write)
→
prev data
0xF1
0xF2
← SP
free ↓
PUSH 0xF2
(SP grows DOWN ↓)
→
prev data
0xF1
0xF2
0xF3
← SP (lowest addr used)
PUSH 0xF3
(SP at lowest addr)

FD: SP always points to the last written item at the lowest address in use. Stack grows toward lower addresses. ARM Cortex-M uses this model exclusively.

Why “Full” Descending?
“Full” means SP always points to a full (occupied) slot — the last word that was written. Contrast with “Empty” where SP points to the next available slot. On Cortex-M, when you push: SP = SP − 4, then write to [SP]. When you pop: read from [SP], then SP = SP + 4. SP is always valid and always points to real data on the stack (never to garbage).

7. PUSH and POP Mechanics on Cortex-M

The Cortex-M instruction set provides PUSH and POP instructions that operate on a register list. They always use the Full Descending convention.

/* PUSH {reg-list} — pushes registers in descending register order
   Each register takes 4 bytes on the stack.
   SP is decremented by 4 × (number of registers) first.            */

PUSH {R4, R5, R6, LR}
; Equivalent to:
; SP = SP - 16        ; make room for 4 registers
; [SP+12] = R4        ; highest register at highest address
; [SP+8]  = R5
; [SP+4]  = R6
; [SP+0]  = LR        ; lowest register at lowest address (SP points here)

/* POP {reg-list} — pops in ascending register order, SP increments */
POP {R4, R5, R6, PC}
; Equivalent to:
; R4 = [SP+12]        ; restore callee-saved registers
; R5 = [SP+8]
; R6 = [SP+4]
; PC = [SP+0]         ; return: PC = saved LR value
; SP = SP + 16        ; release the stack frame
POP {…, PC} — the return mechanism
The standard function return on Cortex-M is POP {R4–R7, PC} rather than a separate branch instruction. The saved LR (return address) is popped directly into PC, causing the processor to jump back to the call site. This saves one instruction and allows the compiler to combine register restore with function return.

SP alignment — always 8 bytes at function call boundary

/* AAPCS requires SP to be 8-byte aligned at any public function call.
   The compiler inserts padding pushes if necessary.                    */

void caller(void)
{
    /* SP is 8-byte aligned here (guaranteed by ABI) */
    callee(1, 2, 3, 4);
    /* After call returns, SP restored to same 8-byte aligned value */
}

/* If callee only saves 1 register (4 bytes), compiler adds a dummy push:
   PUSH {R4, R7}   ; 8 bytes — maintains alignment
   instead of:
   PUSH {R4}       ; 4 bytes — would break alignment               */

8. Stack Frame During a Function Call

Every function call creates a stack frame — a contiguous region on the stack that holds all the data belonging to that invocation of the function. Understanding stack frames lets you read disassembly, debug crashes, and reason about stack depth.

Stack Frame Layout — foo() calls bar()

foo() stack frame
← high address (SP before foo entered)
LR (return to caller of foo)
R4 (saved callee register)
R5 (saved callee register)
local int x
local int buf[3] slot 0
local int buf[3] slot 1
local int buf[3] slot 2
← SP (points here while in foo)
→
foo calls
bar()
bar() stack frame added on top
← foo’s frame above (unchanged)
LR = return into foo
R4 (saved for bar)
local int result
← SP (now lower address)

When bar() returns: its frame is discarded (SP moves up). foo’s frame is intact and execution resumes.

/* Compiled code for a simple function (arm-none-eabi-gcc -O1) */
int sum(int a, int b, int c)
{
    int total = a + b + c;  /* local variable */
    return total;
}

/* Disassembly (total fits in R0 — no stack needed!): */
; ADD R0, R0, R1    ; R0 = a + b
; ADD R0, R0, R2    ; R0 = (a+b) + c
; BX  LR            ; return total in R0

/* With a larger function needing stack allocation: */
void complex(void)
{
    uint8_t buffer[64];   /* 64 bytes on stack */
    int     i;            /* 4 bytes on stack  */
    /* ... */
}

/* Disassembly entry: */
; PUSH {R4, LR}         ; save callee-saved + return address (8 bytes)
; SUB  SP, SP, #68      ; allocate buffer[64] + i + padding = 68 bytes
; ...
; ADD  SP, SP, #68      ; release locals
; POP  {R4, PC}         ; restore + return

9. Exception Entry Stack Frame

When any exception or interrupt fires on Cortex-M, the processor performs an automatic exception entry stacking before the ISR begins. Eight registers are pushed onto the stack by hardware — no compiler code needed.

Hardware-Pushed Exception Frame (Cortex-M4, no FPU)

← SP before exception (higher addr)
xPSRSP+28
PC (return address)SP+24
LR (EXC_RETURN)SP+20
R12SP+16
R3SP+12
R2SP+8
R1SP+4
R0SP+0
← SP after stacking (ISR entry SP)

8 registers × 4 bytes = 32 bytes consumed by every exception entry. On exception return, hardware pops these 8 registers automatically.

CALLER-SAVED REGS

Why R0–R3, R12, LR, PC, xPSR?

These are the registers the interrupted code might have been using at the exact moment of interruption. By saving them automatically, the ISR can use R0–R3 and R12 freely without corrupting the interrupted code’s state. R4–R11 are callee-saved — the ISR is responsible for saving those if it uses them.

FPU EXTENSION

With FPU: 26 registers!

If the Cortex-M4 FPU is enabled and the interrupted code was using floating-point, the hardware pushes an additional 18 FPU registers (S0–S15, FPSCR, reserved word) making the exception frame 26 words (104 bytes). This is controlled by the FPCA bit in the CONTROL register.

/* The EXC_RETURN value placed in LR by hardware encodes the return path.
   Common values for Cortex-M4:                                          */

// 0xFFFFFFF9 — Return to Thread mode using MSP (no FPU state)
// 0xFFFFFFFD — Return to Thread mode using PSP (no FPU state)
// 0xFFFFFFE9 — Return to Thread mode using MSP (with FPU state)
// 0xFFFFFFED — Return to Thread mode using PSP (with FPU state)

void SysTick_Handler(void)
{
    /* On entry: hardware has already pushed R0-R3,R12,LR,PC,xPSR */
    /* SP has moved down by 32 bytes                                */

    system_tick++;   /* safe to use R0-R3 freely here              */

    /* On exit: BX LR (EXC_RETURN) triggers hardware unstack       */
}

10. MSP and PSP — Two Stack Pointers

Cortex-M provides two stack pointers to support RTOS designs where kernel and user tasks need separate stacks:

MSP

Main Stack Pointer

  • Always used in Handler mode (ISRs, fault handlers)
  • Used in Thread mode when CONTROL[1] (SPSEL) = 0
  • Default after reset — startup code uses MSP
  • Points to the OS/kernel stack
  • Initial value loaded from vector table word 0
PSP

Process Stack Pointer

  • Used in Thread mode when CONTROL[1] (SPSEL) = 1
  • Never used in Handler mode
  • Each RTOS task gets its own PSP value
  • Allows the OS to isolate task stacks from the kernel stack
  • Must be set explicitly by software before switching to PSP
/* Reading and writing MSP / PSP via MRS/MSR */

static inline uint32_t get_msp(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, MSP" : "=r"(val));
    return val;
}

static inline void set_psp(uint32_t val)
{
    __asm volatile ("MSR PSP, %0" : : "r"(val) : "memory");
}

/* Switch Thread mode to use PSP (typical RTOS task switch init) */
static inline void switch_to_psp(void)
{
    uint32_t ctrl;
    __asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));
    ctrl |= (1U << 1);   /* set SPSEL = 1 → use PSP */
    __asm volatile ("MSR CONTROL, %0\n\t" "ISB\n\t"
                    : : "r"(ctrl) : "memory");
}

/* FreeRTOS pattern: each task has its own stack array */
static uint32_t task1_stack[256];   /* 1 KB stack for task1 */
static uint32_t task2_stack[256];   /* 1 KB stack for task2 */

void start_task1(void)
{
    /* PSP = top of task1_stack */
    set_psp((uint32_t)(task1_stack + 256));
    switch_to_psp();
    /* Now task1 uses PSP; ISRs and kernel use MSP */
}

11. Sizing the Stack Correctly

Choosing the right stack size is one of the most important and underappreciated tasks in embedded system design. Too small and you overflow; too large and you waste precious SRAM.

What consumes stack space?

SourceBytes consumedNotes
Exception frame (no FPU)32 bytesPer interrupt nesting level
Exception frame (with FPU)104 bytesIf FPCA is set
Callee-saved registers (R4–R11, LR)36 bytes maxPer function that uses them
Local variablesVariesCheck compiler output with -fstack-usage
Function argument passing (>4 args)4 bytes per extra argFirst 4 args in R0–R3
Recursive callsN × frame sizeAvoid recursion in embedded!

Practical sizing strategy

/* Step 1: Let GCC report per-function stack usage */
/* Compile with: arm-none-eabi-gcc -fstack-usage ... */
/* Produces .su files:  foo.c:23:void foo() 48 static */

/* Step 2: Find worst-case call chain manually or with tools like:
   - cppcheck --check-level=exhaustive (limited)
   - pc-lint, Polyspace (commercial)
   - arm-none-eabi-nm + manual analysis                            */

/* Step 3: Add safety margin */
#define MAX_CALL_DEPTH_BYTES  512    /* from analysis               */
#define MAX_IRQ_NESTING       4      /* 4 levels of nested IRQs     */
#define IRQ_FRAME_BYTES       104    /* worst case with FPU         */
#define SAFETY_MARGIN         256    /* 25% headroom                */

#define MIN_STACK_SIZE  (MAX_CALL_DEPTH_BYTES + \
                         MAX_IRQ_NESTING * IRQ_FRAME_BYTES + \
                         SAFETY_MARGIN)
/* = 512 + 416 + 256 = 1184 bytes → round up to 1536 or 2048 bytes */

/* In linker script: */
/* _Min_Stack_Size = 0x800;  /* 2 KB minimum stack */
STM32CubeIDE default stack size
STM32CubeIDE sets _Min_Stack_Size = 0x400 (1 KB) by default. This is often too small for real applications that use printf (which calls _write, which may recurse through the C library with significant stack usage). For projects using USB, FatFS, or printf, consider 4–8 KB minimum stack.

12. Stack Overflow Detection

By default, Cortex-M has no automatic stack overflow detection. When the stack pointer goes below the bottom of the stack region, it silently overwrites .bss or .data — corrupting global variables and producing impossible-to-trace bugs. Here are four techniques to catch overflows.

METHOD 1

Stack watermarking (canary pattern)

Fill the entire stack with a known pattern at startup. After running for a while, inspect how far the pattern has been overwritten to find the high-water mark.

/* Fill stack with 0xDEADBEEF pattern */
extern uint32_t _estack, _Min_Stack_Size;
void stack_paint(void) {
    uint32_t *p = &_estack - (_Min_Stack_Size/4);
    while (p < &_estack) *p++ = 0xDEADBEEFU;
}
/* Check: scan from bottom, count consecutive 0xDEADBEEF */
METHOD 2

MPU stack guard region

Configure the MPU to make the last 32 bytes of the stack region read-only with no execute. Any access triggers a MemManage fault, catching overflow before it corrupts data.

/* MPU region 7: guard page below stack */
MPU_RBAR = (STACK_BOTTOM & ~0x1F) | (7 << 1) | 1;
MPU_RASR = (0b111 << 28)  /* no access */
         | (0b00100 << 1) /* 32 bytes  */
         | 1;              /* enable    */
METHOD 3

PSPLIM / MSPLIM (Cortex-M33/M55)

Cortex-M33 and later add dedicated stack limit registers. The processor raises a fault when SP goes below the configured limit. Cortex-M4 does NOT have these — use MPU instead.

METHOD 4

FreeRTOS stack overflow hooks

FreeRTOS checks the watermark at each context switch when configCHECK_FOR_STACK_OVERFLOW is set. Call vApplicationStackOverflowHook() to catch the overflow in your code.

/* Complete stack watermark check utility for bare-metal STM32 */
#include <stdint.h>

#define STACK_CANARY  0xDEADBEEFU

/* Called from Reset_Handler BEFORE main() */
void stack_paint(void)
{
    extern uint32_t _estack;
    extern uint32_t _Min_Stack_Size;
    uint32_t stack_bottom = (uint32_t)&_estack - (uint32_t)&_Min_Stack_Size;
    uint32_t *p = (uint32_t *)stack_bottom;

    /* Paint from bottom up to current SP */
    while ((uint32_t)p < __get_MSP()) {
        *p++ = STACK_CANARY;
    }
}

/* Call periodically or from HardFault handler */
uint32_t stack_get_unused_bytes(void)
{
    extern uint32_t _estack;
    extern uint32_t _Min_Stack_Size;
    uint32_t stack_bottom = (uint32_t)&_estack - (uint32_t)&_Min_Stack_Size;
    uint32_t *p = (uint32_t *)stack_bottom;
    uint32_t unused = 0;

    while (*p == STACK_CANARY) {
        unused += 4;
        p++;
    }
    return unused;
}

void check_stack_health(void)
{
    uint32_t free_bytes = stack_get_unused_bytes();
    if (free_bytes < 64) {
        /* DANGER: less than 64 bytes of stack remaining! */
        Error_Handler();
    }
}

13. FAQ

Q: Why does the stack grow downward? Wouldn’t upward be more intuitive?
It is a historical convention. Early computer designs placed the stack at the top of memory and it grew downward, while program data started at the bottom and grew upward. This way both could share the same address space without a fixed boundary. ARM inherited this convention and standardised it. Growing downward also means that if a stack frame contains an array, the array indices (0, 1, 2…) map to increasing addresses — which matches C’s memory model for arrays on the stack.
Q: What is the difference between MSP and PSP in a bare-metal (no RTOS) project?
In a bare-metal project with no RTOS, you always use MSP for everything — both the main application code and all ISRs. PSP is optional and only becomes useful when you have an RTOS that wants to give each task its own isolated stack. The startup file sets the initial MSP value from the vector table word 0 (pointing to the top of SRAM), and you never need to configure PSP unless you add an RTOS.
Q: How many bytes does the hardware push on exception entry?
On Cortex-M3/M4 without FPU context: 32 bytes (8 registers × 4 bytes: R0, R1, R2, R3, R12, LR, PC, xPSR). On Cortex-M4 with FPU enabled and the interrupted code was using floating-point (FPCA=1): 104 bytes (8 + 18 additional FPU registers: S0–S15, FPSCR, and one reserved word). On Cortex-M0 (no FPU): always 32 bytes, same as M3/M4 without FPU.
Q: What happens if a local variable is too large for the stack?
The compiler will still allocate it on the stack — it does not check at compile time whether enough stack space exists. At runtime, SP will move below the stack region and silently overwrite whatever is below (heap, .bss, .data). The first symptom is usually a corrupted global variable or a HardFault from an invalid address. To avoid this: use static or global buffers for large arrays, or allocate them with malloc() on the heap (with appropriate heap size). Never put a buffer larger than ~256 bytes on the stack without careful analysis.
Q: Can nested interrupts overflow the stack?
Yes, easily. Each interrupt nesting level adds one exception frame (32–104 bytes) plus whatever the ISR itself uses. If you have 4 levels of nesting, that is at minimum 128 bytes of exception frames alone before the ISR code runs. Configure your NVIC priorities carefully to limit nesting depth, and always include the maximum nesting depth in your stack size calculation.
Q: Is it safe to call printf() inside an ISR?
Generally no, for two reasons. First, printf() may use a large internal buffer on the stack (newlib’s printf can use 1–2 KB depending on format string). Second, printf() is not re-entrant — if the main code is in the middle of a printf() when the ISR fires and the ISR also calls printf(), you will corrupt the internal state. Use ITM direct writes or a lock-free ring buffer for debug output from ISRs instead.

Next: Exceptions and Interrupts — NVIC Deep Dive

The stack is where exception frames land. The next lecture covers the NVIC in depth — priority levels, grouping, enabling and pending interrupts, and writing ISRs that interact correctly with the stack and the exception model.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *