ARM Stack PUSH POP Operations-Embedded C Course Online

ARM Stack PUSH POP Operations | Embedded Systems ARM Cortex-M

EmbeddedPathashala — ARM Cortex-M Course Home All Lectures STM32 Projects

ARM Stack PUSH POP Operations

Master the ARM Cortex-M Full Descending stack through a complete worked example — PUSH LR/R0/R1, POP R2/R3/PC — with cycle-accurate register state tracking, AAPCS conventions, and compiler-generated prologue/epilogue patterns.

Home › ARM Cortex-M Course › Lecture 11 — ARM Stack PUSH POP Operations

What PUSH and POP Actually Do

At their core, PUSH and POP are not special-purpose instructions unique to stacks — they are convenient aliases for the more general Store Multiple (STM) and Load Multiple (LDM) instructions operating through the stack pointer (SP / R13). Understanding them at the instruction level removes any mystery about how function calls preserve register state.

PUSH {Rn, …} is equivalent to STMDB SP!, {Rn, …}
  • STMDB = Store Multiple Decrement Before
  • SP is decremented before each store (Full Descending rule)

POP {Rn, …} is equivalent to LDMIA SP!, {Rn, …}
  • LDMIA = Load Multiple Increment After
  • Each load reads SP’s current location, then SP is incremented

The ! suffix means “writeback” — the updated SP value is written back to the SP register after the operation completes. Without writeback the stack pointer would not actually move, which would be an error for any real push or pop.

On Cortex-M, PUSH and POP always operate in groups of 4-byte words and always on 4-byte-aligned addresses (SP must be 4-byte aligned; 8-byte alignment is required at function call boundaries per AAPCS). Attempting an unaligned stack access triggers a HardFault.

Multi-Register Ordering on the Stack

When you write PUSH {LR, R0, R1} the assembler does not push in the order you list them. The ARM architecture mandates that the register with the highest number is stored at the highest memory address. This rule ensures that POP can unambiguously restore the correct register regardless of the order in the instruction.

Rule: For a PUSH of {Rn, Rm, Rk} where n < m < k:
Rk is stored at the highest address, Rn at the lowest.

For the sequence PUSH {LR, R0, R1}:

  • LR (R14) has the highest register number → stored at the highest address after SP decrements three times.
  • R1 is stored at the middle address.
  • R0 is stored at the lowest address (lowest SP value, the actual new SP).

However, the slides and many real compiler outputs push registers individually in source order — PUSH LR, then PUSH R0, then PUSH R1 — which puts LR at the highest address (pushed first, SP decrements least before it is stored relative to the final SP), then R0 one slot lower, then R1 at the lowest address. This is what we trace in the worked example below. The end result is stack-layout-compatible because the matching POPs are also individual and in reverse order.

Worked Example: Register Initial State

Consider the following initial register state. These values represent a realistic embedded scenario: R0 and R1 hold function arguments (or local computed results), LR holds the return address placed there by a BL instruction that called the current function, R2 and R3 are “scratch” registers that will be overwritten, and PC points to the current instruction.

Initial Register State
R0  = 0x0000FFBC
R1  = 0x0000AB11
R2  = 0x00000000
R3  = 0x00000005
LR  = 0x800211CD
PC  = 0x80010000
SP  = [current top]
 
LR (R14) = return address — where the processor jumps back after this function returns.
“Previously pushed data” is whatever the caller already placed on the stack before calling us.

The instruction sequence to trace is:

PUSH LR      ; save return address
PUSH R0      ; save R0 (will be used for something else)
PUSH R1      ; save R1

; ... function body uses R0, R1, R2, R3 freely ...

POP  R2      ; restore top of stack into R2
POP  R3      ; restore next slot into R3
POP  PC      ; restore into PC — this is the function RETURN
Note: This is an unusual sequence — normally you POP the same registers you PUSHed. Here we POP into R2/R3/PC instead of R1/R0/LR. This is intentional to demonstrate that POP simply loads from wherever SP points — it doesn’t “know” which register originally contributed the value. The value from the top of the stack (last pushed) goes into R2, and the bottom (first pushed, LR) goes into PC, causing a function return.

Step-by-Step Stack Trace

Each step below shows the instruction executing (highlighted in red in the slides), the resulting stack layout, and the updated register state. The stack grows downward (Full Descending); SP always points to the last written (valid) word.

Step 0 — Before Any PUSH

Stack (↓ = lower addr)
SP→
prev pushed data
prev pushed data
prev pushed data
[empty]
[empty]
[empty]
Registers:
R0 = 0x0000FFBC
R1 = 0x0000AB11
LR = 0x800211CD
R2 = 0x00000000
R3 = 0x00000005
PC = 0x80010000
SP points to the last valid word pushed by the caller. The 3 empty slots below will be used next.

Step 1 — PUSH LR

The processor decrements SP by 4, then writes LR (0x800211CD) to the new SP address. SP now points to the slot containing 0x800211CD.

Stack after PUSH LR
prev pushed data
prev pushed data
prev pushed data
SP→
0x800211CD (LR)
[empty]
[empty]
Action:
SP = SP – 4
MEM[SP] = LR

Registers:
R0 = 0x0000FFBC
R1 = 0x0000AB11
LR = 0x800211CD ✓
R2 = 0x00000000
R3 = 0x00000005

Step 2 — PUSH R0

SP decrements by 4 again. R0 (0x0000FFBC) is written to the new SP. LR value on the stack is one slot above SP.

Stack after PUSH R0
prev pushed data
prev pushed data
prev pushed data
0x800211CD (LR)
SP→
0x0000FFBC (R0)
[empty]
Action:
SP = SP – 4
MEM[SP] = R0

SP now points to
0x0000FFBC slot.
LR slot is above at
SP+4.

Step 3 — PUSH R1

SP decrements by 4 one more time. R1 (0x0000AB11) is stored. Now three words are saved on the stack; SP points to the R1 slot (the most recently pushed, lowest address).

Stack after PUSH R1 (all 3 pushed)
prev pushed data
prev pushed data
prev pushed data
0x800211CD (LR) @ SP+8
0x0000FFBC (R0) @ SP+4
SP→
0x0000AB11 (R1) @ SP
Stack layout:
SP+8 = 0x800211CD (LR)
SP+4 = 0x0000FFBC (R0)
SP+0 = 0x0000AB11 (R1)

Function body can now
freely use R0,R1,LR
without losing their
values.

Step 4 — POP R2

POP reads the word at the current SP address and writes it to R2, then increments SP by 4. The top of the stack was 0x0000AB11 (originally R1), so R2 becomes 0x0000AB11.

Stack after POP R2
prev pushed data
prev pushed data
prev pushed data
0x800211CD (LR) @ SP+4
SP→
0x0000FFBC (R0) @ SP
0x0000AB11 (popped)
Action:
R2 = MEM[SP] → 0x0000AB11
SP = SP + 4

R2 = 0x0000AB11 ← was R1
(R2 changed from 0 to AB11)

SP now points to
the old R0 slot.

Step 5 — POP R3

POP reads from SP (which now holds 0x0000FFBC, originally R0) and loads it into R3, then SP increments by 4. R3 becomes 0x0000FFBC.

Stack after POP R3
prev pushed data
prev pushed data
prev pushed data
SP→
0x800211CD (LR) @ SP
0x0000FFBC (popped)
0x0000AB11 (popped)
Action:
R3 = MEM[SP] → 0x0000FFBC
SP = SP + 4

R2 = 0x0000AB11 (R1’s value)
R3 = 0x0000FFBC (R0’s value)

One word remains on stack:
the original LR value.

Step 6 — POP PC (Function Return)

The final POP targets PC directly. The processor reads 0x800211CD from the stack into PC. SP increments by 4, returning to exactly the same value it had before the three PUSHes — the stack is balanced. The processor immediately begins fetching instructions from 0x800211CD — this is the function return.

Stack after POP PC (balanced)
prev pushed data
prev pushed data
SP→
prev pushed data ← back here!
0x800211CD (popped→PC)
0x0000FFBC (popped)
0x0000AB11 (popped)
Final Register State:
R2 = 0x0000AB11
R3 = 0x0000FFBC
PC = 0x800211CD ← returned!

SP restored to original
value (stack balanced).

Execution resumes at
0x800211CD (caller site).
Key insight: POP PC is the most efficient return mechanism on ARM Cortex-M. It combines loading the return address AND branching to it in a single instruction. The alternative — POP {LR}; BX LR — requires two instructions and an extra cycle.

Complete State Summary

Step Instruction SP Change Memory Written/Read Register Changed
0 Initial — — —
1 PUSH LR SP − 4 MEM[new SP] ← 0x800211CD SP updated
2 PUSH R0 SP − 4 MEM[new SP] ← 0x0000FFBC SP updated
3 PUSH R1 SP − 4 MEM[new SP] ← 0x0000AB11 SP updated
4 POP R2 SP + 4 R2 ← MEM[old SP] = 0x0000AB11 R2 = 0x0000AB11 (was 0)
5 POP R3 SP + 4 R3 ← MEM[old SP] = 0x0000FFBC R3 = 0x0000FFBC (was 5)
6 POP PC SP + 4 PC ← MEM[old SP] = 0x800211CD PC = 0x800211CD (function returned)

The net SP change is zero: three decrements of 4 followed by three increments of 4. This is the definition of a balanced stack frame. Any imbalance causes the stack to drift, eventually corrupting data or triggering a fault.

AAPCS: Which Registers Must Be Saved?

The ARM Architecture Procedure Call Standard (AAPCS) — sometimes called the ARM calling convention — defines which registers a function is allowed to destroy (caller-saved) and which it must preserve (callee-saved). Understanding this tells you exactly which registers will appear in PUSH/POP sequences in real compiler output.

Registers AAPCS Role Saved by Notes
R0 – R3 Argument / scratch (caller-saved) Caller (if needed) Function arguments 1–4; return value in R0 (or R0:R1 for 64-bit). Destroyed freely by callee.
R4 – R11 Callee-saved (general purpose) Callee (this function) Must be preserved across a function call. If the function uses them, it must PUSH at entry and POP at exit.
R12 (IP) Intra-procedure scratch Neither (volatile) Used by linker veneers (long-range branch stubs). Treat as destroyed across calls.
R13 (SP) Stack pointer Callee Must be balanced (same value at exit as at entry, 8-byte aligned at call boundary).
R14 (LR) Link register Callee (if it calls others) Overwritten by BL. Must PUSH LR at start if function makes any sub-calls; POP PC to return.
R15 (PC) Program counter Hardware Loading PC causes a branch. POP PC = function return.
S0 – S15 / D0 – D7 FPU caller-saved Caller Single/double precision float scratch. Destroyed by callee freely.
S16 – S31 / D8 – D15 FPU callee-saved Callee Must be preserved if used. VPUSH/VPOP in function prologue/epilogue.
Leaf function optimisation: A function that makes no sub-calls (a “leaf”) never needs to PUSH LR because BL never overwrites it. The compiler can skip the PUSH LR / POP PC pair entirely and return with BX LR directly, saving two instructions and two stack accesses.

What the Compiler Actually Generates

Let’s look at a concrete C function and the assembly GCC (arm-none-eabi-gcc -O1 -mcpu=cortex-m4 -mthumb) generates for it, to see PUSH/POP in a real context.

C source

/* Non-leaf: calls another function, uses callee-saved R4 */
int process(int a, int b)
{
    int temp = a + b;          /* uses R0+R1 → R0 */
    int result = helper(temp); /* BL overwrites LR */
    return result + 1;
}

int helper(int x)
{
    return x * 2;              /* leaf: no BL inside */
}

Assembly output (annotated)

process:
    PUSH    {R4, LR}       ; R4 = callee-saved (will use it)
                           ; LR = must save because we call helper (BL)
    ADD     R4, R0, R1     ; temp = a + b  (R4 callee-saved, safe to keep across call)
    MOV     R0, R4         ; arg0 for helper = temp
    BL      helper         ; call helper → overwrites LR with return addr
    ADD     R0, R0, #1     ; result + 1
    POP     {R4, PC}       ; restore R4; POP PC = return (uses saved LR value)

helper:
    LSL     R0, R0, #1     ; x * 2 = x << 1
    BX      LR             ; leaf: return directly via LR, no PUSH/POP needed
Pattern rules:
1. If the function uses R4–R11, it PUSHes them at the top and POPs at the bottom.
2. If the function calls any sub-function, it must PUSH LR (because BL overwrites LR) and POP PC to return.
3. The two are often combined: PUSH {R4, R5, LR} / POP {R4, R5, PC} — an even number of registers ensures 8-byte alignment.

Why an even number of registers matters

AAPCS requires the stack pointer to be 8-byte aligned at the moment of any function call (the BL instruction). If you PUSH an odd number of 4-byte words, SP becomes only 4-byte aligned. The compiler automatically adds a padding register to the PUSH list (often a register it doesn’t actually need) just to ensure alignment.

; If only LR needed to save, compiler adds a dummy register:
PUSH {R3, LR}   ; R3 not actually used — just for 8-byte alignment
; body...
POP  {R3, PC}   ; R3 restored (value harmless), PC = return

PUSH/POP vs STM/LDM: Under the Hood

The ARM Thumb-2 encoding uses a unified instruction set where PUSH and POP are simply compact encodings of STMDB and LDMIA targeting SP with writeback. You can verify this by disassembling any Cortex-M binary — objdump will show STM/LDM variants in some contexts.

High-level form Equivalent STM/LDM form Operation
PUSH {R0, R1, LR} STMDB SP!, {R0, R1, LR} Decrement SP before each store; highest register to highest address
POP {R0, R1, PC} LDMIA SP!, {R0, R1, PC} Load from current SP then increment; lowest register from lowest address
PUSH {R4} STR R4, [SP, #-4]! Single register: pre-index STR with writeback (identical semantics)
POP {R4} LDR R4, [SP], #4 Single register: post-index LDR with writeback

Knowing the STM/LDM equivalents is useful when reading disassembly of optimised code, where the compiler may emit STMDB/LDMIA directly (especially in Cortex-M4 code with multiple register operations).

Stack Alignment and HardFault Prevention

ARM Cortex-M hardware enforces certain alignment rules on stack accesses. Violating them can lead to subtle, hard-to-reproduce bugs or outright HardFaults.

4-byte alignment — always required

Every PUSH/POP operates on 4-byte words. SP must be 4-byte aligned before any PUSH (it will remain 4-byte aligned since the decrement is always a multiple of 4). If SP is misaligned (e.g., points to an odd address), a UsageFault or HardFault fires immediately.

8-byte alignment — required at call boundaries

The AAPCS says SP must be 8-byte aligned when a BL or BLX executes. This allows the hardware to push exception frames (which are 8-byte-aligned by default on Cortex-M4 with FPU) without extra pad words. GCC enforces this automatically.

STKALIGN bit in CCR

The Configuration and Control Register (CCR, address 0xE000ED14) contains the STKALIGN bit (bit 9). When set (the power-on reset value on Cortex-M3/M4), the hardware automatically aligns the stack to 8 bytes when entering an exception handler, inserting a padding word if necessary. The xPSR bit 9 records whether padding was inserted so the hardware can remove it on exception return.

#define CCR   (*(volatile uint32_t *)0xE000ED14U)
#define STKALIGN_BIT  (1U << 9)

void check_stack_align(void)
{
    if (CCR & STKALIGN_BIT)
        /* hardware auto-aligns exception stack — good */;
    else
        /* only 4-byte aligned — potential issue with FPU context */;
}

Practical Patterns for Firmware Engineers

Pattern 1: Inline assembly PUSH/POP for critical section

/* Save and restore xPSR/PRIMASK around a critical section */
static inline uint32_t enter_critical(void)
{
    uint32_t primask;
    __asm volatile (
        "MRS %0, PRIMASK\n\t"   /* read current interrupt mask */
        "CPSID I\n\t"           /* disable all interrupts */
        : "=r" (primask)
        :
        : "memory"
    );
    return primask;             /* caller saves this on its stack */
}

static inline void exit_critical(uint32_t saved)
{
    __asm volatile (
        "MSR PRIMASK, %0\n\t"   /* restore original mask */
        "ISB\n\t"
        :
        : "r" (saved)
        : "memory"
    );
}

Pattern 2: Manually balanced PUSH/POP in assembly ISR

/* Pure assembly ISR — must balance SP exactly */
    .thumb_func
    .global USART2_IRQHandler
USART2_IRQHandler:
    PUSH {R4, R5, LR}        /* callee-saved + return addr */
    LDR  R4, =0x40004400     /* USART2 base address */
    LDR  R5, [R4, #0x00]     /* SR register */
    TST  R5, #0x20            /* check RXNE bit */
    BEQ  .no_data
    LDR  R0, [R4, #0x04]     /* DR register — read byte */
    BL   usart2_rx_callback   /* process it */
.no_data:
    POP  {R4, R5, PC}        /* restore + return from ISR */

Pattern 3: Checking SP programmatically

/* Read current SP value */
static inline uint32_t get_sp(void)
{
    uint32_t sp;
    __asm volatile ("MOV %0, SP" : "=r"(sp));
    return sp;
}

/* Assert SP is 4-byte aligned before a sensitive operation */
void assert_sp_aligned(void)
{
    uint32_t sp = get_sp();
    if (sp & 0x3U) {
        /* SP misaligned — trap immediately */
        __asm volatile ("BKPT #0");
    }
}

Pattern 4: Tail call with POP PC

/* When a function's last act is calling another,
   the compiler can convert BL + BX LR into B (branch),
   or the programmer can do it in assembly: */

foo:
    PUSH {LR}
    /* ... do work ... */
    POP  {LR}
    B    bar       /* tail call: jump to bar, which returns to foo's caller */
    /* equivalent to:
       BL bar
       BX LR    -- but saves a PUSH/POP */

Relation to Hardware Exception Stacking

The PUSH/POP mechanism is also what the Cortex-M hardware itself does when an interrupt fires, but it does it automatically without any instruction fetch. When an IRQ is accepted, the hardware pushes the exception frame to the current stack pointer — using the exact Full Descending model described above.

Hardware Exception Frame (pushed automatically)
SP+28
xPSR
pushed last (highest addr)
SP+24
PC
interrupted instruction + 1
SP+20
LR
return addr (or EXC_RETURN)
SP+16
R12
 
SP+12
R3
 
SP+8
R2
 
SP+4
R1
 
SP+0
R0
pushed first (lowest addr, new SP)
Hardware pushes R0–R3, R12, LR, PC, xPSR automatically (8 words = 32 bytes). IRQ handler runs with a clean register file — R0-R3/R12/LR/PC/xPSR already saved.

Because the hardware saves R0–R3 and R12, the ISR can use them freely without manual PUSH. Only if the ISR uses R4–R11 does it need to PUSH them manually at the start of the handler. This is why most short ISRs have no PUSH/POP at all — the hardware exception frame provides enough scratch space.

Frequently Asked Questions

Q: Can I POP into a different register than I PUSHed into?

Yes — as the worked example shows. POP simply reads from wherever SP points and writes to the destination register. The stack doesn’t remember which register produced each word. This is how tail-call optimisations and lazy context-save tricks work.

Q: Why does POP PC cause a branch but POP LR does not?

Because PC (R15) is the program counter — any write to PC immediately redirects instruction fetch to that address. LR (R14) is just a general register used by convention; writing to LR has no side effect until you explicitly do BX LR.

Q: What happens if PUSH and POP counts don’t match?

The stack becomes imbalanced — SP has a different value at function exit than entry. The function that called us will then execute with a corrupt SP, causing it to read garbage data on its next POP and eventually corrupt PC, leading to a HardFault or erratic behaviour. The AAPCS makes this a programmer obligation, not a hardware guarantee.

Q: Is PUSH {R0} safe inside an ISR?

Yes — the hardware exception frame has already saved R0-R3 before your ISR starts executing, but pushing R0 again in the ISR is perfectly legal. It just uses extra stack space. You might do this if you need to preserve R0 across a nested function call inside the ISR.

Q: Can I use PUSH/POP to pass data between an ISR and main code?

No — the stack is private to the current execution context. Data shared between ISRs and main code should use global variables declared volatile (or better, use a proper ring buffer with memory barriers). PUSH/POP data is only visible within the same call frame.

Q: What is the maximum PUSH register list size?

Thumb-2 PUSH can include any subset of R0–R12 plus LR in one instruction — up to 14 registers (56 bytes) in a single multi-register push. On Cortex-M4 with FPU, VPUSH can save up to 16 double-precision (S-registers) at once (128 bytes). These multi-register forms are encoded more compactly than individual push instructions.

Q: Does PUSH LR at the start of every function hurt performance?

It adds one memory write and one memory read per function call pair (a stack push and a pop). On Cortex-M4 with the flash cache hitting, this is typically 1–2 cycles. For performance-critical inner loops, write them as leaf functions so the compiler can skip PUSH LR entirely and use BX LR for free.

ARM PUSH POP STMDB LDMIA Cortex-M Stack AAPCS Calling Convention Function Prologue Epilogue Callee Saved Registers Full Descending Stack STM32F411 Embedded Systems Free ARM Course

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *