ARM Stack PUSH POP Operations
Master the ARM Cortex-M Full Descending stack through a complete worked example — PUSH LR/R0/R1, POP R2/R3/PC — with cycle-accurate register state tracking, AAPCS conventions, and compiler-generated prologue/epilogue patterns.
What PUSH and POP Actually Do
At their core, PUSH and POP are not special-purpose instructions unique to stacks — they are convenient aliases for the more general Store Multiple (STM) and Load Multiple (LDM) instructions operating through the stack pointer (SP / R13). Understanding them at the instruction level removes any mystery about how function calls preserve register state.
• STMDB = Store Multiple Decrement Before
• SP is decremented before each store (Full Descending rule)
POP {Rn, …} is equivalent to LDMIA SP!, {Rn, …}
• LDMIA = Load Multiple Increment After
• Each load reads SP’s current location, then SP is incremented
The ! suffix means “writeback” — the updated SP value is written back to the SP register after the operation completes. Without writeback the stack pointer would not actually move, which would be an error for any real push or pop.
On Cortex-M, PUSH and POP always operate in groups of 4-byte words and always on 4-byte-aligned addresses (SP must be 4-byte aligned; 8-byte alignment is required at function call boundaries per AAPCS). Attempting an unaligned stack access triggers a HardFault.
Multi-Register Ordering on the Stack
When you write PUSH {LR, R0, R1} the assembler does not push in the order you list them. The ARM architecture mandates that the register with the highest number is stored at the highest memory address. This rule ensures that POP can unambiguously restore the correct register regardless of the order in the instruction.
Rk is stored at the highest address, Rn at the lowest.
For the sequence PUSH {LR, R0, R1}:
- LR (R14) has the highest register number → stored at the highest address after SP decrements three times.
- R1 is stored at the middle address.
- R0 is stored at the lowest address (lowest SP value, the actual new SP).
However, the slides and many real compiler outputs push registers individually in source order — PUSH LR, then PUSH R0, then PUSH R1 — which puts LR at the highest address (pushed first, SP decrements least before it is stored relative to the final SP), then R0 one slot lower, then R1 at the lowest address. This is what we trace in the worked example below. The end result is stack-layout-compatible because the matching POPs are also individual and in reverse order.
Worked Example: Register Initial State
Consider the following initial register state. These values represent a realistic embedded scenario: R0 and R1 hold function arguments (or local computed results), LR holds the return address placed there by a BL instruction that called the current function, R2 and R3 are “scratch” registers that will be overwritten, and PC points to the current instruction.
R1 = 0x0000AB11
R2 = 0x00000000
R3 = 0x00000005
PC = 0x80010000
SP = [current top]
“Previously pushed data” is whatever the caller already placed on the stack before calling us.
The instruction sequence to trace is:
PUSH LR ; save return address
PUSH R0 ; save R0 (will be used for something else)
PUSH R1 ; save R1
; ... function body uses R0, R1, R2, R3 freely ...
POP R2 ; restore top of stack into R2
POP R3 ; restore next slot into R3
POP PC ; restore into PC — this is the function RETURN
Step-by-Step Stack Trace
Each step below shows the instruction executing (highlighted in red in the slides), the resulting stack layout, and the updated register state. The stack grows downward (Full Descending); SP always points to the last written (valid) word.
Step 0 — Before Any PUSH
R1 = 0x0000AB11
LR = 0x800211CD
R2 = 0x00000000
R3 = 0x00000005
PC = 0x80010000
Step 1 — PUSH LR
The processor decrements SP by 4, then writes LR (0x800211CD) to the new SP address. SP now points to the slot containing 0x800211CD.
R1 = 0x0000AB11
LR = 0x800211CD ✓
R2 = 0x00000000
R3 = 0x00000005
Step 2 — PUSH R0
SP decrements by 4 again. R0 (0x0000FFBC) is written to the new SP. LR value on the stack is one slot above SP.
SP now points to
0x0000FFBC slot.
LR slot is above at
SP+4.
Step 3 — PUSH R1
SP decrements by 4 one more time. R1 (0x0000AB11) is stored. Now three words are saved on the stack; SP points to the R1 slot (the most recently pushed, lowest address).
SP+4 = 0x0000FFBC (R0)
SP+0 = 0x0000AB11 (R1)
Function body can now
freely use R0,R1,LR
without losing their
values.
Step 4 — POP R2
POP reads the word at the current SP address and writes it to R2, then increments SP by 4. The top of the stack was 0x0000AB11 (originally R1), so R2 becomes 0x0000AB11.
R2 = 0x0000AB11 ← was R1
(R2 changed from 0 to AB11)
SP now points to
the old R0 slot.
Step 5 — POP R3
POP reads from SP (which now holds 0x0000FFBC, originally R0) and loads it into R3, then SP increments by 4. R3 becomes 0x0000FFBC.
R2 = 0x0000AB11 (R1’s value)
R3 = 0x0000FFBC (R0’s value)
One word remains on stack:
the original LR value.
Step 6 — POP PC (Function Return)
The final POP targets PC directly. The processor reads 0x800211CD from the stack into PC. SP increments by 4, returning to exactly the same value it had before the three PUSHes — the stack is balanced. The processor immediately begins fetching instructions from 0x800211CD — this is the function return.
SP restored to original
value (stack balanced).
Execution resumes at
0x800211CD (caller site).
Complete State Summary
| Step | Instruction | SP Change | Memory Written/Read | Register Changed |
|---|---|---|---|---|
| 0 | Initial | — | — | — |
| 1 | PUSH LR | SP − 4 | MEM[new SP] ← 0x800211CD | SP updated |
| 2 | PUSH R0 | SP − 4 | MEM[new SP] ← 0x0000FFBC | SP updated |
| 3 | PUSH R1 | SP − 4 | MEM[new SP] ← 0x0000AB11 | SP updated |
| 4 | POP R2 | SP + 4 | R2 ← MEM[old SP] = 0x0000AB11 | R2 = 0x0000AB11 (was 0) |
| 5 | POP R3 | SP + 4 | R3 ← MEM[old SP] = 0x0000FFBC | R3 = 0x0000FFBC (was 5) |
| 6 | POP PC | SP + 4 | PC ← MEM[old SP] = 0x800211CD | PC = 0x800211CD (function returned) |
The net SP change is zero: three decrements of 4 followed by three increments of 4. This is the definition of a balanced stack frame. Any imbalance causes the stack to drift, eventually corrupting data or triggering a fault.
AAPCS: Which Registers Must Be Saved?
The ARM Architecture Procedure Call Standard (AAPCS) — sometimes called the ARM calling convention — defines which registers a function is allowed to destroy (caller-saved) and which it must preserve (callee-saved). Understanding this tells you exactly which registers will appear in PUSH/POP sequences in real compiler output.
| Registers | AAPCS Role | Saved by | Notes |
|---|---|---|---|
| R0 – R3 | Argument / scratch (caller-saved) | Caller (if needed) | Function arguments 1–4; return value in R0 (or R0:R1 for 64-bit). Destroyed freely by callee. |
| R4 – R11 | Callee-saved (general purpose) | Callee (this function) | Must be preserved across a function call. If the function uses them, it must PUSH at entry and POP at exit. |
| R12 (IP) | Intra-procedure scratch | Neither (volatile) | Used by linker veneers (long-range branch stubs). Treat as destroyed across calls. |
| R13 (SP) | Stack pointer | Callee | Must be balanced (same value at exit as at entry, 8-byte aligned at call boundary). |
| R14 (LR) | Link register | Callee (if it calls others) | Overwritten by BL. Must PUSH LR at start if function makes any sub-calls; POP PC to return. |
| R15 (PC) | Program counter | Hardware | Loading PC causes a branch. POP PC = function return. |
| S0 – S15 / D0 – D7 | FPU caller-saved | Caller | Single/double precision float scratch. Destroyed by callee freely. |
| S16 – S31 / D8 – D15 | FPU callee-saved | Callee | Must be preserved if used. VPUSH/VPOP in function prologue/epilogue. |
What the Compiler Actually Generates
Let’s look at a concrete C function and the assembly GCC (arm-none-eabi-gcc -O1 -mcpu=cortex-m4 -mthumb) generates for it, to see PUSH/POP in a real context.
C source
/* Non-leaf: calls another function, uses callee-saved R4 */
int process(int a, int b)
{
int temp = a + b; /* uses R0+R1 → R0 */
int result = helper(temp); /* BL overwrites LR */
return result + 1;
}
int helper(int x)
{
return x * 2; /* leaf: no BL inside */
}
Assembly output (annotated)
process:
PUSH {R4, LR} ; R4 = callee-saved (will use it)
; LR = must save because we call helper (BL)
ADD R4, R0, R1 ; temp = a + b (R4 callee-saved, safe to keep across call)
MOV R0, R4 ; arg0 for helper = temp
BL helper ; call helper → overwrites LR with return addr
ADD R0, R0, #1 ; result + 1
POP {R4, PC} ; restore R4; POP PC = return (uses saved LR value)
helper:
LSL R0, R0, #1 ; x * 2 = x << 1
BX LR ; leaf: return directly via LR, no PUSH/POP needed
1. If the function uses R4–R11, it PUSHes them at the top and POPs at the bottom.
2. If the function calls any sub-function, it must PUSH LR (because BL overwrites LR) and POP PC to return.
3. The two are often combined: PUSH {R4, R5, LR} / POP {R4, R5, PC} — an even number of registers ensures 8-byte alignment.
Why an even number of registers matters
AAPCS requires the stack pointer to be 8-byte aligned at the moment of any function call (the BL instruction). If you PUSH an odd number of 4-byte words, SP becomes only 4-byte aligned. The compiler automatically adds a padding register to the PUSH list (often a register it doesn’t actually need) just to ensure alignment.
; If only LR needed to save, compiler adds a dummy register:
PUSH {R3, LR} ; R3 not actually used — just for 8-byte alignment
; body...
POP {R3, PC} ; R3 restored (value harmless), PC = return
PUSH/POP vs STM/LDM: Under the Hood
The ARM Thumb-2 encoding uses a unified instruction set where PUSH and POP are simply compact encodings of STMDB and LDMIA targeting SP with writeback. You can verify this by disassembling any Cortex-M binary — objdump will show STM/LDM variants in some contexts.
| High-level form | Equivalent STM/LDM form | Operation |
|---|---|---|
PUSH {R0, R1, LR} |
STMDB SP!, {R0, R1, LR} |
Decrement SP before each store; highest register to highest address |
POP {R0, R1, PC} |
LDMIA SP!, {R0, R1, PC} |
Load from current SP then increment; lowest register from lowest address |
PUSH {R4} |
STR R4, [SP, #-4]! |
Single register: pre-index STR with writeback (identical semantics) |
POP {R4} |
LDR R4, [SP], #4 |
Single register: post-index LDR with writeback |
Knowing the STM/LDM equivalents is useful when reading disassembly of optimised code, where the compiler may emit STMDB/LDMIA directly (especially in Cortex-M4 code with multiple register operations).
Stack Alignment and HardFault Prevention
ARM Cortex-M hardware enforces certain alignment rules on stack accesses. Violating them can lead to subtle, hard-to-reproduce bugs or outright HardFaults.
4-byte alignment — always required
Every PUSH/POP operates on 4-byte words. SP must be 4-byte aligned before any PUSH (it will remain 4-byte aligned since the decrement is always a multiple of 4). If SP is misaligned (e.g., points to an odd address), a UsageFault or HardFault fires immediately.
8-byte alignment — required at call boundaries
The AAPCS says SP must be 8-byte aligned when a BL or BLX executes. This allows the hardware to push exception frames (which are 8-byte-aligned by default on Cortex-M4 with FPU) without extra pad words. GCC enforces this automatically.
STKALIGN bit in CCR
The Configuration and Control Register (CCR, address 0xE000ED14) contains the STKALIGN bit (bit 9). When set (the power-on reset value on Cortex-M3/M4), the hardware automatically aligns the stack to 8 bytes when entering an exception handler, inserting a padding word if necessary. The xPSR bit 9 records whether padding was inserted so the hardware can remove it on exception return.
#define CCR (*(volatile uint32_t *)0xE000ED14U)
#define STKALIGN_BIT (1U << 9)
void check_stack_align(void)
{
if (CCR & STKALIGN_BIT)
/* hardware auto-aligns exception stack — good */;
else
/* only 4-byte aligned — potential issue with FPU context */;
}
Practical Patterns for Firmware Engineers
Pattern 1: Inline assembly PUSH/POP for critical section
/* Save and restore xPSR/PRIMASK around a critical section */
static inline uint32_t enter_critical(void)
{
uint32_t primask;
__asm volatile (
"MRS %0, PRIMASK\n\t" /* read current interrupt mask */
"CPSID I\n\t" /* disable all interrupts */
: "=r" (primask)
:
: "memory"
);
return primask; /* caller saves this on its stack */
}
static inline void exit_critical(uint32_t saved)
{
__asm volatile (
"MSR PRIMASK, %0\n\t" /* restore original mask */
"ISB\n\t"
:
: "r" (saved)
: "memory"
);
}
Pattern 2: Manually balanced PUSH/POP in assembly ISR
/* Pure assembly ISR — must balance SP exactly */
.thumb_func
.global USART2_IRQHandler
USART2_IRQHandler:
PUSH {R4, R5, LR} /* callee-saved + return addr */
LDR R4, =0x40004400 /* USART2 base address */
LDR R5, [R4, #0x00] /* SR register */
TST R5, #0x20 /* check RXNE bit */
BEQ .no_data
LDR R0, [R4, #0x04] /* DR register — read byte */
BL usart2_rx_callback /* process it */
.no_data:
POP {R4, R5, PC} /* restore + return from ISR */
Pattern 3: Checking SP programmatically
/* Read current SP value */
static inline uint32_t get_sp(void)
{
uint32_t sp;
__asm volatile ("MOV %0, SP" : "=r"(sp));
return sp;
}
/* Assert SP is 4-byte aligned before a sensitive operation */
void assert_sp_aligned(void)
{
uint32_t sp = get_sp();
if (sp & 0x3U) {
/* SP misaligned — trap immediately */
__asm volatile ("BKPT #0");
}
}
Pattern 4: Tail call with POP PC
/* When a function's last act is calling another,
the compiler can convert BL + BX LR into B (branch),
or the programmer can do it in assembly: */
foo:
PUSH {LR}
/* ... do work ... */
POP {LR}
B bar /* tail call: jump to bar, which returns to foo's caller */
/* equivalent to:
BL bar
BX LR -- but saves a PUSH/POP */
Relation to Hardware Exception Stacking
The PUSH/POP mechanism is also what the Cortex-M hardware itself does when an interrupt fires, but it does it automatically without any instruction fetch. When an IRQ is accepted, the hardware pushes the exception frame to the current stack pointer — using the exact Full Descending model described above.
Because the hardware saves R0–R3 and R12, the ISR can use them freely without manual PUSH. Only if the ISR uses R4–R11 does it need to PUSH them manually at the start of the handler. This is why most short ISRs have no PUSH/POP at all — the hardware exception frame provides enough scratch space.
Frequently Asked Questions
Yes — as the worked example shows. POP simply reads from wherever SP points and writes to the destination register. The stack doesn’t remember which register produced each word. This is how tail-call optimisations and lazy context-save tricks work.
Because PC (R15) is the program counter — any write to PC immediately redirects instruction fetch to that address. LR (R14) is just a general register used by convention; writing to LR has no side effect until you explicitly do BX LR.
The stack becomes imbalanced — SP has a different value at function exit than entry. The function that called us will then execute with a corrupt SP, causing it to read garbage data on its next POP and eventually corrupt PC, leading to a HardFault or erratic behaviour. The AAPCS makes this a programmer obligation, not a hardware guarantee.
Yes — the hardware exception frame has already saved R0-R3 before your ISR starts executing, but pushing R0 again in the ISR is perfectly legal. It just uses extra stack space. You might do this if you need to preserve R0 across a nested function call inside the ISR.
No — the stack is private to the current execution context. Data shared between ISRs and main code should use global variables declared volatile (or better, use a proper ring buffer with memory barriers). PUSH/POP data is only visible within the same call frame.
Thumb-2 PUSH can include any subset of R0–R12 plus LR in one instruction — up to 14 registers (56 bytes) in a single multi-register push. On Cortex-M4 with FPU, VPUSH can save up to 16 double-precision (S-registers) at once (128 bytes). These multi-register forms are encoded more compactly than individual push instructions.
It adds one memory write and one memory read per function call pair (a stack push and a pop). On Cortex-M4 with the flash cache hitting, this is typically 1–2 cycles. For performance-critical inner loops, write them as leaf functions so the compiler can skip PUSH LR entirely and use BX LR for free.

2 Comments