AAPCS ARM Calling Convention
The ARM Architecture Procedure Call Standard (AAPCS) defines the contract between every caller and every callee in ARM software. Master the register roles, argument passing rules, return value conventions, and see real compiler output that follows — or deliberately deviates from — the standard.
What Is AAPCS and Why Does It Exist?
The ARM Architecture Procedure Call Standard (AAPCS) is the ARM-specific part of the larger ARM ABI (Application Binary Interface). It was published by ARM Ltd. and is maintained as a living specification (current release: 2019Q1.1 at time of writing). Every C compiler targeting ARM — GCC, Clang/LLVM, IAR, Keil armcc — must follow AAPCS to produce code that interoperates correctly.
The AAPCS defines how subroutines can be separately written, separately compiled, and separately assembled to work together. It describes a contract between a calling routine and a called routine that defines:
• Obligations on the caller to create a program state in which the called routine may start to execute.
• Obligations on the called routine to preserve the program state of the caller across the call.
• The rights of the called routine to alter the program state of its caller.
Without AAPCS, code compiled by GCC could not call a library function compiled by Keil, because each compiler might use registers in incompatible ways. AAPCS makes this “separate compilation” model work reliably. On STM32 projects, AAPCS is why your application code can call STM32 HAL functions (compiled separately by ST) without any glue code.
AAPCS operates at the register level. It does not specify memory layout, calling sequences beyond registers, or OS-specific ABI concerns. For Cortex-M it governs the 16 integer registers (R0–R15) and, when the FPU is present, the floating-point registers (S0–S31 / D0–D15).
Register Classification: Caller-Saved vs Callee-Saved
AAPCS divides the ARM core registers into two categories based on who is responsible for preserving their values across a function call:
R1 — arg2 / return value hi
R2 — arg3
R3 — arg4
R12 — intra-procedure scratch
R14 (LR) — link register
xPSR — processor status
R5 — general purpose
R6 — general purpose
R7 — general purpose
R8 — general purpose
R9 — general purpose (or PIC base)
R10 — general purpose
R11 — general purpose (or FP)
R15 (PC) — Program counter. Loaded by BL (call) and BX LR / POP PC (return). Not a general register.
Argument Passing: Up to 4 in Registers
When the caller invokes a function, it places arguments in registers R0, R1, R2, R3 in order. The callee reads its parameters from those same registers. No stack access is needed for up to four 32-bit arguments — the call is very fast.
Example: 4-argument function call
/* Caller (fun_x) calls callee (fun_y) with arguments 1, 2, 4, 5 */
/* C source */
void fun_x(void)
{
int ret;
ret = fun_y(1, 2, 4, 5); /* four int arguments */
}
int fun_y(int a, int b, int c, int d)
{
return (a + b + c + d);
}
/* Compiler output (arm-none-eabi-gcc -O1 -mthumb -mcpu=cortex-m4) */
fun_x:
PUSH {R3, LR} ; 8-byte align + save LR (will call fun_y)
MOV R0, #1 ; arg1 a = 1 → R0
MOV R1, #2 ; arg2 b = 2 → R1
MOV R2, #4 ; arg3 c = 4 → R2
MOV R3, #5 ; arg4 d = 5 → R3
BL fun_y ; call — LR ← return addr, PC ← fun_y
; R0 now contains the return value (12)
POP {R3, PC} ; restore R3 (dummy for alignment), return
fun_y:
; R0=a=1, R1=b=2, R2=c=4, R3=d=5 on entry (no PUSH needed — leaf function)
ADD R0, R0, R1 ; R0 = a + b = 3
ADD R0, R0, R2 ; R0 = 3 + c = 7
ADD R0, R0, R3 ; R0 = 7 + d = 12
BX LR ; return; R0 = 12 = result
What happens with more than 4 arguments?
When a function takes more than 4 integer arguments, the first 4 go in R0–R3 as usual. Arguments 5, 6, … are pushed onto the stack by the caller before the BL, in reverse order (last argument pushed first). The callee accesses them via SP-relative loads.
/* Six-argument function */
int add6(int a, int b, int c, int d, int e, int f)
{
return a + b + c + d + e + f;
}
void caller(void)
{
int r = add6(1, 2, 3, 4, 5, 6);
}
/* Compiler output for caller: */
caller:
PUSH {R7, LR}
SUB SP, SP, #8 ; make room for args 5 and 6 on stack
MOV R7, #6 ; f = 6
STR R7, [SP, #4] ; push f onto stack (arg 6, higher address)
MOV R7, #5
STR R7, [SP, #0] ; push e onto stack (arg 5, lower address)
MOV R0, #1 ; a → R0
MOV R1, #2 ; b → R1
MOV R2, #3 ; c → R2
MOV R3, #4 ; d → R3
BL add6
ADD SP, SP, #8 ; clean up stack args
POP {R7, PC}
Return Values
The callee returns its result in registers, not on the stack. AAPCS defines the return register(s) based on the result size:
| Return type | Register(s) used | Example |
|---|---|---|
| 8-bit / 16-bit / 32-bit integer, pointer | R0 | int foo(), char *bar() |
64-bit integer (long long, uint64_t) |
R0 (low 32 bits) + R1 (high 32 bits) | uint64_t get_timestamp() |
32-bit float (float) — soft-float ABI |
R0 | Float bits packed as integer |
32-bit float (float) — hard-float ABI |
S0 | Cortex-M4F with -mfloat-abi=hard |
64-bit float (double) — hard-float ABI |
D0 (= S0:S1 pair) | Cortex-M4F with -mfloat-abi=hard |
| Struct ≤ 4 bytes | R0 | Packed into a single register |
| Struct > 4 bytes | Caller passes hidden pointer in R0; callee writes result there | Large structs returned by reference |
64-bit return example
#include <stdint.h>
uint64_t make_u64(uint32_t hi, uint32_t lo)
{
return ((uint64_t)hi << 32) | lo;
}
/* Compiler output:
make_u64 receives hi in R0, lo in R1.
It must return the 64-bit result in R0 (low) and R1 (high).
make_u64:
MOV R2, R0 ; R2 = hi
MOV R0, R1 ; R0 = lo (low 32 bits of result)
MOV R1, R2 ; R1 = hi (high 32 bits of result)
BX LR
*/
Full Example: Callee Saving R4–R11
The following example shows a function that uses R4 and R5 for its own computation. Because R4 and R5 are callee-saved, the function must save them at the start and restore them at the end, regardless of what values the caller placed there.
/* Compute dot product: sum of a[i]*b[i] for i=0..3 */
int dot4(const int *a, const int *b)
{
int s0 = a[0] * b[0];
int s1 = a[1] * b[1];
int s2 = a[2] * b[2];
int s3 = a[3] * b[3];
return s0 + s1 + s2 + s3;
}
/* Annotated GCC -O1 output */
dot4:
; R0 = pointer to a[], R1 = pointer to b[]
PUSH {R4, R5, LR} ; save R4, R5 (will use them), LR (non-leaf)
LDR R2, [R0, #0] ; R2 = a[0]
LDR R3, [R1, #0] ; R3 = b[0]
MUL R4, R2, R3 ; R4 = a[0]*b[0] ← R4 used here, already saved
LDR R2, [R0, #4] ; R2 = a[1]
LDR R3, [R1, #4] ; R3 = b[1]
MUL R5, R2, R3 ; R5 = a[1]*b[1] ← R5 used here, already saved
ADD R4, R4, R5 ; R4 = s0 + s1
LDR R2, [R0, #8]
LDR R3, [R1, #8]
MUL R5, R2, R3 ; R5 = a[2]*b[2]
ADD R4, R4, R5 ; R4 = s0+s1+s2
LDR R2, [R0, #12]
LDR R3, [R1, #12]
MUL R5, R2, R3 ; R5 = a[3]*b[3]
ADD R0, R4, R5 ; R0 = total = return value
POP {R4, R5, PC} ; restore R4 and R5 to caller's values; return
What a Function May Freely Modify
According to AAPCS, a C function can freely modify the following without saving them first:
- R0, R1, R2, R3 — argument/scratch registers. Always treated as destroyed after any function call.
- R12 (IP) — intra-procedure call scratch register. Used by compiler-generated veneers (long-range branch stubs). Treat as clobbered by any BL.
- R14 (LR) — the BL instruction itself overwrites LR with the return address. If a function makes any sub-call, its original LR is gone unless saved first.
- xPSR condition flags — any instruction that affects flags (CMP, ADD with S suffix, etc.) modifies xPSR. The called function has no obligation to restore the N/Z/C/V flags.
/* WRONG: CMP sets flags, then function call clobbers them */
__asm volatile ("CMP R0, #0");
some_function(); /* destroys N/Z/C/V flags! */
__asm volatile ("BEQ label"); /* reads stale flags — undefined behaviour */
/* CORRECT: read the compare result into a register before the call */
int cmp_result = (value == 0);
some_function();
if (cmp_result) { ... }
AAPCS in Inline Assembly: The Clobber List
When you write GCC inline assembly, the clobber list tells the compiler which registers and memory your assembly code destroys. This is AAPCS reasoning applied at the instruction level — you are explicitly informing the compiler which registers it must not rely on after your inline block.
/* Example: inline assembly that uses R0, R1 and modifies memory */
void atomic_set_bit(volatile uint32_t *reg, int bit)
{
__asm volatile (
"LDR R0, [%0] \n\t" /* read register */
"ORR R0, R0, %1 \n\t" /* set bit */
"STR R0, [%0] \n\t" /* write back */
: /* no output operands */
: "r"(reg), "r"(1U << bit) /* input: reg pointer, bit mask */
: "R0", "memory" /* clobbers: R0 (scratch), memory */
);
}
/* Better version using named operands and constraints: */
void atomic_set_bit_v2(volatile uint32_t *reg, uint32_t mask)
{
uint32_t tmp;
__asm volatile (
"LDR %[t], [%[r]] \n\t"
"ORR %[t], %[t], %[m] \n\t"
"STR %[t], [%[r]] \n\t"
: [t] "=&r" (tmp) /* output: early-clobber temp */
: [r] "r" (reg),
[m] "r" (mask)
: "memory"
);
}
AAPCS with FPU: Hard-Float vs Soft-Float ABI
The Cortex-M4F (and Cortex-M33/M55) includes a hardware floating-point unit. AAPCS defines two ABI variants for FPU-equipped cores:
| ABI variant | GCC flag | Float args/return | Performance |
|---|---|---|---|
| Soft-float | -mfloat-abi=soft |
Integer registers (R0–R3). FPU not used at all. | Slowest. Software emulation for all float ops. |
| Softfp | -mfloat-abi=softfp |
Integer registers (R0–R3). FPU used for computation but not argument passing. | Compatible with soft-float libraries; partial speed. |
| Hard-float | -mfloat-abi=hard |
S0–S15 / D0–D7 for args, S0/D0 for return. Full FPU ABI. | Fastest. No integer ↔ float register moves needed. |
Hard-float caller-saved and callee-saved FPU registers
| FPU Registers | AAPCS Role | Who saves? |
|---|---|---|
| S0–S15 / D0–D7 | Caller-saved (volatile) | Caller saves if needed across a call |
| S16–S31 / D8–D15 | Callee-saved (non-volatile) | Callee saves with VPUSH / restores with VPOP |
| FPSCR | Caller-saved | Condition flags / rounding mode destroyed by callee |
AAPCS Practical Checklist for Firmware Engineers
When writing a C function:
- Use R4–R11 for variables that must survive across sub-calls; the compiler handles PUSH/POP automatically.
- Keep frequently-called functions as leaves (no sub-calls) to eliminate PUSH LR / POP PC overhead.
- Pass up to 4 arguments by value; use a struct pointer for more to keep the call in registers.
- For 64-bit return values, declare the return type as uint64_t; the compiler packs it into R0:R1.
When writing assembly that calls C functions:
my_asm_caller:
PUSH {R4, LR} ; save R4 (callee-saved) and LR (non-leaf)
MOV R0, #42 ; first argument in R0
BL c_function ; call — clobbers R0-R3, R12, LR
; DO NOT read R1, R2, R3 after BL — they may be garbage
MOV R4, R0 ; save return value in R4 (safe across next call)
MOV R0, R4 ; pass it as argument to next function
BL another_c_func
POP {R4, PC} ; restore R4 to caller's value; return
When writing C that calls assembly:
/* Declare the assembly function with correct C prototype —
AAPCS governs the interface */
extern int asm_multiply(int a, int b); /* a→R0, b→R1, return→R0 */
void use_asm(void)
{
int result = asm_multiply(7, 8); /* compiler sets R0=7, R1=8 automatically */
/* result is in R0 = 56 */
}
/* In assembly: */
/* asm_multiply.s */
.thumb_func
.global asm_multiply
asm_multiply:
MUL R0, R0, R1 /* R0 = a * b */
BX LR /* return R0 */
Frequently Asked Questions
The compiler must follow AAPCS at the ABI boundary — externally visible functions (no static, exported symbols). For static functions called only within one translation unit, the compiler may use custom calling conventions (e.g., pass arguments in R4–R7 if it can prove no external code calls the function). This is called “intra-procedure call optimisation.” You will see this in -O2/-O3 output sometimes.
When a BL target is more than 16 MB away (impossible in Thumb-2 with 24-bit offset), the linker inserts a “veneer” — a small trampoline routine. The veneer uses R12 to load the full 32-bit target address without disturbing any AAPCS-protected register. Since R12 is caller-saved, the caller expects it to be destroyed, so the veneer is free to use it. This is why R12 is sometimes called the “linker scratch register.”
Yes, if you write assembly that modifies a register not listed as an output operand. Failing to declare a clobber allows the compiler to keep live values in that register across your inline block, leading to silent data corruption. Always list any register you modify in the clobber section: : “R4”, “memory” etc.
To maintain 8-byte stack alignment at the BL boundary. AAPCS requires SP to be 8-byte aligned when BL executes. If only one register (e.g., LR) needs saving, that gives an odd number of 4-byte words on the stack. GCC adds a dummy register (often R3) to make it an even number — two words = 8 bytes. The value pushed for R3 is meaningless; only the alignment matters.
Yes, with care. The __attribute__((regparm)) or custom calling conventions are GCC extensions. However, violating AAPCS is only safe for functions that will never be called from separately compiled code or from ISRs entered via the hardware vector table. Any public API must be AAPCS-compliant.
They can be destroyed. If you set a flag with CMP or TST and then call any function before reading the flag, the function may change N/Z/C/V. Never place a function call between a flag-setting instruction and the branch that depends on it. Store the comparison result in a register or variable before the call.

1 Comment