AAPCS ARM Calling Convention-Embedded Systems Training in Hyderabad

AAPCS ARM Calling Convention | Embedded Systems Course

EmbeddedPathashala — ARM Cortex-M Course Home All Lectures STM32 Projects

AAPCS ARM Calling Convention

The ARM Architecture Procedure Call Standard (AAPCS) defines the contract between every caller and every callee in ARM software. Master the register roles, argument passing rules, return value conventions, and see real compiler output that follows — or deliberately deviates from — the standard.

Home › ARM Cortex-M Course › Lecture 13 — AAPCS ARM Calling Convention

What Is AAPCS and Why Does It Exist?

The ARM Architecture Procedure Call Standard (AAPCS) is the ARM-specific part of the larger ARM ABI (Application Binary Interface). It was published by ARM Ltd. and is maintained as a living specification (current release: 2019Q1.1 at time of writing). Every C compiler targeting ARM — GCC, Clang/LLVM, IAR, Keil armcc — must follow AAPCS to produce code that interoperates correctly.

AAPCS Scope (from the specification):
The AAPCS defines how subroutines can be separately written, separately compiled, and separately assembled to work together. It describes a contract between a calling routine and a called routine that defines:
• Obligations on the caller to create a program state in which the called routine may start to execute.
• Obligations on the called routine to preserve the program state of the caller across the call.
• The rights of the called routine to alter the program state of its caller.

Without AAPCS, code compiled by GCC could not call a library function compiled by Keil, because each compiler might use registers in incompatible ways. AAPCS makes this “separate compilation” model work reliably. On STM32 projects, AAPCS is why your application code can call STM32 HAL functions (compiled separately by ST) without any glue code.

AAPCS operates at the register level. It does not specify memory layout, calling sequences beyond registers, or OS-specific ABI concerns. For Cortex-M it governs the 16 integer registers (R0–R15) and, when the FPU is present, the floating-point registers (S0–S31 / D0–D15).

Register Classification: Caller-Saved vs Callee-Saved

AAPCS divides the ARM core registers into two categories based on who is responsible for preserving their values across a function call:

ARM Register Roles under AAPCS
CALLER-SAVED (Volatile)
R0 — arg1 / return value
R1 — arg2 / return value hi
R2 — arg3
R3 — arg4
R12 — intra-procedure scratch
R14 (LR) — link register
xPSR — processor status
Callee may destroy these freely. If caller needs them after the call, caller must save them first.
CALLEE-SAVED (Non-volatile)
R4 — general purpose
R5 — general purpose
R6 — general purpose
R7 — general purpose
R8 — general purpose
R9 — general purpose (or PIC base)
R10 — general purpose
R11 — general purpose (or FP)
Callee MUST preserve these. If callee uses them, it must PUSH at start and POP at end.
SPECIAL-PURPOSE (Hardware-managed)
R13 (SP) — Stack pointer (MSP or PSP). Must be balanced across call. Must be 8-byte aligned at call boundary.
R15 (PC) — Program counter. Loaded by BL (call) and BX LR / POP PC (return). Not a general register.
Memory trick: Caller-saved = R0–R3, R12, LR (think: the registers involved directly in a call — arguments, scratch, and the link register). Callee-saved = R4–R11 (the “worker” registers a subroutine uses for its own long-lived computations).

Argument Passing: Up to 4 in Registers

When the caller invokes a function, it places arguments in registers R0, R1, R2, R3 in order. The callee reads its parameters from those same registers. No stack access is needed for up to four 32-bit arguments — the call is very fast.

Example: 4-argument function call

/* Caller (fun_x) calls callee (fun_y) with arguments 1, 2, 4, 5 */

/* C source */
void fun_x(void)
{
    int ret;
    ret = fun_y(1, 2, 4, 5);   /* four int arguments */
}

int fun_y(int a, int b, int c, int d)
{
    return (a + b + c + d);
}
/* Compiler output (arm-none-eabi-gcc -O1 -mthumb -mcpu=cortex-m4) */

fun_x:
    PUSH    {R3, LR}        ; 8-byte align + save LR (will call fun_y)
    MOV     R0, #1          ; arg1 a = 1  → R0
    MOV     R1, #2          ; arg2 b = 2  → R1
    MOV     R2, #4          ; arg3 c = 4  → R2
    MOV     R3, #5          ; arg4 d = 5  → R3
    BL      fun_y           ; call — LR ← return addr, PC ← fun_y
    ; R0 now contains the return value (12)
    POP     {R3, PC}        ; restore R3 (dummy for alignment), return

fun_y:
    ; R0=a=1, R1=b=2, R2=c=4, R3=d=5 on entry (no PUSH needed — leaf function)
    ADD     R0, R0, R1      ; R0 = a + b = 3
    ADD     R0, R0, R2      ; R0 = 3 + c = 7
    ADD     R0, R0, R3      ; R0 = 7 + d = 12
    BX      LR              ; return; R0 = 12 = result
Register State at the BL fun_y Instruction
R0
1
arg1 (a)
R1
2
arg2 (b)
R2
4
arg3 (c)
R3
5
arg4 (d)
After BL returns, R0 = 12 (the return value). R1, R2, R3 may have been destroyed by fun_y — they are caller-saved.

What happens with more than 4 arguments?

When a function takes more than 4 integer arguments, the first 4 go in R0–R3 as usual. Arguments 5, 6, … are pushed onto the stack by the caller before the BL, in reverse order (last argument pushed first). The callee accesses them via SP-relative loads.

/* Six-argument function */
int add6(int a, int b, int c, int d, int e, int f)
{
    return a + b + c + d + e + f;
}

void caller(void)
{
    int r = add6(1, 2, 3, 4, 5, 6);
}

/* Compiler output for caller: */
caller:
    PUSH    {R7, LR}
    SUB     SP, SP, #8         ; make room for args 5 and 6 on stack
    MOV     R7, #6             ; f = 6
    STR     R7, [SP, #4]       ; push f onto stack (arg 6, higher address)
    MOV     R7, #5
    STR     R7, [SP, #0]       ; push e onto stack (arg 5, lower address)
    MOV     R0, #1             ; a → R0
    MOV     R1, #2             ; b → R1
    MOV     R2, #3             ; c → R2
    MOV     R3, #4             ; d → R3
    BL      add6
    ADD     SP, SP, #8         ; clean up stack args
    POP     {R7, PC}

Return Values

The callee returns its result in registers, not on the stack. AAPCS defines the return register(s) based on the result size:

Return type Register(s) used Example
8-bit / 16-bit / 32-bit integer, pointer R0 int foo(), char *bar()
64-bit integer (long long, uint64_t) R0 (low 32 bits) + R1 (high 32 bits) uint64_t get_timestamp()
32-bit float (float) — soft-float ABI R0 Float bits packed as integer
32-bit float (float) — hard-float ABI S0 Cortex-M4F with -mfloat-abi=hard
64-bit float (double) — hard-float ABI D0 (= S0:S1 pair) Cortex-M4F with -mfloat-abi=hard
Struct ≤ 4 bytes R0 Packed into a single register
Struct > 4 bytes Caller passes hidden pointer in R0; callee writes result there Large structs returned by reference

64-bit return example

#include <stdint.h>

uint64_t make_u64(uint32_t hi, uint32_t lo)
{
    return ((uint64_t)hi << 32) | lo;
}

/* Compiler output:
   make_u64 receives hi in R0, lo in R1.
   It must return the 64-bit result in R0 (low) and R1 (high).

   make_u64:
       MOV  R2, R0       ; R2 = hi
       MOV  R0, R1       ; R0 = lo  (low 32 bits of result)
       MOV  R1, R2       ; R1 = hi  (high 32 bits of result)
       BX   LR
*/

Full Example: Callee Saving R4–R11

The following example shows a function that uses R4 and R5 for its own computation. Because R4 and R5 are callee-saved, the function must save them at the start and restore them at the end, regardless of what values the caller placed there.

/* Compute dot product: sum of a[i]*b[i] for i=0..3 */
int dot4(const int *a, const int *b)
{
    int s0 = a[0] * b[0];
    int s1 = a[1] * b[1];
    int s2 = a[2] * b[2];
    int s3 = a[3] * b[3];
    return s0 + s1 + s2 + s3;
}
/* Annotated GCC -O1 output */
dot4:
    ; R0 = pointer to a[], R1 = pointer to b[]
    PUSH    {R4, R5, LR}    ; save R4, R5 (will use them), LR (non-leaf)

    LDR     R2, [R0, #0]    ; R2 = a[0]
    LDR     R3, [R1, #0]    ; R3 = b[0]
    MUL     R4, R2, R3      ; R4 = a[0]*b[0]   ← R4 used here, already saved

    LDR     R2, [R0, #4]    ; R2 = a[1]
    LDR     R3, [R1, #4]    ; R3 = b[1]
    MUL     R5, R2, R3      ; R5 = a[1]*b[1]   ← R5 used here, already saved
    ADD     R4, R4, R5      ; R4 = s0 + s1

    LDR     R2, [R0, #8]
    LDR     R3, [R1, #8]
    MUL     R5, R2, R3      ; R5 = a[2]*b[2]
    ADD     R4, R4, R5      ; R4 = s0+s1+s2

    LDR     R2, [R0, #12]
    LDR     R3, [R1, #12]
    MUL     R5, R2, R3      ; R5 = a[3]*b[3]
    ADD     R0, R4, R5      ; R0 = total = return value

    POP     {R4, R5, PC}    ; restore R4 and R5 to caller's values; return
The caller that invoked dot4 can rely on R4 and R5 being unchanged after the call. Even though dot4 used them extensively inside, it restored them before returning. This is the AAPCS guarantee in action.

What a Function May Freely Modify

According to AAPCS, a C function can freely modify the following without saving them first:

  • R0, R1, R2, R3 — argument/scratch registers. Always treated as destroyed after any function call.
  • R12 (IP) — intra-procedure call scratch register. Used by compiler-generated veneers (long-range branch stubs). Treat as clobbered by any BL.
  • R14 (LR) — the BL instruction itself overwrites LR with the return address. If a function makes any sub-call, its original LR is gone unless saved first.
  • xPSR condition flags — any instruction that affects flags (CMP, ADD with S suffix, etc.) modifies xPSR. The called function has no obligation to restore the N/Z/C/V flags.
Common bug: Calling a function inside an if condition that relies on flags set before the call.
/* WRONG: CMP sets flags, then function call clobbers them */
__asm volatile ("CMP R0, #0");
some_function();   /* destroys N/Z/C/V flags! */
__asm volatile ("BEQ label");   /* reads stale flags — undefined behaviour */

/* CORRECT: read the compare result into a register before the call */
int cmp_result = (value == 0);
some_function();
if (cmp_result) { ... }

AAPCS in Inline Assembly: The Clobber List

When you write GCC inline assembly, the clobber list tells the compiler which registers and memory your assembly code destroys. This is AAPCS reasoning applied at the instruction level — you are explicitly informing the compiler which registers it must not rely on after your inline block.

/* Example: inline assembly that uses R0, R1 and modifies memory */
void atomic_set_bit(volatile uint32_t *reg, int bit)
{
    __asm volatile (
        "LDR  R0, [%0]      \n\t"   /* read register */
        "ORR  R0, R0, %1    \n\t"   /* set bit */
        "STR  R0, [%0]      \n\t"   /* write back */
        :                            /* no output operands */
        : "r"(reg), "r"(1U << bit)  /* input: reg pointer, bit mask */
        : "R0", "memory"             /* clobbers: R0 (scratch), memory */
    );
}

/* Better version using named operands and constraints: */
void atomic_set_bit_v2(volatile uint32_t *reg, uint32_t mask)
{
    uint32_t tmp;
    __asm volatile (
        "LDR  %[t], [%[r]]  \n\t"
        "ORR  %[t], %[t], %[m] \n\t"
        "STR  %[t], [%[r]]  \n\t"
        : [t] "=&r" (tmp)            /* output: early-clobber temp */
        : [r]  "r"  (reg),
          [m]  "r"  (mask)
        : "memory"
    );
}

AAPCS with FPU: Hard-Float vs Soft-Float ABI

The Cortex-M4F (and Cortex-M33/M55) includes a hardware floating-point unit. AAPCS defines two ABI variants for FPU-equipped cores:

ABI variant GCC flag Float args/return Performance
Soft-float -mfloat-abi=soft Integer registers (R0–R3). FPU not used at all. Slowest. Software emulation for all float ops.
Softfp -mfloat-abi=softfp Integer registers (R0–R3). FPU used for computation but not argument passing. Compatible with soft-float libraries; partial speed.
Hard-float -mfloat-abi=hard S0–S15 / D0–D7 for args, S0/D0 for return. Full FPU ABI. Fastest. No integer ↔ float register moves needed.
Mixing ABI warning: Never link soft-float compiled code with hard-float compiled code. The function signatures are register-incompatible. If you call a soft-float library from hard-float application code, you must use -mfloat-abi=softfp throughout, or use wrapper functions that convert between the two conventions.

Hard-float caller-saved and callee-saved FPU registers

FPU Registers AAPCS Role Who saves?
S0–S15 / D0–D7 Caller-saved (volatile) Caller saves if needed across a call
S16–S31 / D8–D15 Callee-saved (non-volatile) Callee saves with VPUSH / restores with VPOP
FPSCR Caller-saved Condition flags / rounding mode destroyed by callee

AAPCS Practical Checklist for Firmware Engineers

When writing a C function:

  • Use R4–R11 for variables that must survive across sub-calls; the compiler handles PUSH/POP automatically.
  • Keep frequently-called functions as leaves (no sub-calls) to eliminate PUSH LR / POP PC overhead.
  • Pass up to 4 arguments by value; use a struct pointer for more to keep the call in registers.
  • For 64-bit return values, declare the return type as uint64_t; the compiler packs it into R0:R1.

When writing assembly that calls C functions:

my_asm_caller:
    PUSH {R4, LR}           ; save R4 (callee-saved) and LR (non-leaf)
    MOV  R0, #42            ; first argument in R0
    BL   c_function         ; call — clobbers R0-R3, R12, LR
    ; DO NOT read R1, R2, R3 after BL — they may be garbage
    MOV  R4, R0             ; save return value in R4 (safe across next call)
    MOV  R0, R4             ; pass it as argument to next function
    BL   another_c_func
    POP  {R4, PC}           ; restore R4 to caller's value; return

When writing C that calls assembly:

/* Declare the assembly function with correct C prototype —
   AAPCS governs the interface */
extern int asm_multiply(int a, int b);  /* a→R0, b→R1, return→R0 */

void use_asm(void)
{
    int result = asm_multiply(7, 8);    /* compiler sets R0=7, R1=8 automatically */
    /* result is in R0 = 56 */
}

/* In assembly: */
/* asm_multiply.s */
    .thumb_func
    .global asm_multiply
asm_multiply:
    MUL R0, R0, R1    /* R0 = a * b */
    BX  LR             /* return R0 */

Frequently Asked Questions

Q: Does the compiler always follow AAPCS, or can it optimise away register saves?

The compiler must follow AAPCS at the ABI boundary — externally visible functions (no static, exported symbols). For static functions called only within one translation unit, the compiler may use custom calling conventions (e.g., pass arguments in R4–R7 if it can prove no external code calls the function). This is called “intra-procedure call optimisation.” You will see this in -O2/-O3 output sometimes.

Q: Why is R12 designated as “intra-procedure call scratch” (IP)?

When a BL target is more than 16 MB away (impossible in Thumb-2 with 24-bit offset), the linker inserts a “veneer” — a small trampoline routine. The veneer uses R12 to load the full 32-bit target address without disturbing any AAPCS-protected register. Since R12 is caller-saved, the caller expects it to be destroyed, so the veneer is free to use it. This is why R12 is sometimes called the “linker scratch register.”

Q: Must I declare all clobbered registers in inline assembly?

Yes, if you write assembly that modifies a register not listed as an output operand. Failing to declare a clobber allows the compiler to keep live values in that register across your inline block, leading to silent data corruption. Always list any register you modify in the clobber section: : “R4”, “memory” etc.

Q: Why does GCC sometimes push R3 even when it doesn’t use R3?

To maintain 8-byte stack alignment at the BL boundary. AAPCS requires SP to be 8-byte aligned when BL executes. If only one register (e.g., LR) needs saving, that gives an odd number of 4-byte words on the stack. GCC adds a dummy register (often R3) to make it an even number — two words = 8 bytes. The value pushed for R3 is meaningless; only the alignment matters.

Q: Can AAPCS be violated intentionally?

Yes, with care. The __attribute__((regparm)) or custom calling conventions are GCC extensions. However, violating AAPCS is only safe for functions that will never be called from separately compiled code or from ISRs entered via the hardware vector table. Any public API must be AAPCS-compliant.

Q: What happens to PSR flags when I call a function?

They can be destroyed. If you set a flag with CMP or TST and then call any function before reading the flag, the function may change N/Z/C/V. Never place a function call between a flag-setting instruction and the branch that depends on it. Store the comparison result in a register or variable before the call.

AAPCS ARM ABI Caller Saved Registers Callee Saved Registers ARM Argument Passing R0 R1 R2 R3 R4 to R11 Hard Float Soft Float Cortex-M4 FPU Free Embedded Course

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *