ARM GCC Inline Assembly:
Operands, Constraints & Pointers
Move from theory to practice. Learn exactly how GCC constraint characters, operand indexing, and modifier prefixes work — with four fully explained examples that compile and run on any STM32 Cortex-M target.
- Quick Recap — Inline Assembly Syntax
- Operand Indexing: %0, %1, %2…
- GCC Constraint Characters for ARM
- Constraint Modifiers: =, +, &
- Example 1 — Move C Variable to ARM Register
- Example 2 — Read CONTROL Register into C Variable
- Example 3 — Copy One C Variable to Another
- Example 4 — Dereference a Pointer in Assembly
- Practical Patterns Reference
- Common Mistakes
- FAQ
1. Quick Recap — Inline Assembly Syntax
Lecture 04 introduced the general form of ARM GCC inline assembly. Here is a quick reminder before we go deeper into how operands and constraints work.
__asm volatile (
"assembly instructions\n\t" /* code section */
: output operands /* written by asm */
: input operands /* read by asm */
: clobber list /* registers/mem modified */
);
Every section is separated by a colon :. If a section is empty you still write
the colon as a placeholder so GCC knows the structure. If there are no clobbers at all you
can stop after the input list.
Why volatile?
Without volatile, GCC may decide the assembly block has no visible side
effects and remove it during optimisation. Always use volatile when reading
or writing hardware registers, or when the instruction itself has a side effect
(e.g. CPSID I, DSB, WFI).
Newline + tab after each instruction
GCC concatenates all strings in the code section into a single assembler text. Adding
\n\t at the end of each instruction line ensures the assembler sees
properly separated lines — otherwise all instructions run together and assembly fails.
2. Operand Indexing: %0, %1, %2…
The assembly code string cannot directly name C variables. Instead, GCC replaces placeholder tokens with the actual register (or memory address) it assigns to each operand. These placeholders are written as a percent sign followed by a digit.
Operand Numbering Rules
Outputs are always numbered first, inputs second — regardless of their position in the code string.
/* Counts how GCC numbers operands */
__asm volatile (
"ADD %0, %1, %2\n\t" /* result = a + b */
: "=r" (result) /* %0 = output 0 */
: "r" (a), /* %1 = input 0 */
"r" (b) /* %2 = input 1 */
);
%0/%1 you can use symbolic names:
"ADD %[res], %[x], %[y]" with the operand list written as
[res] "=r"(result), [x] "r"(a), [y] "r"(b). Named operands make long
blocks far easier to read and maintain — use them whenever you have more than two operands.
The advantage of operand indexing is that GCC controls register allocation.
You write %0 and GCC picks any free register (say R3) and
substitutes it everywhere that %0 appears. This avoids conflicts with the rest
of the compiled C code.
3. GCC Constraint Characters for ARM
A constraint is a short string that tells GCC what kind of storage an operand
must live in. If the constraint is "r", GCC must put the value in a
general-purpose ARM register before the asm runs. Different constraints enable different
instruction encodings.
ARM has two execution states relevant here: ARM state (32-bit A32 instructions) and Thumb state (16-bit / 32-bit T32 instructions). Cortex-M processors run only in Thumb state, so the Thumb column in the table below is the one that applies to you.
| Constraint | Meaning in ARM State | Meaning in Thumb State (Cortex-M) |
|---|---|---|
r |
General-purpose register r0–r15 | Same — r0–r15 (most useful one) |
l |
Same as r |
Low registers r0–r7 only (needed by some 16-bit Thumb encodings) |
h |
Not available in ARM state | High registers r8–r15 |
I |
Immediate 0–255 (data-processing instructions) | Immediate 0–255 (e.g. SWI operand) |
J |
Indexing constants −4095 to 4095 | Immediate −255 to −1 |
K |
Same as I but inverted (ARM barrel-shift trick) | Same as I, shifted |
L |
Same as I but negated | Constant in range −7 to 7 |
M |
Constant 0–32 or a power of 2 (shift amounts) | Constant multiple of 4 in range 0–1020 |
m |
Any valid memory address | Any valid memory address (pointer dereference) |
N |
Not available | Constant 0–31 (bit-shift count) |
O |
Not available | Constant multiple of 4 in range −508 to 508 |
f |
Floating-point registers f0–f7 (legacy FPA) | Not available |
w |
VFP/NEON vector registers s0–s31 | Not available (use CMSIS intrinsics instead) |
G |
Immediate floating-point constant | Not available |
X |
Any operand | Any operand (rarely needed) |
"r" for almost everything.
The other constraints become relevant when writing heavily optimised tight loops or when
an instruction requires a specific immediate range that GCC must check at compile time.
Why constraint characters matter
Consider the ARM LSL (logical shift left) instruction:
/* This will fail at assemble time if val > 31: */
__asm volatile ("LSL %0, %0, %1" : "+r"(data) : "r"(shift_amount));
/* Better: tell GCC the shift must fit in 0..31 */
__asm volatile ("LSL %0, %0, %1" : "+r"(data) : "N"(shift_amount));
With "N", GCC produces a compile error if shift_amount is
not a constant in range 0–31, catching the bug before it reaches the target hardware.
With "r", the value is placed in a register and the assembler never checks
the range — the shift silently wraps or produces garbage at runtime.
4. Constraint Modifiers: =, +, &
A modifier is a character that precedes the constraint character to provide extra information about how the operand is used. Without a modifier, GCC assumes the operand is read-only (input) or write-only (output, always needed for outputs).
Write-only output
The operand is written to by the assembly. Its initial value is meaningless.
Used for all output operands that are not also inputs.
"=r"(result)
Read-write operand
The operand is both read and written. GCC loads it into a register before the
asm block and stores it back after. Must be listed in the output section.
"+r"(counter)
Early-clobber
Tells GCC this output register may be written before all inputs are consumed.
GCC must not allocate this register to any input. Prevents subtle data corruption
in multi-instruction blocks.
"=&r"(temp)
/* Without & — DANGEROUS if GCC reuses register */
__asm volatile (
"LDR %0, [%1]\n\t" /* read *ptr */
"STR %0, [%2]\n\t" /* write *dst */
: "=r" (tmp) /* output — %0 */
: "r" (ptr), /* input — %1 */
"r" (dst) /* input — %2 */
);
/* GCC may assign tmp and ptr to the same register!
After LDR, tmp overwrites what was ptr. STR then
uses garbage as the destination address. */
/* Correct version using early-clobber */
__asm volatile (
"LDR %0, [%1]\n\t"
"STR %0, [%2]\n\t"
: "=&r" (tmp) /* & = don't share with inputs */
: "r" (ptr),
"r" (dst)
);
&
to that output constraint. If you have a single-instruction block it is never needed.
5. Example 1 — Move C Variable to ARM Register
The simplest possible inline assembly task: take a value stored in a C variable and
load it into a named ARM register. The instruction is MOV, the source is a
C integer, and the destination is R0.
MOV
C variable
val (INPUT)R0 (ARM register)
int val = 50;
/* Move val into ARM register R0.
No output needed — we are just demonstrating the move. */
__asm volatile ("MOV R0, %0"
: /* no output operands */
: "r" (val) /* %0 = input operand; GCC loads val into a register */
);
Notice %0 in the instruction string refers to the first operand in the
combined list. Since there are no outputs, the first (and only) operand is the input
"r"(val).
What GCC Does Behind the Scenes
in SRAM
load val into Rn (e.g. R2)
Rn substituted for %0
ARM core register
GCC chose R2 for val — you wrote %0, GCC substituted R2. The MOV R0,R2 is the resulting instruction.
Why this is useful
In AAPCS (ARM Procedure Call Standard), R0 is the first function argument register.
Occasionally you need to ensure a specific value is in R0 before calling a function or
triggering a software interrupt. Inline assembly with the "r" constraint
is the correct way to do this without hard-coding a register that GCC might be using.
"MOV R0, %0" and R0 is already being used by GCC to hold some
other variable, you just silently corrupted it. Always add "R0" to the
clobber list when you explicitly name a register: : : "r"(val) : "R0"
/* Correct version — declares R0 as clobbered */
int val = 50;
__asm volatile ("MOV R0, %0"
:
: "r" (val)
: "R0" /* tell GCC we trash R0 */
);
6. Example 2 — Read CONTROL Register into C Variable
The CONTROL register is a non-memory-mapped ARM special register. You cannot
access it via a C pointer — there is no address for it. The only way to read it is with
the MRS (Move from Special Register to General Register) instruction.
MRS
CONTROL register (special reg)
C variable
control_reg (OUTPUT)uint32_t control_reg;
__asm volatile ("MRS %0, CONTROL"
: "=r" (control_reg) /* %0 = output; GCC picks a register, copies to control_reg after */
: /* no input operands */
: /* no clobbers */
);
Here the modifier is = (write-only output). GCC allocates a free register
(say R1), substitutes it for %0, producing MRS R1, CONTROL,
then generates a STR R1, [sp, #offset] to save it into the control_reg
variable on the stack.
MRS Instruction Flow
bit 0: nPRIV
bit 1: SPSEL
bit 2: FPCA
general register
C variable in SRAM
Practical use — check privilege level
#include <stdint.h>
static inline int is_privileged(void)
{
uint32_t ctrl;
__asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));
/* CONTROL bit 0 (nPRIV): 0 = privileged, 1 = unprivileged */
return (ctrl & 0x01U) == 0U;
}
void demo(void)
{
if (is_privileged()) {
/* Safe to access SCB, NVIC, MPU etc. */
} else {
/* Restricted — cannot touch privileged peripherals */
}
}
MSR (Move to Special Register) — writes a general-purpose register into a special register. Used when you want to change CONTROL, PRIMASK, etc.
/* Write to CONTROL: switch to unprivileged Thread mode */
static inline void drop_to_unprivileged(void)
{
uint32_t ctrl;
__asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));
ctrl |= 0x01U; /* set nPRIV bit */
__asm volatile ("MSR CONTROL, %0" : : "r"(ctrl) : "memory");
__asm volatile ("ISB"); /* instruction sync barrier — required after CONTROL write */
}
7. Example 3 — Copy One C Variable to Another
This example demonstrates both an input operand and an output operand in the
same asm block. The task is to copy the integer value of var1 into
var2 using a MOV instruction at the assembly level.
MOV
var1 — INPUT operand (%1)
var2 — OUTPUT operand (%0)
int var1 = 10;
int var2;
/* Copy var1 to var2 via MOV.
%0 = first operand (var2, output, numbered first).
%1 = second operand (var1, input). */
__asm ("MOV %0, %1"
: "=r" (var2) /* output: %0 */
: "r" (var1) /* input: %1 */
);
/* After this, var2 == 10 */
Why use inline assembly for something a C assignment can do? In real code you would not.
But this example teaches the two-operand pattern: the output operand is
always %0 and the input follows as %1. Understanding this
numbering is essential for more complex blocks.
How GCC Translates the Asm Block
GCC reused R3 for both operands since they do not overlap. With optimisation enabled, the whole thing may collapse into a single STR.
Extending to arithmetic — ADD with two inputs and one output
int a = 7, b = 3, result;
/* result = a + b using ADD instruction */
__asm volatile (
"ADD %[res], %[x], %[y]"
: [res] "=r" (result) /* named output operand */
: [x] "r" (a), /* named input operand */
[y] "r" (b) /* named input operand */
);
/* result == 10 */
8. Example 4 — Dereference a Pointer in Assembly
This is where inline assembly starts to feel genuinely useful for embedded work. Sometimes you need to read a value from a specific memory address — for example, a hardware register or a fixed SRAM location — and the C compiler’s volatile pointer dereference is not giving you the exact instruction sequence you need.
LDR
*p2 — memory at address held in p2
p1 — OUTPUT C variable
int p1;
int *p2;
/* Make p2 point to a specific SRAM address */
p2 = (int *)0x20000008;
/* p1 = *p2 — load the 32-bit word at address 0x20000008 into p1.
LDR %0, [%1] means: load the word at the address in register %1
into register %0. */
__asm volatile ("LDR %0, [%1]"
: "=r" (p1) /* output: GCC puts result into p1 after asm */
: "r" (p2) /* input: GCC puts the pointer value in a reg */
);
/* p1 now holds the 32-bit value stored at address 0x20000008 */
Memory Dereference via LDR
holds 0x20000008
32-bit word in memory
C variable updated
Real-world application: read a peripheral register
/* Read the GPIO port A input data register (IDR) at 0x40020010.
This is equivalent to: val = *((volatile uint32_t*)0x40020010) */
uint32_t gpioa_idr;
uint32_t *reg_ptr = (uint32_t *)0x40020010U; /* GPIOA_IDR */
__asm volatile ("LDR %0, [%1]"
: "=r" (gpioa_idr)
: "r" (reg_ptr)
);
/* Check pin PA5 */
if (gpioa_idr & (1U << 5)) {
/* PA5 is HIGH */
}
*((volatile uint32_t *)0x40020010) is fine and preferred.
The inline LDR becomes necessary when:
- You need the load/store to be part of a larger atomic asm sequence (e.g. LDREX/STREX)
- You need to control which exact addressing mode is used
- You are implementing low-level primitives like atomic compare-and-swap
STR — the opposite: write through a pointer
uint32_t value = 0xDEADBEEF;
uint32_t *dst = (uint32_t *)0x20000100;
/* Store value at address held in dst */
__asm volatile ("STR %1, [%0]"
: /* no C output — side effect goes to memory */
: "r" (dst), /* %0 = pointer */
"r" (value) /* %1 = data to write */
: "memory" /* tell GCC: memory has been modified */
);
"memory" in the clobber list, GCC might reorder loads/stores around
the asm block or keep a cached register value instead of re-reading from memory after
the asm. The "memory" clobber acts as a full compiler memory barrier.
9. Practical Patterns Reference
The following are ready-to-use inline assembly patterns for common Cortex-M tasks. Copy and use them directly — they have been validated for arm-none-eabi-gcc.
/* ---------------------------------------------------------------
Pattern 1: Read any special register into a uint32_t
Replace PRIMASK with: CONTROL, BASEPRI, FAULTMASK, PSP, MSP
--------------------------------------------------------------- */
static inline uint32_t read_special(void)
{
uint32_t val;
__asm volatile ("MRS %0, PRIMASK" : "=r"(val));
return val;
}
/* ---------------------------------------------------------------
Pattern 2: Write any special register from a uint32_t
--------------------------------------------------------------- */
static inline void write_special(uint32_t val)
{
__asm volatile ("MSR PRIMASK, %0" : : "r"(val) : "memory");
}
/* ---------------------------------------------------------------
Pattern 3: Disable and re-enable interrupts (critical section)
--------------------------------------------------------------- */
static inline uint32_t enter_critical(void)
{
uint32_t saved_primask;
__asm volatile (
"MRS %0, PRIMASK\n\t"
"CPSID I\n\t"
: "=r"(saved_primask)
:
: "memory"
);
return saved_primask;
}
static inline void exit_critical(uint32_t saved_primask)
{
__asm volatile (
"MSR PRIMASK, %0\n\t"
:
: "r"(saved_primask)
: "memory"
);
}
/* ---------------------------------------------------------------
Pattern 4: Memory barriers (required around DMA / peripheral I/O)
--------------------------------------------------------------- */
static inline void dsb(void) { __asm volatile ("DSB" ::: "memory"); }
static inline void dmb(void) { __asm volatile ("DMB" ::: "memory"); }
static inline void isb(void) { __asm volatile ("ISB" ::: "memory"); }
/* ---------------------------------------------------------------
Pattern 5: Count leading zeros (CLZ) — fast log2 / priority
--------------------------------------------------------------- */
static inline uint32_t clz(uint32_t val)
{
uint32_t result;
__asm volatile ("CLZ %0, %1" : "=r"(result) : "r"(val));
return result;
}
/* ---------------------------------------------------------------
Pattern 6: Reverse bytes in a 32-bit word (REV)
--------------------------------------------------------------- */
static inline uint32_t bswap32(uint32_t val)
{
uint32_t result;
__asm volatile ("REV %0, %1" : "=r"(result) : "r"(val));
return result;
}
/* ---------------------------------------------------------------
Pattern 7: Software breakpoint (halt debugger at this line)
--------------------------------------------------------------- */
static inline void bkpt(void)
{
__asm volatile ("BKPT #0");
}
/* ---------------------------------------------------------------
Pattern 8: Wait For Interrupt (low-power sleep)
--------------------------------------------------------------- */
static inline void wfi(void)
{
__asm volatile ("WFI" ::: "memory");
}
/* ---------------------------------------------------------------
Pattern 9: No-operation (delay or alignment padding)
--------------------------------------------------------------- */
static inline void nop(void)
{
__asm volatile ("NOP");
}
__get_PRIMASK(), __set_PRIMASK(), __disable_irq(),
__enable_irq(), __DSB(), __CLZ(), __BKPT(),
__WFI(), __NOP(). If CMSIS is available in your project, prefer
these — they are portable across GCC, Clang, IAR, and Keil, and are reviewed by ARM.
10. Common Mistakes
Forgetting \n\t between instructions
/* Wrong — assembler sees one long line */
__asm volatile (
"MOV R0, %0 MRS R1, PRIMASK"
: : "r"(val)
);
/* Correct */
__asm volatile (
"MOV R0, %0\n\t"
"MRS R1, PRIMASK\n\t"
: : "r"(val) : "R0","R1"
);
Missing volatile on side-effect asm
/* GCC may delete this at -O2! */
__asm ("CPSID I");
/* Correct */
__asm volatile ("CPSID I" ::: "memory");
Not clobbering modified registers
/* GCC doesn't know R4 is changed */
__asm volatile ("MOV R4, #0xFF");
/* Correct: declare R4 clobbered */
__asm volatile (
"MOV R4, #0xFF"
::: "R4"
);
Missing “memory” after memory writes
/* GCC may cache register values
after this store — BUG */
__asm volatile (
"STR %1,[%0]"
: : "r"(ptr),"r"(val)
);
/* Correct */
__asm volatile (
"STR %1,[%0]"
: : "r"(ptr),"r"(val)
: "memory"
);
Missing ISB after CONTROL write
/* Writing CONTROL without ISB can
cause unpredictable behaviour */
__asm volatile ("MSR CONTROL, %0" : : "r"(v));
/* Add this immediately after: */
__asm volatile ("ISB" ::: "memory");
Using wrong constraint for literal immediates
/* "r" puts value in a register,
but ORR expects an immediate here */
uint32_t mask = 0xFF;
__asm volatile (
"ORR %0, %0, %1"
: "+r"(reg) : "r"(mask) /* OK */
);
/* For compile-time constants, use "I"
to get direct encoding: */
__asm volatile (
"ORR %0, %0, #0xFF"
: "+r"(reg)
);
11. FAQ
"I" for a runtime
variable (not a compile-time constant) will produce a cryptic assembler error because
the constraint requires a literal number. Always test your asm blocks with optimisation
enabled (-O2) to catch issues early.
"memory" clobber is especially important to force GCC to re-read any
variables from memory after the asm block completes.
.s assembly file and calling it
as an external function — it will be far more readable.
__asm without volatile allows GCC to delete the block if it
determines the outputs are never used, or to move the block for scheduling purposes.
__asm volatile prevents both. For any block that has side effects beyond
writing to output operands — interrupt enable/disable, barrier instructions, WFI,
hardware register writes — always use volatile.
__get_PRIMASK(), __DSB(),
__CLZ() etc.) are reviewed by ARM engineers, work correctly across all
compilers, and are far easier to read. Use inline assembly when you need something CMSIS
cannot express.
"memory" clobber or missing
volatile. At -O0 GCC is conservative and re-reads everything
from memory constantly. At -O2 it aggressively caches values in registers
and skips blocks it considers side-effect-free. Fix: add volatile to every
asm block that writes hardware, and add "memory" to every block that reads
or writes memory that GCC does not know about through its operand list.
Next: Memory Map and Bus Architecture
Now that you can write assembly-level register access, the next lecture maps out the full Cortex-M memory space — flash, SRAM, peripherals, and system regions — and explains how the AHB/APB bus hierarchy determines which peripherals you can reach, and at what speed.

2 Comments