ARM GCC Inline Assembly:Operands, Constraints & Pointers-Embedded Systems Training in Hyderabad

ARM Inline Assembly Operands and Constraints | Cortex-M Programming

Lecture 05 — ARM Cortex-M Programming

ARM GCC Inline Assembly:
Operands, Constraints & Pointers

Move from theory to practice. Learn exactly how GCC constraint characters, operand indexing, and modifier prefixes work — with four fully explained examples that compile and run on any STM32 Cortex-M target.

Covers source pages 41–50  |  STM32F411  |  arm-none-eabi-gcc

1. Quick Recap — Inline Assembly Syntax

Lecture 04 introduced the general form of ARM GCC inline assembly. Here is a quick reminder before we go deeper into how operands and constraints work.

__asm volatile (
    "assembly instructions\n\t"   /* code section   */
    : output operands             /* written by asm */
    : input operands              /* read by asm    */
    : clobber list                /* registers/mem modified */
);

Every section is separated by a colon :. If a section is empty you still write the colon as a placeholder so GCC knows the structure. If there are no clobbers at all you can stop after the input list.

VOLATILE

Why volatile?

Without volatile, GCC may decide the assembly block has no visible side effects and remove it during optimisation. Always use volatile when reading or writing hardware registers, or when the instruction itself has a side effect (e.g. CPSID I, DSB, WFI).

IMPORTANT

Newline + tab after each instruction

GCC concatenates all strings in the code section into a single assembler text. Adding \n\t at the end of each instruction line ensures the assembler sees properly separated lines — otherwise all instructions run together and assembly fails.

2. Operand Indexing: %0, %1, %2…

The assembly code string cannot directly name C variables. Instead, GCC replaces placeholder tokens with the actual register (or memory address) it assigns to each operand. These placeholders are written as a percent sign followed by a digit.

Operand Numbering Rules

%0
First operand in the combined list (outputs listed first, then inputs)
%1
Second operand — counting continues across both output and input sections
%N
N+1-th operand. Limit is 30 operands per asm block in practice

Outputs are always numbered first, inputs second — regardless of their position in the code string.

/* Counts how GCC numbers operands */

__asm volatile (
    "ADD %0, %1, %2\n\t"   /* result = a + b */
    : "=r" (result)         /* %0 = output 0  */
    : "r"  (a),             /* %1 = input 0   */
      "r"  (b)              /* %2 = input 1   */
);
Named operands (GCC extension)
Instead of %0/%1 you can use symbolic names: "ADD %[res], %[x], %[y]" with the operand list written as [res] "=r"(result), [x] "r"(a), [y] "r"(b). Named operands make long blocks far easier to read and maintain — use them whenever you have more than two operands.

The advantage of operand indexing is that GCC controls register allocation. You write %0 and GCC picks any free register (say R3) and substitutes it everywhere that %0 appears. This avoids conflicts with the rest of the compiled C code.

3. GCC Constraint Characters for ARM

A constraint is a short string that tells GCC what kind of storage an operand must live in. If the constraint is "r", GCC must put the value in a general-purpose ARM register before the asm runs. Different constraints enable different instruction encodings.

ARM has two execution states relevant here: ARM state (32-bit A32 instructions) and Thumb state (16-bit / 32-bit T32 instructions). Cortex-M processors run only in Thumb state, so the Thumb column in the table below is the one that applies to you.

Constraint Meaning in ARM State Meaning in Thumb State (Cortex-M)
r General-purpose register r0–r15 Same — r0–r15 (most useful one)
l Same as r Low registers r0–r7 only (needed by some 16-bit Thumb encodings)
h Not available in ARM state High registers r8–r15
I Immediate 0–255 (data-processing instructions) Immediate 0–255 (e.g. SWI operand)
J Indexing constants −4095 to 4095 Immediate −255 to −1
K Same as I but inverted (ARM barrel-shift trick) Same as I, shifted
L Same as I but negated Constant in range −7 to 7
M Constant 0–32 or a power of 2 (shift amounts) Constant multiple of 4 in range 0–1020
m Any valid memory address Any valid memory address (pointer dereference)
N Not available Constant 0–31 (bit-shift count)
O Not available Constant multiple of 4 in range −508 to 508
f Floating-point registers f0–f7 (legacy FPA) Not available
w VFP/NEON vector registers s0–s31 Not available (use CMSIS intrinsics instead)
G Immediate floating-point constant Not available
X Any operand Any operand (rarely needed)
90% Rule
For Cortex-M embedded work, you will use "r" for almost everything. The other constraints become relevant when writing heavily optimised tight loops or when an instruction requires a specific immediate range that GCC must check at compile time.

Why constraint characters matter

Consider the ARM LSL (logical shift left) instruction:

/* This will fail at assemble time if val > 31: */
__asm volatile ("LSL %0, %0, %1" : "+r"(data) : "r"(shift_amount));

/* Better: tell GCC the shift must fit in 0..31 */
__asm volatile ("LSL %0, %0, %1" : "+r"(data) : "N"(shift_amount));

With "N", GCC produces a compile error if shift_amount is not a constant in range 0–31, catching the bug before it reaches the target hardware. With "r", the value is placed in a register and the assembler never checks the range — the shift silently wraps or produces garbage at runtime.

4. Constraint Modifiers: =, +, &

A modifier is a character that precedes the constraint character to provide extra information about how the operand is used. Without a modifier, GCC assumes the operand is read-only (input) or write-only (output, always needed for outputs).

= (equals)

Write-only output

The operand is written to by the assembly. Its initial value is meaningless. Used for all output operands that are not also inputs.

"=r"(result)

+ (plus)

Read-write operand

The operand is both read and written. GCC loads it into a register before the asm block and stores it back after. Must be listed in the output section.

"+r"(counter)

& (ampersand)

Early-clobber

Tells GCC this output register may be written before all inputs are consumed. GCC must not allocate this register to any input. Prevents subtle data corruption in multi-instruction blocks.

"=&r"(temp)

/* Without & — DANGEROUS if GCC reuses register */
__asm volatile (
    "LDR %0, [%1]\n\t"   /* read *ptr */
    "STR %0, [%2]\n\t"   /* write *dst */
    : "=r" (tmp)          /* output — %0 */
    : "r"  (ptr),         /* input  — %1 */
      "r"  (dst)          /* input  — %2 */
);
/* GCC may assign tmp and ptr to the same register!
   After LDR, tmp overwrites what was ptr. STR then
   uses garbage as the destination address.           */

/* Correct version using early-clobber */
__asm volatile (
    "LDR %0, [%1]\n\t"
    "STR %0, [%2]\n\t"
    : "=&r" (tmp)         /* & = don't share with inputs */
    : "r"   (ptr),
      "r"   (dst)
);
Rule of thumb for &
Whenever your asm block writes an output register in an instruction that appears before the last instruction that reads an input register, add & to that output constraint. If you have a single-instruction block it is never needed.

5. Example 1 — Move C Variable to ARM Register

The simplest possible inline assembly task: take a value stored in a C variable and load it into a named ARM register. The instruction is MOV, the source is a C integer, and the destination is R0.

Instruction
MOV
Source
C variable val (INPUT)
Destination
R0 (ARM register)
int val = 50;

/* Move val into ARM register R0.
   No output needed — we are just demonstrating the move. */
__asm volatile ("MOV R0, %0"
    :               /* no output operands */
    : "r" (val)     /* %0 = input operand; GCC loads val into a register */
);

Notice %0 in the instruction string refers to the first operand in the combined list. Since there are no outputs, the first (and only) operand is the input "r"(val).

What GCC Does Behind the Scenes

val = 50
in SRAM
→
LDR Rn, [addr]
load val into Rn (e.g. R2)
→
MOV R0, R2
Rn substituted for %0
→
R0 = 50
ARM core register

GCC chose R2 for val — you wrote %0, GCC substituted R2. The MOV R0,R2 is the resulting instruction.

Why this is useful

In AAPCS (ARM Procedure Call Standard), R0 is the first function argument register. Occasionally you need to ensure a specific value is in R0 before calling a function or triggering a software interrupt. Inline assembly with the "r" constraint is the correct way to do this without hard-coding a register that GCC might be using.

Do not hard-code register names without clobbers
If you write "MOV R0, %0" and R0 is already being used by GCC to hold some other variable, you just silently corrupted it. Always add "R0" to the clobber list when you explicitly name a register: : : "r"(val) : "R0"
/* Correct version — declares R0 as clobbered */
int val = 50;
__asm volatile ("MOV R0, %0"
    :
    : "r" (val)
    : "R0"              /* tell GCC we trash R0 */
);

6. Example 2 — Read CONTROL Register into C Variable

The CONTROL register is a non-memory-mapped ARM special register. You cannot access it via a C pointer — there is no address for it. The only way to read it is with the MRS (Move from Special Register to General Register) instruction.

Instruction
MRS
Source
CONTROL register (special reg)
Destination
C variable control_reg (OUTPUT)
uint32_t control_reg;

__asm volatile ("MRS %0, CONTROL"
    : "=r" (control_reg)   /* %0 = output; GCC picks a register, copies to control_reg after */
    :                      /* no input operands */
    :                      /* no clobbers */
);

Here the modifier is = (write-only output). GCC allocates a free register (say R1), substitutes it for %0, producing MRS R1, CONTROL, then generates a STR R1, [sp, #offset] to save it into the control_reg variable on the stack.

MRS Instruction Flow

CONTROL
bit 0: nPRIV
bit 1: SPSEL
bit 2: FPCA
→ MRS R1, CONTROL →
R1
general register
→ STR R1, [sp,#n] →
control_reg
C variable in SRAM

Practical use — check privilege level

#include <stdint.h>

static inline int is_privileged(void)
{
    uint32_t ctrl;
    __asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));

    /* CONTROL bit 0 (nPRIV): 0 = privileged, 1 = unprivileged */
    return (ctrl & 0x01U) == 0U;
}

void demo(void)
{
    if (is_privileged()) {
        /* Safe to access SCB, NVIC, MPU etc. */
    } else {
        /* Restricted — cannot touch privileged peripherals */
    }
}
MRS vs MSR
MRS (Move from Special Register) — reads a special register into a general-purpose register (and then into a C variable via the output constraint).
MSR (Move to Special Register) — writes a general-purpose register into a special register. Used when you want to change CONTROL, PRIMASK, etc.
/* Write to CONTROL: switch to unprivileged Thread mode */
static inline void drop_to_unprivileged(void)
{
    uint32_t ctrl;
    __asm volatile ("MRS %0, CONTROL" : "=r"(ctrl));
    ctrl |= 0x01U;   /* set nPRIV bit */
    __asm volatile ("MSR CONTROL, %0" : : "r"(ctrl) : "memory");
    __asm volatile ("ISB");   /* instruction sync barrier — required after CONTROL write */
}

7. Example 3 — Copy One C Variable to Another

This example demonstrates both an input operand and an output operand in the same asm block. The task is to copy the integer value of var1 into var2 using a MOV instruction at the assembly level.

Instruction
MOV
Source
var1 — INPUT operand (%1)
Destination
var2 — OUTPUT operand (%0)
int var1 = 10;
int var2;

/* Copy var1 to var2 via MOV.
   %0 = first operand (var2, output, numbered first).
   %1 = second operand (var1, input).                 */
__asm ("MOV %0, %1"
    : "=r" (var2)    /* output: %0 */
    : "r"  (var1)    /* input:  %1 */
);

/* After this, var2 == 10 */

Why use inline assembly for something a C assignment can do? In real code you would not. But this example teaches the two-operand pattern: the output operand is always %0 and the input follows as %1. Understanding this numbering is essential for more complex blocks.

How GCC Translates the Asm Block

; C: int var1=10, var2;
; var1 is at [sp, #0], var2 is at [sp, #4]

LDR R3, [sp, #0] ; load var1 into R3 (satisfies “r”(var1) → %1)
MOV R3, R3 ; the actual inline asm instruction (%0=R3, %1=R3)
STR R3, [sp, #4] ; store R3 into var2 (satisfies “=r”(var2) → %0)

GCC reused R3 for both operands since they do not overlap. With optimisation enabled, the whole thing may collapse into a single STR.

Extending to arithmetic — ADD with two inputs and one output

int a = 7, b = 3, result;

/* result = a + b using ADD instruction */
__asm volatile (
    "ADD %[res], %[x], %[y]"
    : [res] "=r" (result)   /* named output operand */
    : [x]   "r"  (a),       /* named input operand  */
      [y]   "r"  (b)        /* named input operand  */
);

/* result == 10 */

8. Example 4 — Dereference a Pointer in Assembly

This is where inline assembly starts to feel genuinely useful for embedded work. Sometimes you need to read a value from a specific memory address — for example, a hardware register or a fixed SRAM location — and the C compiler’s volatile pointer dereference is not giving you the exact instruction sequence you need.

Instruction
LDR
Source
*p2 — memory at address held in p2
Destination
p1 — OUTPUT C variable
int  p1;
int *p2;

/* Make p2 point to a specific SRAM address */
p2 = (int *)0x20000008;

/* p1 = *p2 — load the 32-bit word at address 0x20000008 into p1.
   LDR %0, [%1] means: load the word at the address in register %1
   into register %0.                                               */
__asm volatile ("LDR %0, [%1]"
    : "=r" (p1)    /* output: GCC puts result into p1 after asm */
    : "r"  (p2)    /* input:  GCC puts the pointer value in a reg */
);

/* p1 now holds the 32-bit value stored at address 0x20000008 */

Memory Dereference via LDR

p2
holds 0x20000008
→ R1 ←
SRAM[0x20000008]
32-bit word in memory
LDR R0, [R1]
R0 → p1
C variable updated

Real-world application: read a peripheral register

/* Read the GPIO port A input data register (IDR) at 0x40020010.
   This is equivalent to: val = *((volatile uint32_t*)0x40020010) */

uint32_t gpioa_idr;
uint32_t *reg_ptr = (uint32_t *)0x40020010U;   /* GPIOA_IDR */

__asm volatile ("LDR %0, [%1]"
    : "=r" (gpioa_idr)
    : "r"  (reg_ptr)
);

/* Check pin PA5 */
if (gpioa_idr & (1U << 5)) {
    /* PA5 is HIGH */
}
When is this better than a volatile pointer?
In most cases *((volatile uint32_t *)0x40020010) is fine and preferred. The inline LDR becomes necessary when:
  • You need the load/store to be part of a larger atomic asm sequence (e.g. LDREX/STREX)
  • You need to control which exact addressing mode is used
  • You are implementing low-level primitives like atomic compare-and-swap

STR — the opposite: write through a pointer

uint32_t value = 0xDEADBEEF;
uint32_t *dst  = (uint32_t *)0x20000100;

/* Store value at address held in dst */
__asm volatile ("STR %1, [%0]"
    :                     /* no C output — side effect goes to memory */
    : "r"  (dst),         /* %0 = pointer */
      "r"  (value)        /* %1 = data to write */
    : "memory"            /* tell GCC: memory has been modified */
);
Always add “memory” clobber for memory-modifying asm
Without "memory" in the clobber list, GCC might reorder loads/stores around the asm block or keep a cached register value instead of re-reading from memory after the asm. The "memory" clobber acts as a full compiler memory barrier.

9. Practical Patterns Reference

The following are ready-to-use inline assembly patterns for common Cortex-M tasks. Copy and use them directly — they have been validated for arm-none-eabi-gcc.

/* ---------------------------------------------------------------
   Pattern 1: Read any special register into a uint32_t
   Replace PRIMASK with: CONTROL, BASEPRI, FAULTMASK, PSP, MSP
   --------------------------------------------------------------- */
static inline uint32_t read_special(void)
{
    uint32_t val;
    __asm volatile ("MRS %0, PRIMASK" : "=r"(val));
    return val;
}

/* ---------------------------------------------------------------
   Pattern 2: Write any special register from a uint32_t
   --------------------------------------------------------------- */
static inline void write_special(uint32_t val)
{
    __asm volatile ("MSR PRIMASK, %0" : : "r"(val) : "memory");
}

/* ---------------------------------------------------------------
   Pattern 3: Disable and re-enable interrupts (critical section)
   --------------------------------------------------------------- */
static inline uint32_t enter_critical(void)
{
    uint32_t saved_primask;
    __asm volatile (
        "MRS   %0, PRIMASK\n\t"
        "CPSID I\n\t"
        : "=r"(saved_primask)
        :
        : "memory"
    );
    return saved_primask;
}

static inline void exit_critical(uint32_t saved_primask)
{
    __asm volatile (
        "MSR PRIMASK, %0\n\t"
        :
        : "r"(saved_primask)
        : "memory"
    );
}

/* ---------------------------------------------------------------
   Pattern 4: Memory barriers (required around DMA / peripheral I/O)
   --------------------------------------------------------------- */
static inline void dsb(void) { __asm volatile ("DSB" ::: "memory"); }
static inline void dmb(void) { __asm volatile ("DMB" ::: "memory"); }
static inline void isb(void) { __asm volatile ("ISB" ::: "memory"); }

/* ---------------------------------------------------------------
   Pattern 5: Count leading zeros (CLZ) — fast log2 / priority
   --------------------------------------------------------------- */
static inline uint32_t clz(uint32_t val)
{
    uint32_t result;
    __asm volatile ("CLZ %0, %1" : "=r"(result) : "r"(val));
    return result;
}

/* ---------------------------------------------------------------
   Pattern 6: Reverse bytes in a 32-bit word (REV)
   --------------------------------------------------------------- */
static inline uint32_t bswap32(uint32_t val)
{
    uint32_t result;
    __asm volatile ("REV %0, %1" : "=r"(result) : "r"(val));
    return result;
}

/* ---------------------------------------------------------------
   Pattern 7: Software breakpoint (halt debugger at this line)
   --------------------------------------------------------------- */
static inline void bkpt(void)
{
    __asm volatile ("BKPT #0");
}

/* ---------------------------------------------------------------
   Pattern 8: Wait For Interrupt (low-power sleep)
   --------------------------------------------------------------- */
static inline void wfi(void)
{
    __asm volatile ("WFI" ::: "memory");
}

/* ---------------------------------------------------------------
   Pattern 9: No-operation (delay or alignment padding)
   --------------------------------------------------------------- */
static inline void nop(void)
{
    __asm volatile ("NOP");
}
CMSIS equivalents
ARM CMSIS-Core provides intrinsic functions for all of the above patterns: __get_PRIMASK(), __set_PRIMASK(), __disable_irq(), __enable_irq(), __DSB(), __CLZ(), __BKPT(), __WFI(), __NOP(). If CMSIS is available in your project, prefer these — they are portable across GCC, Clang, IAR, and Keil, and are reviewed by ARM.

10. Common Mistakes

MISTAKE 1

Forgetting \n\t between instructions

/* Wrong — assembler sees one long line */
__asm volatile (
  "MOV R0, %0 MRS R1, PRIMASK"
  : : "r"(val)
);

/* Correct */
__asm volatile (
  "MOV R0, %0\n\t"
  "MRS R1, PRIMASK\n\t"
  : : "r"(val) : "R0","R1"
);
MISTAKE 2

Missing volatile on side-effect asm

/* GCC may delete this at -O2! */
__asm ("CPSID I");

/* Correct */
__asm volatile ("CPSID I" ::: "memory");
MISTAKE 3

Not clobbering modified registers

/* GCC doesn't know R4 is changed */
__asm volatile ("MOV R4, #0xFF");

/* Correct: declare R4 clobbered */
__asm volatile (
  "MOV R4, #0xFF"
  ::: "R4"
);
MISTAKE 4

Missing “memory” after memory writes

/* GCC may cache register values
   after this store — BUG */
__asm volatile (
  "STR %1,[%0]"
  : : "r"(ptr),"r"(val)
);

/* Correct */
__asm volatile (
  "STR %1,[%0]"
  : : "r"(ptr),"r"(val)
  : "memory"
);
MISTAKE 5

Missing ISB after CONTROL write

/* Writing CONTROL without ISB can
   cause unpredictable behaviour   */
__asm volatile ("MSR CONTROL, %0" : : "r"(v));
/* Add this immediately after: */
__asm volatile ("ISB" ::: "memory");
MISTAKE 6

Using wrong constraint for literal immediates

/* "r" puts value in a register,
   but ORR expects an immediate here */
uint32_t mask = 0xFF;
__asm volatile (
  "ORR %0, %0, %1"
  : "+r"(reg) : "r"(mask)  /* OK */
);
/* For compile-time constants, use "I"
   to get direct encoding: */
__asm volatile (
  "ORR %0, %0, #0xFF"
  : "+r"(reg)
);

11. FAQ

Q: What happens if I use the wrong constraint character?
GCC will either reject your code with an error like “impossible constraint”, or silently generate incorrect code. For example, using "I" for a runtime variable (not a compile-time constant) will produce a cryptic assembler error because the constraint requires a literal number. Always test your asm blocks with optimisation enabled (-O2) to catch issues early.
Q: Can I mix C statements with inline assembly in the same function?
Yes, and that is the entire point of inline assembly. GCC handles register allocation so that the assembly block does not accidentally destroy values that the surrounding C code is using — as long as you correctly declare all clobbers. The "memory" clobber is especially important to force GCC to re-read any variables from memory after the asm block completes.
Q: Is there a limit to how many operands I can have?
GCC supports up to 30 operands per asm block in practice (limited by the maximum number of registers). Beyond 10 operands, consider splitting the block into multiple smaller asm blocks or writing a separate .s assembly file and calling it as an external function — it will be far more readable.
Q: What is the difference between __asm and __asm volatile?
__asm without volatile allows GCC to delete the block if it determines the outputs are never used, or to move the block for scheduling purposes. __asm volatile prevents both. For any block that has side effects beyond writing to output operands — interrupt enable/disable, barrier instructions, WFI, hardware register writes — always use volatile.
Q: When should I prefer CMSIS intrinsics over inline assembly?
Almost always — unless you need an instruction sequence that CMSIS does not expose, or you need very fine-grained control over register allocation in a performance-critical loop. CMSIS intrinsics (__get_PRIMASK(), __DSB(), __CLZ() etc.) are reviewed by ARM engineers, work correctly across all compilers, and are far easier to read. Use inline assembly when you need something CMSIS cannot express.
Q: Why does my asm block work at -O0 but break at -O2?
The most common cause is a missing "memory" clobber or missing volatile. At -O0 GCC is conservative and re-reads everything from memory constantly. At -O2 it aggressively caches values in registers and skips blocks it considers side-effect-free. Fix: add volatile to every asm block that writes hardware, and add "memory" to every block that reads or writes memory that GCC does not know about through its operand list.

Next: Memory Map and Bus Architecture

Now that you can write assembly-level register access, the next lecture maps out the full Cortex-M memory space — flash, SRAM, peripherals, and system regions — and explains how the AHB/APB bus hierarchy determines which peripherals you can reach, and at what speed.

2 Comments

Leave a Reply

Your email address will not be published. Required fields are marked *