Bit-Banding: AtomicSingle-Bit Memory Access – Embedded C Training in Hyderabad

ARM Cortex-M Bit-Banding Explained | STM32 Embedded

Lecture 09 — ARM Cortex-M Programming

Bit-Banding: Atomic
Single-Bit Memory Access

Bit-banding is an ARM Cortex-M feature that gives every individual bit in SRAM and peripheral memory its own word-aligned address. Write a 1 or 0 to that address and exactly one bit changes — atomically, without a read-modify-write sequence, and without disabling interrupts.

Covers source pages 71–80  |  STM32F411  |  ARM Cortex-M TRM

1. The Problem Bit-Banding Solves

In embedded systems you constantly need to change a single bit inside a byte or word without disturbing any other bits. A classic example is toggling one GPIO output while leaving the others unchanged, or setting a flag in a status register.

The traditional approach is a read-modify-write (RMW) sequence:

/* Traditional: set bit 3 of a SRAM variable */
volatile uint8_t flags;

/* Step 1: read the current value */
uint8_t temp = flags;       /* LDR / LDRB */

/* Step 2: set bit 3 */
temp |= (1U << 3);         /* ORR        */

/* Step 3: write it back */
flags = temp;               /* STR / STRB */

This looks harmless, but it has a fatal flaw in interrupt-driven systems:

RMW Race Condition — The Interrupt Problem

Thread code │ Interrupt fires here │ Thread resumes

temp = flags; │ │
│ ISR: flags |= (1<<5); │ ← ISR sets bit 5
temp |= (1<<3); │ │ ← Thread still has OLD temp
flags = temp; │ │ ← ISR’s bit 5 OVERWRITTEN!

Result: bit 3 is set, but bit 5 that the ISR set is LOST.

The conventional fix is to disable interrupts around the RMW, accept the latency penalty, or use atomic intrinsics. Bit-banding offers an alternative: a single-instruction atomic write that changes exactly one bit with no window for an interrupt to corrupt the operation.

2. What is Bit-Banding?

Bit-banding is a Cortex-M hardware feature that maps every individual bit in a designated memory region to its own unique word address in a larger alias region. Writing 0x00000001 to the alias address sets that one bit. Writing 0x00000000 clears it. The hardware performs the actual read-modify-write at the memory level — atomically, invisible to the bus.

KEY PROPERTY

Hardware-atomic bit access

The bit-band alias write translates to a single indivisible hardware operation. No interrupt can see an intermediate state. This is true atomicity — not achieved by disabling interrupts in software.

OPTIONAL FEATURE

Not on every Cortex-M chip

Bit-banding is defined in the Cortex-M3 and Cortex-M4 architecture but is not present on Cortex-M0 or Cortex-M0+. Even on M3/M4, some chip vendors choose not to implement it. Always check the chip reference manual.

STM32F411: bit-banding is supported for both SRAM and peripheral regions.

Why “bit-banding”?
The name comes from the concept of each bit having its own “band” of address space in the alias region. The feature was first introduced in ARM processors of the 1980s and was carried forward into Cortex-M to support legacy code and provide atomic bit access without requiring the LDREX/STREX exclusive access instructions.

3. Bit-Band Regions in the Cortex-M Memory Map

There are exactly two bit-band regions in the Cortex-M memory map. Each has a corresponding alias region that is 32× larger:

Region type Bit-Band Region Size Alias Region Alias Size
SRAM 0x20000000 – 0x200FFFFF 1 MB 0x22000000 – 0x23FFFFFF 32 MB
Peripheral 0x40000000 – 0x400FFFFF 1 MB 0x42000000 – 0x43FFFFFF 32 MB

Bit-Band Regions in the Cortex-M Memory Map

SRAM Bit-Band Region
0x20000000
↕ 1 MB (1,048,576 bytes = 8,388,608 bits)
0x200FFFFF
⟷
SRAM Alias Region
0x22000000
↕ 32 MB (8,388,608 words)
0x23FFFFFF
Each word = 1 bit in the band
Peripheral Bit-Band Region
0x40000000
↕ 1 MB (8,388,608 bits)
0x400FFFFF
⟷
Peripheral Alias Region
0x42000000
↕ 32 MB (8,388,608 words)
0x43FFFFFF
Each word = 1 bit in the band

1 MB × 8 bits/byte = 8,388,608 bits. Each bit gets its own 4-byte (word) slot in the 32 MB alias. 32 MB / 4 bytes = 8,388,608 words. The math checks out.

Only the first 1 MB of each region is bit-band capable
On STM32F411, total SRAM is 128 KB (0x20000000–0x2001FFFF), which is entirely within the 1 MB bit-band region — so all SRAM is bit-band accessible. The peripheral bit-band region covers the first 1 MB of APB1 peripherals (0x40000000–0x400FFFFF), which includes TIM2–5, RTC, WWDG, IWDG, SPI2/3, USART2/3, I2C1–3, and the PWR/DAC block. GPIOs are on AHB1 at 0x40020000, which is outside the peripheral bit-band region.

4. How the Alias Mapping Works

Each byte in the bit-band region holds 8 bits. Each of those 8 bits gets its own 32-bit word in the alias region. Because a word is 4 bytes, 8 bits × 4 bytes = 32 bytes of alias space per byte of bit-band memory.

Here is the exact mapping for the first few bits of SRAM at 0x20000000:

Bit-Band → Alias Address Mapping (SRAM)

0x20000000 bit[0]
←→
0x22000000 word (bit[0])
0x20000000 bit[1]
←→
0x22000004 word (bit[0])
0x20000000 bit[2]
←→
0x22000008 word (bit[0])
⋮ (each subsequent bit adds 4 to alias address)
0x20000000 bit[7]
←→
0x2200001C word (bit[0])
0x20000000 bit[8]
(= bit[0] of next byte)
←→
0x22000020 word (bit[0])
⋮
0x20000000 bit[31]
←→
0x2200007C word (bit[0])
0x20000004 bit[0]
←→
0x22000080 word (bit[0])
0x20000004 bit[1]
←→
0x22000084 word (bit[0])

Pattern: each bit in the bit-band region gets a 4-byte (word) slot in the alias. Reading the alias word returns 0 or 1 (not the full byte value). Writing 1 sets the bit; writing 0 clears it. All other bits of the written word are ignored.

Why words in the alias, not bytes?
The AHB bus on Cortex-M is 32 bits wide. A word-sized alias means every bit access is a naturally aligned 32-bit AHB transfer — the simplest and most efficient transaction type. Using byte-sized alias slots would complicate the bus logic and potentially require byte-enable signals not present on all slave interfaces.

5. Bit-to-Word Expansion Diagram

Think of it this way: take a 32-bit word in SRAM and “expand” each of its 32 bits into its own 4-byte word in the alias. The first word at 0x20000000 expands into 32 alias words occupying 0x22000000–0x2200007C (128 bytes). The next word at 0x20000004 maps to 0x22000080–0x220000FC, and so on.

Word Expansion: 32-bit SRAM Word → 32 Alias Words

BIT-BAND (0x20000000)
31
30
…
6
5
4
3
2
1
0
One 32-bit word in SRAM
⟹
ALIAS REGION (0x22000000 – 0x2200007C)
0x2200007C
bit[31] = 1 (set)
0x22000078
bit[30] = 0
⋮
0x22000018
bit[6] = 1 (set)
…
bits 5–1 = 0
0x22000000
bit[0] = 1 (set)

6. The Alias Address Formula

Given a byte address in the bit-band region and a bit position (0–7 for byte access, 0–31 for word access), you can calculate the exact alias word address using the following formula:

alias_addr = alias_base + (32 × (byte_addr − band_base)) + (bit_num × 4)
Where:
alias_base = 0x22000000 (SRAM) or 0x42000000 (Peripheral)
band_base = 0x20000000 (SRAM) or 0x40000000 (Peripheral)
byte_addr = actual byte address in the bit-band region
bit_num = bit position within the byte/word (0 = LSB)

Why 32 and 4?

Each byte in the bit-band region has 8 bits. Each bit gets a 4-byte (word) alias slot. So one byte uses 8 × 4 = 32 bytes of alias space. The factor of 32 accounts for moving to the next byte; the factor of 4 accounts for moving to the next bit within the same byte.

Formula Derivation — Where Do 32 and 4 Come From?

/* 1 byte = 8 bits, each bit → 1 alias word (4 bytes) */
1 byte in band → 8 alias words = 8 × 4 bytes = 32 bytes in alias

/* Moving from byte N to byte N+1 in the band = +32 in alias */
byte_offset = byte_addr – band_base
alias_byte_offset = 32 × byte_offset

/* Within one byte, each bit is +4 bytes apart in alias */
bit_byte_offset = bit_num × 4

/* Combine: */
alias_addr = alias_base + alias_byte_offset + bit_byte_offset
= alias_base + (32 × (byte_addr – band_base)) + (bit_num × 4)
Alternative form using bit shifts
The formula can be written more efficiently in C using shifts instead of multiply: 32 × x = x << 5 and 4 × y = y << 2, which the compiler will do automatically when you write the formula with multiplies and constants.

7. Worked Examples — Calculating Alias Addresses

Example A: bit 0 of 0x20000000 (SRAM)

alias_base = 0x22000000
band_base  = 0x20000000
byte_addr  = 0x20000000
bit_num    = 0

alias = 0x22000000 + (32 × (0x20000000 - 0x20000000)) + (0 × 4)
      = 0x22000000 + (32 × 0) + 0
      = 0x22000000   ✓

Example B: bit 1 of 0x20000000

alias = 0x22000000 + (32 × 0) + (1 × 4)
      = 0x22000000 + 0 + 4
      = 0x22000004   ✓

Example C: bit 31 of 0x20000000 (highest bit of first word)

alias = 0x22000000 + (32 × 0) + (31 × 4)
      = 0x22000000 + 0 + 124
      = 0x22000000 + 0x7C
      = 0x2200007C   ✓

Example D: bit 0 of 0x20000004 (next word in SRAM)

alias = 0x22000000 + (32 × (0x20000004 - 0x20000000)) + (0 × 4)
      = 0x22000000 + (32 × 4) + 0
      = 0x22000000 + 128
      = 0x22000000 + 0x80
      = 0x22000080   ✓

Example E: bit 6 of 0x20000000 (matching the diagram)

alias = 0x22000000 + (32 × 0) + (6 × 4)
      = 0x22000000 + 0 + 24
      = 0x22000000 + 0x18
      = 0x22000018   ✓

8. Exercise: Clear Bit 7 of Address 0x20000200

This is the exercise from the slides. We want to:

  1. Store the value 0xFF into SRAM address 0x20000200
  2. Clear bit 7 of that byte using the bit-band alias
  3. Compare with the traditional RMW approach

Step 1: Calculate the alias address for bit 7 of 0x20000200

alias_base = 0x22000000
band_base = 0x20000000
byte_addr = 0x20000200
bit_num = 7

alias = 0x22000000 + (32 × (0x20000200 − 0x20000000)) + (7 × 4)
= 0x22000000 + (32 × 0x200) + 28
= 0x22000000 + 0x4000 + 0x1C
= 0x2200401C

Step 2: Write the C code — traditional vs bit-band

#include <stdint.h>

/* ---- Traditional read-modify-write approach ---- */
void clear_bit7_traditional(void)
{
    volatile uint8_t *p = (volatile uint8_t *)0x20000200U;

    /* Store 0xFF */
    *p = 0xFF;

    /* Clear bit 7 — 3 instructions: LDR, BIC, STR */
    *p &= ~(1U << 7);

    /* Result: 0x7F
       Problem: an ISR between LDR and STR could corrupt other bits */
}

/* ---- Bit-band approach ---- */
void clear_bit7_bitband(void)
{
    volatile uint8_t  *p      = (volatile uint8_t  *)0x20000200U;
    volatile uint32_t *bit7   = (volatile uint32_t *)0x2200401CU;

    /* Store 0xFF */
    *p = 0xFF;

    /* Clear bit 7 — 1 instruction: STR (atomic) */
    *bit7 = 0U;

    /* Result: 0x7F
       Atomic: no ISR can interrupt between read and write —
       the hardware does the RMW internally in one bus cycle */
}

/* ---- Verify both produce the same result ---- */
void demo(void)
{
    volatile uint8_t *p = (volatile uint8_t *)0x20000200U;

    /* Traditional */
    *p = 0xFF;
    *p &= ~(1U << 7);
    /* *p == 0x7F */

    /* Reset */
    *p = 0xFF;

    /* Bit-band */
    volatile uint32_t *alias = (volatile uint32_t *)0x2200401CU;
    *alias = 0U;
    /* *p == 0x7F — same result, but atomically */
}

Disassembly comparison

/* Traditional (*p &= ~(1<<7)) disassembly at -O2: */
  LDRB  R0, [R1]      ; read byte from 0x20000200
  BIC   R0, R0, #128  ; clear bit 7
  STRB  R0, [R1]      ; write back
  ; ← interrupt window between LDRB and STRB!

/* Bit-band (*alias = 0) disassembly at -O2: */
  MOV   R0, #0
  STR   R0, [R1]      ; single write to alias address 0x2200401C
  ; ← no interrupt window — one indivisible bus transaction

9. Implementing Bit-Banding in C

Calculating alias addresses by hand every time is error-prone. The standard approach is a macro or inline function that computes the alias address from the byte address and bit number at compile time.

/* ----------------------------------------------------------------
   Generic bit-band macros for Cortex-M3/M4
   ---------------------------------------------------------------- */

/* SRAM bit-band */
#define SRAM_BB_BASE    0x22000000UL
#define SRAM_BB_REGION  0x20000000UL

/* Peripheral bit-band */
#define PERIPH_BB_BASE    0x42000000UL
#define PERIPH_BB_REGION  0x40000000UL

/* Compute alias address for any bit-band byte and bit number.
   Returns a volatile uint32_t pointer ready to read or write. */
#define BB_SRAM_ALIAS(byte_addr, bit)  \
    ((volatile uint32_t *)(SRAM_BB_BASE   + (32U * ((uint32_t)(byte_addr) - SRAM_BB_REGION))   + ((bit) * 4U)))

#define BB_PERIPH_ALIAS(byte_addr, bit) \
    ((volatile uint32_t *)(PERIPH_BB_BASE + (32U * ((uint32_t)(byte_addr) - PERIPH_BB_REGION)) + ((bit) * 4U)))

/* ----------------------------------------------------------------
   Usage examples
   ---------------------------------------------------------------- */

/* Set bit 3 of a SRAM variable atomically */
volatile uint32_t flags = 0;

static inline void flags_set_bit3(void)   { *BB_SRAM_ALIAS(&flags, 3) = 1U; }
static inline void flags_clear_bit3(void) { *BB_SRAM_ALIAS(&flags, 3) = 0U; }
static inline uint32_t flags_get_bit3(void) { return *BB_SRAM_ALIAS(&flags, 3); }

/* Toggle bit using alias — XOR is not atomic, but set/clear is */
static inline void flags_set(uint32_t bit)   { *BB_SRAM_ALIAS(&flags, bit) = 1U; }
static inline void flags_clear(uint32_t bit) { *BB_SRAM_ALIAS(&flags, bit) = 0U; }
static inline uint32_t flags_get(uint32_t bit) { return *BB_SRAM_ALIAS(&flags, bit); }

/* ----------------------------------------------------------------
   Practical example: shared flag between main loop and ISR
   ---------------------------------------------------------------- */

volatile uint8_t event_flags = 0;  /* bit 0: UART received, bit 1: timer expired */

/* ISR — set bit 0 atomically, no need to disable interrupts */
void USART2_IRQHandler(void)
{
    *BB_SRAM_ALIAS(&event_flags, 0) = 1U;  /* set "UART received" flag */
}

/* Main loop — clear bit 0 atomically after handling */
void process_events(void)
{
    if (*BB_SRAM_ALIAS(&event_flags, 0)) {
        handle_uart_data();
        *BB_SRAM_ALIAS(&event_flags, 0) = 0U;  /* clear atomically */
    }
}
Reading the alias — only bit 0 is meaningful
When you read an alias address, the hardware returns 0x00000001 if the corresponding bit is set, or 0x00000000 if it is clear. The upper 31 bits are always zero. Do not try to extract any other bits from the alias read value.

10. Bit-Banding for Peripheral Registers

The peripheral bit-band region covers 0x40000000–0x400FFFFF, which on STM32F411 contains the APB1 bus peripherals. This lets you set or clear individual bits in peripheral control registers atomically — useful for enabling/disabling peripheral functions without affecting other control bits.

GPIO registers are NOT in the peripheral bit-band region
STM32F411 GPIO registers start at 0x40020000 (AHB1 bus), which is outside the 1 MB peripheral bit-band region (0x40000000–0x400FFFFF). You cannot use bit-band aliasing for GPIO ODR/BSRR — use the BSRR register instead, which provides atomic set/clear as part of its hardware design.
/* ----------------------------------------------------------------
   Peripheral bit-band example: USART2 on APB1 (within bit-band region)
   USART2_CR1 is at 0x4000440C
   bit 13 (UE) = USART Enable
   bit 3  (TE) = Transmitter Enable
   bit 2  (RE) = Receiver Enable
   ---------------------------------------------------------------- */

#define USART2_CR1_ADDR  0x4000440CUL

/* Enable USART2 (UE bit 13) atomically */
#define USART2_UE_BIT    (*BB_PERIPH_ALIAS(USART2_CR1_ADDR, 13))

/* Enable Transmitter (TE bit 3) atomically */
#define USART2_TE_BIT    (*BB_PERIPH_ALIAS(USART2_CR1_ADDR, 3))

/* Enable Receiver (RE bit 2) atomically */
#define USART2_RE_BIT    (*BB_PERIPH_ALIAS(USART2_CR1_ADDR, 2))

void usart2_enable(void)
{
    USART2_TE_BIT = 1U;   /* Enable TX — single atomic write */
    USART2_RE_BIT = 1U;   /* Enable RX — single atomic write */
    USART2_UE_BIT = 1U;   /* Enable USART — single atomic write */
}

void usart2_disable(void)
{
    USART2_UE_BIT = 0U;   /* Disable USART atomically */
}

/* ----------------------------------------------------------------
   Alias address verification for USART2_CR1 bit 13:
   alias = 0x42000000 + (32 * (0x4000440C - 0x40000000)) + (13 * 4)
         = 0x42000000 + (32 * 0x440C) + 52
         = 0x42000000 + 0x88180 + 0x34
         = 0x420881B4
   ---------------------------------------------------------------- */

RCC register bit-banding — enable peripheral clocks atomically

/* RCC_AHB1ENR is at 0x40023830 — is this in the peripheral bit-band region?
   0x40023830 < 0x400FFFFF → YES, it is within the first 1 MB.              */

#define RCC_AHB1ENR_ADDR  0x40023830UL

/* Bit 0 of RCC_AHB1ENR = GPIOAEN */
#define RCC_GPIOAEN  (*BB_PERIPH_ALIAS(RCC_AHB1ENR_ADDR, 0))

/* Enable GPIOA clock atomically — no risk of disturbing other clock bits */
void enable_gpioa_clock(void)
{
    RCC_GPIOAEN = 1U;
}

11. When to Use (and Not Use) Bit-Banding

USE IT WHEN
  • You need to set or clear a single bit in SRAM that an ISR also touches
  • You want interrupt-safe flag manipulation without disabling interrupts
  • You are clearing/setting a bit in an APB1 peripheral register and need atomicity
  • You want the simplest possible bit manipulation — one write instruction
  • Porting legacy ARM7 code that relied on bit-banding
AVOID IT WHEN
  • You need to manipulate multiple bits atomically — each alias write is independent
  • The target address is outside the bit-band region (GPIOs on AHB1, for example)
  • You are on a Cortex-M0/M0+ — bit-banding is not available
  • CMSIS atomic intrinsics (__LDREXW/__STREXW) give you more portable multi-bit atomicity
  • Readability matters more than the minor performance gain

Modern alternative: C11 atomic operations

/* C11 _Atomic gives portable atomic single-bit operations on any architecture */
#include <stdatomic.h>

_Atomic uint8_t event_flags = 0;

/* ISR */
void USART2_IRQHandler(void)
{
    atomic_fetch_or(&event_flags, (1U << 0));  /* set bit 0 atomically */
}

/* Main */
void process(void)
{
    if (atomic_load(&event_flags) & (1U << 0)) {
        handle_uart();
        atomic_fetch_and(&event_flags, ~(1U << 0));  /* clear bit 0 atomically */
    }
}
Bit-banding vs C11 atomics vs LDREX/STREX
Bit-banding: hardware-enforced, zero overhead, single instruction — but limited to the bit-band regions and not portable to Cortex-M0.
LDREX/STREX: available on M3/M4, works anywhere in SRAM, can handle multi-bit operations but requires a retry loop.
C11 atomics: fully portable, compiler picks the best implementation (LDREX/STREX on Cortex-M3/M4), recommended for new code.

12. FAQ

Q: Is writing to the alias address truly atomic — even with DMA running?
Yes. The bit-band write is atomic at the bus level. The Cortex-M memory system performs the read-modify-write as a single locked transaction. A DMA access targeting the same byte would be blocked by the bus matrix until the bit-band operation completes. No partial state is visible to any bus master.
Q: What happens if I write a value other than 0 or 1 to an alias address?
Only bit 0 of the written value matters. Writing 0xFFFFFFFF sets the corresponding bit (because bit 0 = 1). Writing 0xFFFFFFFE clears it (bit 0 = 0). Writing 0x00000002 also clears it (bit 0 = 0). The hardware reads only the LSB of the alias write data.
Q: Can I read through the alias to check if a bit is set?
Yes. Reading an alias address returns 0x00000001 if the bit is set, 0x00000000 if it is clear. The read is not atomic in the same sense as the write — it is just a normal memory read of the current bit value. Use it freely for polling.
Q: Does bit-banding work for bits 8–31 of a 32-bit word?
Yes. The formula works for any bit position 0–31 of any byte address within the bit-band region. For a 32-bit word at address A, bit N has alias address: alias_base + (32 × (A − band_base)) + (N × 4). The formula treats the memory as a flat array of bytes, so bit 8 of a word is actually bit 0 of the next byte (A+1), giving: alias_base + (32 × (A+1 − band_base)) + 0.
Q: STM32F411 GPIOs are at 0x40020000 — can I use bit-banding for them?
No. The peripheral bit-band region only covers 0x40000000–0x400FFFFF. GPIO registers start at 0x40020000 which is 0x20000 bytes into the peripheral region — beyond the 1 MB limit (0x100000 = 1,048,576 bytes). For atomic GPIO control, use the BSRR register (write to high 16 bits to clear, low 16 bits to set) which is designed for atomic bit manipulation.
Q: Should I use bit-banding in new projects, or is there a better way?
For new Cortex-M3/M4 projects, bit-banding is a perfectly valid technique for interrupt-safe flag manipulation. However, the C11 _Atomic type and atomic_fetch_or/atomic_fetch_and operations are more portable and readable. CMSIS also provides __LDREXW/__STREXW for exclusive access. Choose based on your portability requirements and team familiarity. Bit-banding is most useful when you need guaranteed single-instruction behaviour and are certain your target supports it.

Next: Stack Memory and Stack Operations

With memory regions fully mapped, the next lecture focuses on the stack — how it grows, what gets pushed during function calls and exceptions, and how to size it correctly for your application.

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *