AMBA Bus Protocols:
AHB Lite, APB & Bus Interfaces
Every data transfer inside a Cortex-M chip — instruction fetch, peripheral write, DMA burst — travels over one of ARM’s AMBA buses. This lecture covers the XN memory region properties, the AMBA specification, AHB Lite vs APB differences, and the four dedicated bus interfaces that connect the CPU core to the rest of the chip.
- Execute Never (XN) — Region Memory Attributes
- External RAM Region (0x60000000–0x9FFFFFFF)
- External Device Region (0xA0000000–0xDFFFFFFF)
- Private Peripheral Bus Region (0xE0000000–0xE00FFFFF)
- What is AMBA?
- AHB Lite — High-performance Bus
- APB — Peripheral Bus
- The AHB-to-APB Bridge
- Four Cortex-M Bus Interfaces
- Bus Matrix on STM32F411
- Performance Implications
- FAQ
1. Execute Never (XN) — Region Memory Attributes
Not all memory regions are equal. The ARM architecture tags certain regions with an XN (eXecute Never) attribute, which tells the processor’s Memory Protection Unit (MPU) and architecture rules that code must never be fetched for execution from those addresses.
If the processor’s program counter ever points into an XN region — either because a bug overwrote the PC or because an attacker tried to execute data as code — the hardware raises a MemManage fault (or HardFault if MemManage is not enabled) before a single instruction from that region executes.
Regions where code CAN run
- Code (0x00000000–0x1FFFFFFF) — flash, ROM
- SRAM (0x20000000–0x3FFFFFFF) — RAM functions allowed
- External RAM (0x60000000–0x9FFFFFFF) — XIP from SDRAM/NOR
Regions where code CANNOT run
- Peripheral (0x40000000–0x5FFFFFFF) — hardware registers only
- External Device (0xA0000000–0xDFFFFFFF) — NOR/LCD in data mode
- PPB (0xE0000000–0xE00FFFFF) — NVIC/SCB/ITM registers
- Vendor-specific (0xE0100000–0xFFFFFFFF) — vendor extensions
Memory Map with XN Attributes
XN regions cause a fault if the CPU tries to fetch instructions from them — regardless of what is stored there.
2. External RAM Region (0x60000000 – 0x9FFFFFFF)
This 1 GB region is reserved for external memory chips connected to the processor through a memory controller. Unlike the peripheral region, code can be executed from here — making it suitable for Execute-in-Place (XIP) applications where the program is too large to fit in on-chip flash.
What connects here
- External SDRAM — e.g. 8 MB IS42S16400 on STM32F429 Discovery
- External SRAM / PSRAM — pseudo-static RAM, byte-accessible
- External NOR flash — execute-in-place program storage
- NAND flash — bulk storage (not XIP; must copy to RAM first)
Requires FMC / FSMC controller
Accessing the External RAM region requires the Flexible Memory Controller (FMC) or its older variant FSMC to be configured. The STM32F411 does not have FMC. Chips like STM32F407, STM32F429, and STM32H743 do. Without FMC, accesses into this region generate a bus fault.
Configuring external SDRAM on STM32F4 with FMC
/* Conceptual: FMC initialisation for IS42S16400J SDRAM (STM32F4 with FMC) */
/* NOT available on STM32F411 — shown for reference on larger STM32F4 devices */
#define FMC_BASE 0xA0000000UL /* FMC register base */
#define SDRAM_BANK1 0xC0000000UL /* External RAM bank 1 mapped here */
/* After FMC SDRAM init, use external memory like normal SRAM: */
uint32_t *ext_ram = (uint32_t *)SDRAM_BANK1;
ext_ram[0] = 0xDEADBEEF; /* write to external SDRAM */
uint32_t val = ext_ram[0]; /* read back */
/* Place large arrays in external RAM with linker section attribute: */
__attribute__((section(".sdram")))
uint8_t video_frame_buffer[480 * 272 * 2]; /* 261 KB — too big for on-chip SRAM */
3. External Device Region (0xA0000000 – 0xDFFFFFFF)
This 1 GB region is designated for external devices and shared memory — equipment that behaves like a peripheral rather than addressable memory. It is permanently marked XN (eXecute Never) and is also non-cacheable and non-bufferable by default, meaning every access goes directly to the external device with no memory-system optimisation.
External Device region applications
- NOR flash in data mode — reading flash ID / status registers
- Parallel LCD interfaces — Intel 8080 / Motorola 6800 mode
- Shared memory — dual-port RAM shared with another processor
- FPGA-mapped registers — FPGA connected via FMC
Never executable
Even if you wrote valid ARM Thumb-2 machine code into a device in this region, the processor would fault before executing it. The XN attribute is enforced by the hardware bus decoder before the instruction reaches the decode stage.
/* Parallel LCD using FMC in 8080 mode — External Device region */
/* Command register at bank 1, A16=0 (RS line tied to FMC address bit 16) */
#define LCD_CMD_ADDR 0x60000000UL /* External Device region, bank 1 */
#define LCD_DATA_ADDR 0x60020000UL /* A16=1 → data register */
#define LCD_CMD (*(volatile uint16_t *)LCD_CMD_ADDR)
#define LCD_DATA (*(volatile uint16_t *)LCD_DATA_ADDR)
void lcd_write_reg(uint16_t reg, uint16_t value)
{
LCD_CMD = reg; /* select register */
LCD_DATA = value; /* write value */
}
void lcd_set_pixel(uint16_t color)
{
LCD_DATA = color; /* stream pixel data to GRAM */
}
4. Private Peripheral Bus Region (0xE0000000 – 0xE00FFFFF)
The Private Peripheral Bus (PPB) region is 1 MB and contains the ARM-defined system control registers that are identical on every Cortex-M chip. It is XN and always requires privileged access.
System Control Space (SCS)
NVIC, SCB, SysTick, MPU, FPU control — the core OS-level control registers. 4 KB block starting at 0xE000E000.
ITM, DWT, FPB
Instrumentation Trace Macrocell (ITM) for printf over SWO, Data Watchpoint and Trace (DWT) for cycle counting, and Flash Patch and Breakpoint (FPB) for hardware breakpoints.
Access rules
Unprivileged code that attempts to access any PPB address will cause a MemManage fault (or HardFault if MPU is off). Only Handler mode and Privileged Thread mode code can touch these registers.
/* Key PPB registers — same addresses on ALL Cortex-M devices */
/* SysTick (at 0xE000E010) */
#define SYST_CSR (*(volatile uint32_t *)0xE000E010U) /* Control & Status */
#define SYST_RVR (*(volatile uint32_t *)0xE000E014U) /* Reload Value */
#define SYST_CVR (*(volatile uint32_t *)0xE000E018U) /* Current Value */
/* NVIC (at 0xE000E100) */
#define NVIC_ISER0 (*(volatile uint32_t *)0xE000E100U) /* Interrupt Set Enable */
#define NVIC_ICER0 (*(volatile uint32_t *)0xE000E180U) /* Interrupt Clear Enable */
#define NVIC_ISPR0 (*(volatile uint32_t *)0xE000E200U) /* Interrupt Set Pending */
#define NVIC_IPR0 (*(volatile uint32_t *)0xE000E400U) /* Interrupt Priority */
/* SCB (at 0xE000ED00) */
#define SCB_CPUID (*(volatile uint32_t *)0xE000ED00U) /* CPU ID register */
#define SCB_AIRCR (*(volatile uint32_t *)0xE000ED0CU) /* App Int & Reset Control*/
#define SCB_SHCSR (*(volatile uint32_t *)0xE000ED24U) /* System Handler Control */
/* DWT Cycle Counter (at 0xE0001000) */
#define DWT_CTRL (*(volatile uint32_t *)0xE0001000U)
#define DWT_CYCCNT (*(volatile uint32_t *)0xE0001004U)
/* Read CPU ID: e.g. 0x410FC241 for Cortex-M4 r0p1 */
void print_cpu_id(void)
{
uint32_t id = SCB_CPUID;
/* bits [19:16] = Constant = 0xF
bits [15:4] = Part number = 0xC24 (M4) or 0xC23 (M3)
bits [3:0] = Revision */
}
5. What is AMBA?
AMBA stands for Advanced Microcontroller Bus Architecture. It is an open specification developed by ARM that defines standard on-chip bus protocols for connecting processor cores, memory, and peripherals inside a System-on-Chip (SoC).
Because AMBA is a published open standard, any chip designer (ST, NXP, Broadcom, Qualcomm, Apple) can use it to interconnect ARM IP blocks — the CPU, DMA controller, memory controller, and peripherals — without needing custom bus logic. This is a major reason why ARM-based chips dominate the embedded and mobile markets: the ecosystem of compatible IP is enormous.
AMBA Version History
| Version | Year | Key Protocols | Used in |
|---|---|---|---|
| AMBA 1 | 1996 | ASB (Advanced System Bus), APB | ARM7TDMI era chips |
| AMBA 2 | 1999 | AHB (full), APB rev2 | ARM9, ARM11 SoCs |
| AMBA 3 | 2003 | AHB-Lite, APB3, AXI | Cortex-M, Cortex-A9 |
| AMBA 4 | 2010 | AXI4, AXI4-Lite, AXI4-Stream, ACE | Cortex-A15, A57, modern SoCs |
| AMBA 5 | 2013 | CHI (Coherent Hub Interface) | Cortex-A72, server-class chips |
Cortex-M processors use AMBA 3: AHB-Lite for high-speed interfaces and APB for slower peripherals.
6. AHB Lite — High-performance Bus
AHB Lite (AMBA High-performance Bus — Lite) is the primary bus used for the Cortex-M core’s main interfaces. It carries instruction fetches, data reads/writes to SRAM, and accesses to high-speed peripherals like GPIO, DMA, and RCC.
AHB Lite properties
- 32-bit address bus + 32-bit data bus
- Single master per bus segment
- Pipelined transfers (address phase + data phase)
- Burst transfers: 4-beat, 8-beat, 16-beat
- Operates at core clock frequency (up to 100 MHz on STM32F411)
- Supports 8-, 16-, and 32-bit data widths
- Low latency: typically 1–2 clock cycles per transfer
Key AHB Lite signals
HADDR[31:0]— 32-bit addressHWRITE— read (0) or write (1)HSIZE[2:0]— transfer size (byte/half/word)HBURST[2:0]— single or burst typeHWDATA[31:0]— write dataHRDATA[31:0]— read dataHREADY— slave ready signal (wait states)HRESP— OKAY or ERROR response
AHB Lite pipelined transfer timing
Two-transfer AHB Lite pipeline (no wait states)
On STM32F411, the AHB1 bus runs at the full system clock (100 MHz), giving a peak throughput of 400 MB/s for 32-bit transfers (100 MHz × 4 bytes). This is why GPIO toggling can be very fast on Cortex-M — a single 32-bit write to the GPIOA_BSRR register over AHB1 completes in one clock cycle.
/* GPIO BSRR: fastest way to set/clear GPIO bits — single AHB write */
#define GPIOA_BSRR (*(volatile uint32_t *)0x40020018U)
/* Set PA5 HIGH — one 32-bit AHB write, one clock cycle */
GPIOA_BSRR = (1U << 5);
/* Set PA5 LOW — write to upper 16 bits (reset bits) */
GPIOA_BSRR = (1U << (5 + 16));
/* Toggle PA5 at maximum speed using ODR XOR */
#define GPIOA_ODR (*(volatile uint32_t *)0x40020014U)
while (1) {
GPIOA_ODR ^= (1U << 5); /* Each iteration = 1 AHB write + loop overhead */
}
7. APB — Advanced Peripheral Bus
APB (Advanced Peripheral Bus) is a simpler, lower-speed bus designed for peripherals that do not need the full bandwidth of AHB. It has a smaller signal count, lower gate count, and lower power consumption — ideal for UART, I2C, basic timers, and watchdogs.
How APB differs from AHB
- No pipelining — address and data in same cycle (simpler slaves)
- No burst mode — one transfer at a time
- Runs at a fraction of AHB clock (APB1 = AHB/2 on STM32F411)
- Narrower signal set — no HBURST, HPROT, HMASTER
- Minimum 2 clock cycles per transfer (setup + access)
- Cannot be a bus master — only AHB is a master
Simplified APB signal set
PCLK— peripheral clock (slower than HCLK)PADDR[31:0]— addressPWRITE— read or writePWDATA[31:0]— write dataPRDATA[31:0]— read dataPSEL— select this peripheralPENABLE— enable strobe (2nd cycle)PREADY— slave done (optional wait states)
APB Transfer Timing — Minimum 2 Cycles
PPRE1 (APB1 prescaler) and PPRE2 (APB2 prescaler).
At 100 MHz HCLK with PPRE1=4 (divide by 2): APB1 PCLK = 50 MHz.
With PPRE2=2 (no divide): APB2 PCLK = 100 MHz.
Timer peripherals have an automatic ×2 multiplier if their APB prescaler is not 1,
so TIM2–5 actually tick at 100 MHz even though APB1 runs at 50 MHz.
8. The AHB-to-APB Bridge
Slow APB peripherals cannot connect directly to the fast AHB bus — their timing requirements are incompatible. The AHB-to-APB bridge is a hardware module that translates between the two bus protocols, handling the clock domain crossing and protocol conversion transparently.
AHB to APB Bridge — Protocol Translation
AHB Master
Bridge
Protocol converter
SPI, TIM2–5
APB Slaves
The bridge holds the AHB transaction in the setup phase while it completes the slower APB transaction. The CPU sees this as a wait state — HREADY goes low until the APB transfer is done.
From the programmer’s perspective, the bridge is completely invisible. You write to a UART register address and the hardware handles the protocol conversion. The only observable effect is that APB peripheral accesses take more clock cycles than AHB peripheral accesses.
/* USART2 is on APB1 at 0x40004400 — access goes through AHB-APB bridge */
#define USART2_SR (*(volatile uint32_t *)0x40004400U)
#define USART2_DR (*(volatile uint32_t *)0x40004404U)
#define USART2_BRR (*(volatile uint32_t *)0x40004408U)
#define USART2_CR1 (*(volatile uint32_t *)0x4000440CU)
/* Writing to USART2_DR is an AHB transaction that the bridge translates
to an APB write at 50 MHz. The CPU may insert 1 wait state. */
void uart_send_byte(uint8_t byte)
{
while (!(USART2_SR & (1U << 7))); /* wait TXE (TX Empty) */
USART2_DR = byte;
}
/* Compare: GPIOA on AHB1 — no bridge, no wait states */
#define GPIOA_ODR (*(volatile uint32_t *)0x40020014U)
void gpio_toggle(void) { GPIOA_ODR ^= (1U << 5); } /* 1 cycle */
9. Four Cortex-M Bus Interfaces
The Cortex-M processor core exposes four distinct AHB Lite bus interfaces to the outside world. Each has a specific purpose. This is why the core can do more than one thing at a time — instruction fetch and data access can happen simultaneously because they use different buses.
Cortex-M Core — Four Bus Interfaces
Core
| Bus Interface | Protocol | Direction | Targets | Notes |
|---|---|---|---|---|
| I-CODE | AHB-Lite | Read only | Code region (0x00–0x1F) | Instruction fetches + vector table reads. Can burst 8 half-words. |
| D-CODE | AHB-Lite | Read only | Code region (0x00–0x1F) | Literal pool and rodata reads. Simultaneous with I-CODE. |
| System | AHB-Lite | Read/Write | SRAM, Peripheral, Ext RAM, Ext Device, Vendor | Main data bus. Also handles debug accesses from JTAG/SWD via AHB-AP. |
| PPB | APB subset | Read/Write | PPB region (0xE0000000–0xE00FFFFF) | System control only. Privileged access required. XN region. |
10. Bus Matrix on STM32F411
On the STM32F411, a multi-layer AHB bus matrix connects multiple bus masters (CPU, DMA1, DMA2) to multiple slaves (flash, SRAM, peripheral buses). The matrix allows simultaneous accesses as long as two masters target different slaves.
STM32F411 Multi-layer AHB Bus Matrix (simplified)
✓ = can access this slave. Two masters on different columns can transfer simultaneously — CPU fetching from flash while DMA copies from SRAM to UART, for example.
Practical example — DMA + CPU simultaneous access
/* DMA2 transfers data from SRAM to SPI1 (on APB2)
while CPU continues executing from flash.
No conflict: CPU uses I-CODE (flash), DMA2 uses System bus (SRAM→APB2). */
#define DMA2_S3CR (*(volatile uint32_t *)0x40026458U) /* DMA2 stream3 CR */
#define DMA2_S3NDTR (*(volatile uint32_t *)0x4002645CU) /* transfer count */
#define DMA2_S3PAR (*(volatile uint32_t *)0x40026460U) /* peripheral addr */
#define DMA2_S3M0AR (*(volatile uint32_t *)0x40026464U) /* memory addr */
static uint8_t spi_tx_buf[64]; /* in SRAM — filled by CPU */
void start_spi_dma(void)
{
DMA2_S3PAR = 0x4001300CU; /* SPI1->DR address (APB2) */
DMA2_S3M0AR = (uint32_t)spi_tx_buf; /* source in SRAM */
DMA2_S3NDTR = 64;
DMA2_S3CR = (3U << 25) /* channel 3 */
| (1U << 6) /* memory-to-peripheral */
| (1U << 10) /* memory increment */
| (1U << 0); /* enable stream */
/* CPU is now free — DMA handles the transfer via the bus matrix */
}
11. Performance Implications for Embedded Code
Understanding the bus architecture helps you write faster code and avoid hidden bottlenecks.
Zero-wait-state operations
- Instruction fetch from flash (with ART accelerator enabled)
- SRAM reads and writes over System bus
- GPIO write via AHB1 (GPIOA_ODR, GPIOA_BSRR)
- DMA transfers that do not contend with CPU
Higher latency operations
- APB peripheral access (UART, I2C, SPI registers) — min 2 PCLK cycles
- Flash reads without ART accelerator at high clock speeds
- External RAM via FMC — depends on timing config (many wait states)
- Bus contention between CPU and DMA on same slave
Clock enable: gating matters for APB peripherals
/* Peripheral clocks are DISABLED by default to save power.
You must enable the clock before the first register access.
Accessing a peripheral with its clock off returns 0x00000000 on reads
and the write is silently ignored — a common beginner bug. */
#define RCC_BASE 0x40023800UL
#define RCC_AHB1ENR (*(volatile uint32_t *)(RCC_BASE + 0x30U))
#define RCC_APB1ENR (*(volatile uint32_t *)(RCC_BASE + 0x40U))
#define RCC_APB2ENR (*(volatile uint32_t *)(RCC_BASE + 0x44U))
void enable_clocks(void)
{
/* AHB1: enable GPIOA (bit 0) and DMA2 (bit 22) */
RCC_AHB1ENR |= (1U << 0) | (1U << 22);
/* APB1: enable USART2 (bit 17) and I2C1 (bit 21) */
RCC_APB1ENR |= (1U << 17) | (1U << 21);
/* APB2: enable SPI1 (bit 12) and TIM1 (bit 0) */
RCC_APB2ENR |= (1U << 12) | (1U << 0);
/* After enabling, a small delay is required for clock propagation.
Reading the ENR register once is sufficient on STM32F4: */
(void)RCC_AHB1ENR;
}
12. FAQ
Next: STM32 Peripheral Register Programming
Now that the bus architecture is clear, the next lecture applies it directly — working through the full bit-field structure of GPIO, RCC, and NVIC registers, and showing how to configure peripherals from scratch without HAL.

2 Comments