AMBA Bus Protocols:AHB Lite, APB & Bus Interfaces-Embedded Systems Course for Freshers in Hyderabad

AMBA Bus Protocols AHB APB Cortex-M | STM32 Embedded

Lecture 08 — ARM Cortex-M Programming

AMBA Bus Protocols:
AHB Lite, APB & Bus Interfaces

Every data transfer inside a Cortex-M chip — instruction fetch, peripheral write, DMA burst — travels over one of ARM’s AMBA buses. This lecture covers the XN memory region properties, the AMBA specification, AHB Lite vs APB differences, and the four dedicated bus interfaces that connect the CPU core to the rest of the chip.

Covers source pages 61–70  |  STM32F411  |  AMBA 3 Spec

1. Execute Never (XN) — Region Memory Attributes

Not all memory regions are equal. The ARM architecture tags certain regions with an XN (eXecute Never) attribute, which tells the processor’s Memory Protection Unit (MPU) and architecture rules that code must never be fetched for execution from those addresses.

If the processor’s program counter ever points into an XN region — either because a bug overwrote the PC or because an attacker tried to execute data as code — the hardware raises a MemManage fault (or HardFault if MemManage is not enabled) before a single instruction from that region executes.

EXECUTABLE

Regions where code CAN run

  • Code (0x00000000–0x1FFFFFFF) — flash, ROM
  • SRAM (0x20000000–0x3FFFFFFF) — RAM functions allowed
  • External RAM (0x60000000–0x9FFFFFFF) — XIP from SDRAM/NOR
XN — EXECUTE NEVER

Regions where code CANNOT run

  • Peripheral (0x40000000–0x5FFFFFFF) — hardware registers only
  • External Device (0xA0000000–0xDFFFFFFF) — NOR/LCD in data mode
  • PPB (0xE0000000–0xE00FFFFF) — NVIC/SCB/ITM registers
  • Vendor-specific (0xE0100000–0xFFFFFFFF) — vendor extensions

Memory Map with XN Attributes

0xFFFFFFFF
↓
0xE0100000
Vendor-specific
XN · 511 MB
0xE00FFFFF
↓
0xE0000000
PPB — NVIC, SCB, SysTick, ITM
XN · 1 MB
0xDFFFFFFF
↓
0xA0000000
External Device — NOR/LCD
XN · 1 GB
0x9FFFFFFF
↓
0x60000000
External RAM — SDRAM/PSRAM
EXECUTABLE · 1 GB
0x5FFFFFFF
↓
0x40000000
Peripheral — GPIO, UART, SPI…
XN · 512 MB
0x3FFFFFFF
↓
0x20000000
SRAM — stack, heap, data
EXECUTABLE · 512 MB
0x1FFFFFFF
↓
0x00000000
Code — Flash, ROM, vector table
EXECUTABLE · 512 MB

XN regions cause a fault if the CPU tries to fetch instructions from them — regardless of what is stored there.

XN and security
XN is the hardware basis for Data Execution Prevention (DEP) — the same concept used in desktop operating systems to prevent buffer-overflow exploits from injecting shellcode into data memory. On embedded systems it prevents a corrupted function pointer from accidentally executing peripheral register data as instructions.

2. External RAM Region (0x60000000 – 0x9FFFFFFF)

This 1 GB region is reserved for external memory chips connected to the processor through a memory controller. Unlike the peripheral region, code can be executed from here — making it suitable for Execute-in-Place (XIP) applications where the program is too large to fit in on-chip flash.

COMMON USES

What connects here

  • External SDRAM — e.g. 8 MB IS42S16400 on STM32F429 Discovery
  • External SRAM / PSRAM — pseudo-static RAM, byte-accessible
  • External NOR flash — execute-in-place program storage
  • NAND flash — bulk storage (not XIP; must copy to RAM first)
IMPORTANT

Requires FMC / FSMC controller

Accessing the External RAM region requires the Flexible Memory Controller (FMC) or its older variant FSMC to be configured. The STM32F411 does not have FMC. Chips like STM32F407, STM32F429, and STM32H743 do. Without FMC, accesses into this region generate a bus fault.

Configuring external SDRAM on STM32F4 with FMC

/* Conceptual: FMC initialisation for IS42S16400J SDRAM (STM32F4 with FMC) */
/* NOT available on STM32F411 — shown for reference on larger STM32F4 devices */

#define FMC_BASE     0xA0000000UL   /* FMC register base */
#define SDRAM_BANK1  0xC0000000UL   /* External RAM bank 1 mapped here */

/* After FMC SDRAM init, use external memory like normal SRAM: */
uint32_t *ext_ram = (uint32_t *)SDRAM_BANK1;
ext_ram[0] = 0xDEADBEEF;   /* write to external SDRAM */
uint32_t val = ext_ram[0]; /* read back */

/* Place large arrays in external RAM with linker section attribute: */
__attribute__((section(".sdram")))
uint8_t video_frame_buffer[480 * 272 * 2];  /* 261 KB — too big for on-chip SRAM */
Execute-in-Place from external NOR flash
If you connect a NOR flash chip (not NAND — NAND cannot be XIP) to the FMC in NOR/PSRAM mode and map it into the External RAM region, the CPU can fetch instructions directly from it. This is called XIP. Throughput is limited by the FMC timing configuration — typically much slower than on-chip flash with the ART accelerator, so performance-critical code should still reside in on-chip flash.

3. External Device Region (0xA0000000 – 0xDFFFFFFF)

This 1 GB region is designated for external devices and shared memory — equipment that behaves like a peripheral rather than addressable memory. It is permanently marked XN (eXecute Never) and is also non-cacheable and non-bufferable by default, meaning every access goes directly to the external device with no memory-system optimisation.

TYPICAL USES

External Device region applications

  • NOR flash in data mode — reading flash ID / status registers
  • Parallel LCD interfaces — Intel 8080 / Motorola 6800 mode
  • Shared memory — dual-port RAM shared with another processor
  • FPGA-mapped registers — FPGA connected via FMC
XN

Never executable

Even if you wrote valid ARM Thumb-2 machine code into a device in this region, the processor would fault before executing it. The XN attribute is enforced by the hardware bus decoder before the instruction reaches the decode stage.

/* Parallel LCD using FMC in 8080 mode — External Device region */
/* Command register at bank 1, A16=0 (RS line tied to FMC address bit 16) */
#define LCD_CMD_ADDR   0x60000000UL   /* External Device region, bank 1 */
#define LCD_DATA_ADDR  0x60020000UL   /* A16=1 → data register           */

#define LCD_CMD    (*(volatile uint16_t *)LCD_CMD_ADDR)
#define LCD_DATA   (*(volatile uint16_t *)LCD_DATA_ADDR)

void lcd_write_reg(uint16_t reg, uint16_t value)
{
    LCD_CMD  = reg;     /* select register */
    LCD_DATA = value;   /* write value     */
}

void lcd_set_pixel(uint16_t color)
{
    LCD_DATA = color;   /* stream pixel data to GRAM */
}

4. Private Peripheral Bus Region (0xE0000000 – 0xE00FFFFF)

The Private Peripheral Bus (PPB) region is 1 MB and contains the ARM-defined system control registers that are identical on every Cortex-M chip. It is XN and always requires privileged access.

0xE000E000

System Control Space (SCS)

NVIC, SCB, SysTick, MPU, FPU control — the core OS-level control registers. 4 KB block starting at 0xE000E000.

0xE0000000

ITM, DWT, FPB

Instrumentation Trace Macrocell (ITM) for printf over SWO, Data Watchpoint and Trace (DWT) for cycle counting, and Flash Patch and Breakpoint (FPB) for hardware breakpoints.

XN + PRIV

Access rules

Unprivileged code that attempts to access any PPB address will cause a MemManage fault (or HardFault if MPU is off). Only Handler mode and Privileged Thread mode code can touch these registers.

/* Key PPB registers — same addresses on ALL Cortex-M devices */

/* SysTick (at 0xE000E010) */
#define SYST_CSR    (*(volatile uint32_t *)0xE000E010U)  /* Control & Status */
#define SYST_RVR    (*(volatile uint32_t *)0xE000E014U)  /* Reload Value     */
#define SYST_CVR    (*(volatile uint32_t *)0xE000E018U)  /* Current Value    */

/* NVIC (at 0xE000E100) */
#define NVIC_ISER0  (*(volatile uint32_t *)0xE000E100U)  /* Interrupt Set Enable   */
#define NVIC_ICER0  (*(volatile uint32_t *)0xE000E180U)  /* Interrupt Clear Enable */
#define NVIC_ISPR0  (*(volatile uint32_t *)0xE000E200U)  /* Interrupt Set Pending  */
#define NVIC_IPR0   (*(volatile uint32_t *)0xE000E400U)  /* Interrupt Priority     */

/* SCB (at 0xE000ED00) */
#define SCB_CPUID   (*(volatile uint32_t *)0xE000ED00U)  /* CPU ID register        */
#define SCB_AIRCR   (*(volatile uint32_t *)0xE000ED0CU)  /* App Int & Reset Control*/
#define SCB_SHCSR   (*(volatile uint32_t *)0xE000ED24U)  /* System Handler Control */

/* DWT Cycle Counter (at 0xE0001000) */
#define DWT_CTRL    (*(volatile uint32_t *)0xE0001000U)
#define DWT_CYCCNT  (*(volatile uint32_t *)0xE0001004U)

/* Read CPU ID: e.g. 0x410FC241 for Cortex-M4 r0p1 */
void print_cpu_id(void)
{
    uint32_t id = SCB_CPUID;
    /* bits [19:16] = Constant = 0xF
       bits [15:4]  = Part number = 0xC24 (M4) or 0xC23 (M3)
       bits [3:0]   = Revision                                */
}

5. What is AMBA?

AMBA stands for Advanced Microcontroller Bus Architecture. It is an open specification developed by ARM that defines standard on-chip bus protocols for connecting processor cores, memory, and peripherals inside a System-on-Chip (SoC).

Because AMBA is a published open standard, any chip designer (ST, NXP, Broadcom, Qualcomm, Apple) can use it to interconnect ARM IP blocks — the CPU, DMA controller, memory controller, and peripherals — without needing custom bus logic. This is a major reason why ARM-based chips dominate the embedded and mobile markets: the ecosystem of compatible IP is enormous.

AMBA Version History

VersionYearKey ProtocolsUsed in
AMBA 1 1996 ASB (Advanced System Bus), APB ARM7TDMI era chips
AMBA 2 1999 AHB (full), APB rev2 ARM9, ARM11 SoCs
AMBA 3 2003 AHB-Lite, APB3, AXI Cortex-M, Cortex-A9
AMBA 4 2010 AXI4, AXI4-Lite, AXI4-Stream, ACE Cortex-A15, A57, modern SoCs
AMBA 5 2013 CHI (Coherent Hub Interface) Cortex-A72, server-class chips

Cortex-M processors use AMBA 3: AHB-Lite for high-speed interfaces and APB for slower peripherals.

Why “Lite”?
Full AMBA AHB supports multiple bus masters competing for the bus (arbitration), split transactions, and retry responses — complex features needed for high-end SoCs. AHB Lite removes the multi-master arbitration logic, keeping only a single master (the CPU or DMA). This is sufficient for microcontrollers and results in much smaller silicon area.

6. AHB Lite — High-performance Bus

AHB Lite (AMBA High-performance Bus — Lite) is the primary bus used for the Cortex-M core’s main interfaces. It carries instruction fetches, data reads/writes to SRAM, and accesses to high-speed peripherals like GPIO, DMA, and RCC.

CHARACTERISTICS

AHB Lite properties

  • 32-bit address bus + 32-bit data bus
  • Single master per bus segment
  • Pipelined transfers (address phase + data phase)
  • Burst transfers: 4-beat, 8-beat, 16-beat
  • Operates at core clock frequency (up to 100 MHz on STM32F411)
  • Supports 8-, 16-, and 32-bit data widths
  • Low latency: typically 1–2 clock cycles per transfer
SIGNAL SUMMARY

Key AHB Lite signals

  • HADDR[31:0] — 32-bit address
  • HWRITE — read (0) or write (1)
  • HSIZE[2:0] — transfer size (byte/half/word)
  • HBURST[2:0] — single or burst type
  • HWDATA[31:0] — write data
  • HRDATA[31:0] — read data
  • HREADY — slave ready signal (wait states)
  • HRESP — OKAY or ERROR response

AHB Lite pipelined transfer timing

Two-transfer AHB Lite pipeline (no wait states)

Clock: │ 1 │ 2 │ 3 │ 4 │
Transfer A: │ ADDR │ DATA │ │ │
Transfer B: │ │ ADDR │ DATA │ │
Transfer C: │ │ │ ADDR │ DATA │

• Address and data phases overlap (pipelining)
• Each transfer takes 1 cycle when HREADY=1 (no wait states)
• SRAM and GPIO registers on STM32F411 respond in 1 cycle

On STM32F411, the AHB1 bus runs at the full system clock (100 MHz), giving a peak throughput of 400 MB/s for 32-bit transfers (100 MHz × 4 bytes). This is why GPIO toggling can be very fast on Cortex-M — a single 32-bit write to the GPIOA_BSRR register over AHB1 completes in one clock cycle.

/* GPIO BSRR: fastest way to set/clear GPIO bits — single AHB write */
#define GPIOA_BSRR  (*(volatile uint32_t *)0x40020018U)

/* Set PA5 HIGH  — one 32-bit AHB write, one clock cycle */
GPIOA_BSRR = (1U << 5);

/* Set PA5 LOW   — write to upper 16 bits (reset bits) */
GPIOA_BSRR = (1U << (5 + 16));

/* Toggle PA5 at maximum speed using ODR XOR */
#define GPIOA_ODR   (*(volatile uint32_t *)0x40020014U)
while (1) {
    GPIOA_ODR ^= (1U << 5);   /* Each iteration = 1 AHB write + loop overhead */
}

7. APB — Advanced Peripheral Bus

APB (Advanced Peripheral Bus) is a simpler, lower-speed bus designed for peripherals that do not need the full bandwidth of AHB. It has a smaller signal count, lower gate count, and lower power consumption — ideal for UART, I2C, basic timers, and watchdogs.

APB KEY FACTS

How APB differs from AHB

  • No pipelining — address and data in same cycle (simpler slaves)
  • No burst mode — one transfer at a time
  • Runs at a fraction of AHB clock (APB1 = AHB/2 on STM32F411)
  • Narrower signal set — no HBURST, HPROT, HMASTER
  • Minimum 2 clock cycles per transfer (setup + access)
  • Cannot be a bus master — only AHB is a master
APB SIGNALS

Simplified APB signal set

  • PCLK — peripheral clock (slower than HCLK)
  • PADDR[31:0] — address
  • PWRITE — read or write
  • PWDATA[31:0] — write data
  • PRDATA[31:0] — read data
  • PSEL — select this peripheral
  • PENABLE — enable strobe (2nd cycle)
  • PREADY — slave done (optional wait states)

APB Transfer Timing — Minimum 2 Cycles

PCLK: │ 1 (SETUP) │ 2 (ACCESS) │ idle │
PADDR: │──── valid ───│──── valid ───│──────────────│
PSEL: │ 1 │ 1 │ 0 │
PENABLE: │ 0 │ 1 │ 0 │
PREADY: │ – │ 1 │ – │

Setup phase: address and PSEL asserted, PENABLE=0
Access phase: PENABLE=1, slave samples/drives data, PREADY signals completion
APB clock prescaler on STM32F411
The RCC_CFGR register has two APB prescaler fields: PPRE1 (APB1 prescaler) and PPRE2 (APB2 prescaler). At 100 MHz HCLK with PPRE1=4 (divide by 2): APB1 PCLK = 50 MHz. With PPRE2=2 (no divide): APB2 PCLK = 100 MHz. Timer peripherals have an automatic ×2 multiplier if their APB prescaler is not 1, so TIM2–5 actually tick at 100 MHz even though APB1 runs at 50 MHz.

8. The AHB-to-APB Bridge

Slow APB peripherals cannot connect directly to the fast AHB bus — their timing requirements are incompatible. The AHB-to-APB bridge is a hardware module that translates between the two bus protocols, handling the clock domain crossing and protocol conversion transparently.

AHB to APB Bridge — Protocol Translation

CPU Core
AHB Master
AHB Lite
100 MHz
AHB-APB
Bridge
Protocol converter
APB
50 MHz
UART, I2C,
SPI, TIM2–5
APB Slaves

The bridge holds the AHB transaction in the setup phase while it completes the slower APB transaction. The CPU sees this as a wait state — HREADY goes low until the APB transfer is done.

From the programmer’s perspective, the bridge is completely invisible. You write to a UART register address and the hardware handles the protocol conversion. The only observable effect is that APB peripheral accesses take more clock cycles than AHB peripheral accesses.

/* USART2 is on APB1 at 0x40004400 — access goes through AHB-APB bridge */
#define USART2_SR   (*(volatile uint32_t *)0x40004400U)
#define USART2_DR   (*(volatile uint32_t *)0x40004404U)
#define USART2_BRR  (*(volatile uint32_t *)0x40004408U)
#define USART2_CR1  (*(volatile uint32_t *)0x4000440CU)

/* Writing to USART2_DR is an AHB transaction that the bridge translates
   to an APB write at 50 MHz. The CPU may insert 1 wait state.         */
void uart_send_byte(uint8_t byte)
{
    while (!(USART2_SR & (1U << 7)));  /* wait TXE (TX Empty) */
    USART2_DR = byte;
}

/* Compare: GPIOA on AHB1 — no bridge, no wait states */
#define GPIOA_ODR  (*(volatile uint32_t *)0x40020014U)
void gpio_toggle(void) { GPIOA_ODR ^= (1U << 5); }  /* 1 cycle */

9. Four Cortex-M Bus Interfaces

The Cortex-M processor core exposes four distinct AHB Lite bus interfaces to the outside world. Each has a specific purpose. This is why the core can do more than one thing at a time — instruction fetch and data access can happen simultaneously because they use different buses.

Cortex-M Core — Four Bus Interfaces

ARM Cortex-Mx
Core
I-CODE
D-CODE
System
PPB
I-CODE Bus
Instruction fetch + vector table reads from CODE region (0x00000000–0x1FFFFFFF). Read-only. 32-bit AHB.
D-CODE Bus
Data reads from CODE region — literal pools, constant tables in flash. Read-only. 32-bit AHB. Works simultaneously with I-CODE.
System Bus
All accesses to SRAM, Peripheral, External RAM, External Device, and vendor regions. Read/write. 32-bit AHB.
PPB Bus
Private Peripheral Bus — NVIC, SCB, SysTick, ITM, DWT, FPB. Uses APB protocol internally. Privileged access only.
Bus Interface Protocol Direction Targets Notes
I-CODE AHB-Lite Read only Code region (0x00–0x1F) Instruction fetches + vector table reads. Can burst 8 half-words.
D-CODE AHB-Lite Read only Code region (0x00–0x1F) Literal pool and rodata reads. Simultaneous with I-CODE.
System AHB-Lite Read/Write SRAM, Peripheral, Ext RAM, Ext Device, Vendor Main data bus. Also handles debug accesses from JTAG/SWD via AHB-AP.
PPB APB subset Read/Write PPB region (0xE0000000–0xE00FFFFF) System control only. Privileged access required. XN region.
Harvard-like architecture benefit
Because instruction fetch (I-CODE) and data access (D-CODE / System) use separate buses, the Cortex-M core can fetch the next instruction while simultaneously loading a constant from flash. This is the embedded equivalent of a Harvard architecture — maximising throughput without a full dual-port memory design.

10. Bus Matrix on STM32F411

On the STM32F411, a multi-layer AHB bus matrix connects multiple bus masters (CPU, DMA1, DMA2) to multiple slaves (flash, SRAM, peripheral buses). The matrix allows simultaneous accesses as long as two masters target different slaves.

STM32F411 Multi-layer AHB Bus Matrix (simplified)

Masters: │ Flash │ SRAM │ APB1 │ APB2 │ AHB1 │ USB │ ───────────────────┼────────┼────────┼────────┼────────┼────────┼────────┤ CPU I-CODE │ ✓ │ │ │ │ │ │ CPU D-CODE │ ✓ │ │ │ │ │ │ CPU System │ │ ✓ │ ✓ │ ✓ │ ✓ │ │ DMA1 │ ✓ │ ✓ │ ✓ │ │ │ │ DMA2 │ ✓ │ ✓ │ │ ✓ │ │ ✓ │

✓ = can access this slave. Two masters on different columns can transfer simultaneously — CPU fetching from flash while DMA copies from SRAM to UART, for example.

Practical example — DMA + CPU simultaneous access

/* DMA2 transfers data from SRAM to SPI1 (on APB2)
   while CPU continues executing from flash.
   No conflict: CPU uses I-CODE (flash), DMA2 uses System bus (SRAM→APB2). */

#define DMA2_S3CR    (*(volatile uint32_t *)0x40026458U)  /* DMA2 stream3 CR */
#define DMA2_S3NDTR  (*(volatile uint32_t *)0x4002645CU)  /* transfer count  */
#define DMA2_S3PAR   (*(volatile uint32_t *)0x40026460U)  /* peripheral addr */
#define DMA2_S3M0AR  (*(volatile uint32_t *)0x40026464U)  /* memory addr     */

static uint8_t spi_tx_buf[64];   /* in SRAM — filled by CPU */

void start_spi_dma(void)
{
    DMA2_S3PAR   = 0x4001300CU;           /* SPI1->DR address (APB2) */
    DMA2_S3M0AR  = (uint32_t)spi_tx_buf; /* source in SRAM           */
    DMA2_S3NDTR  = 64;
    DMA2_S3CR    = (3U << 25)            /* channel 3                */
                 | (1U << 6)             /* memory-to-peripheral     */
                 | (1U << 10)            /* memory increment         */
                 | (1U << 0);            /* enable stream            */

    /* CPU is now free — DMA handles the transfer via the bus matrix */
}

11. Performance Implications for Embedded Code

Understanding the bus architecture helps you write faster code and avoid hidden bottlenecks.

FAST

Zero-wait-state operations

  • Instruction fetch from flash (with ART accelerator enabled)
  • SRAM reads and writes over System bus
  • GPIO write via AHB1 (GPIOA_ODR, GPIOA_BSRR)
  • DMA transfers that do not contend with CPU
SLOWER

Higher latency operations

  • APB peripheral access (UART, I2C, SPI registers) — min 2 PCLK cycles
  • Flash reads without ART accelerator at high clock speeds
  • External RAM via FMC — depends on timing config (many wait states)
  • Bus contention between CPU and DMA on same slave

Clock enable: gating matters for APB peripherals

/* Peripheral clocks are DISABLED by default to save power.
   You must enable the clock before the first register access.
   Accessing a peripheral with its clock off returns 0x00000000 on reads
   and the write is silently ignored — a common beginner bug.            */

#define RCC_BASE      0x40023800UL
#define RCC_AHB1ENR  (*(volatile uint32_t *)(RCC_BASE + 0x30U))
#define RCC_APB1ENR  (*(volatile uint32_t *)(RCC_BASE + 0x40U))
#define RCC_APB2ENR  (*(volatile uint32_t *)(RCC_BASE + 0x44U))

void enable_clocks(void)
{
    /* AHB1: enable GPIOA (bit 0) and DMA2 (bit 22) */
    RCC_AHB1ENR |= (1U << 0) | (1U << 22);

    /* APB1: enable USART2 (bit 17) and I2C1 (bit 21) */
    RCC_APB1ENR |= (1U << 17) | (1U << 21);

    /* APB2: enable SPI1 (bit 12) and TIM1 (bit 0) */
    RCC_APB2ENR |= (1U << 12) | (1U << 0);

    /* After enabling, a small delay is required for clock propagation.
       Reading the ENR register once is sufficient on STM32F4:          */
    (void)RCC_AHB1ENR;
}
Bus clock must be enabled before first peripheral access
Forgetting to enable the peripheral clock is one of the most common bugs in bare-metal STM32 programming. The register address is valid, the CPU access succeeds from the bus’s perspective, but the peripheral’s internal logic is not clocked — it never sees the write. Always enable RCC clock before configuring any peripheral.

12. FAQ

Q: Can the CPU execute code from the Peripheral region (0x40000000)?
No. The Peripheral region is permanently XN (eXecute Never). If the program counter ever reaches a peripheral address — due to a bug like a corrupted function pointer — the processor raises a MemManage fault before executing anything. This is a hardware protection that cannot be disabled.
Q: What is the difference between AHB-Lite and full AHB?
Full AHB (AMBA 2) supports multiple bus masters competing for the same bus, requiring an arbiter that handles split transactions (a slave says “try again later”) and retry responses. AHB-Lite simplifies this by supporting only a single master per bus segment. The Cortex-M processor’s own bus interfaces are AHB-Lite. When multiple masters exist (CPU + DMA), the chip vendor adds a bus matrix that gives each master its own AHB-Lite channel to the slaves — no arbitration needed.
Q: Why does I-CODE exist separately from the System bus?
If instruction fetch shared the System bus with data accesses, the CPU would stall every time it needed to read a variable from SRAM while also trying to fetch the next instruction from flash. Having separate I-CODE and D-CODE buses for the Code region, plus the System bus for data memory, means the pipeline can stay full. This is the key reason Cortex-M processors can run efficiently at 100+ MHz without a cache.
Q: If APB is slower, why use it at all instead of putting everything on AHB?
Silicon area and power. Every AHB-capable slave needs a full set of AHB handshake logic (HREADY, HRESP, HTRANS, HBURST, HPROT…) which takes more transistors than the simplified APB interface. A UART that transfers at 115200 baud does not need 100 MHz burst access — putting it on APB saves area and reduces dynamic power consumption. The AHB-APB bridge adds two cycles of latency but that is negligible for slow peripherals.
Q: Can two DMA streams access the same peripheral simultaneously?
No. If two DMA streams try to access the same slave (e.g. both DMA1 and DMA2 writing to SRAM), the bus matrix arbitrates between them. One transfer completes first; the other waits. This is transparent to your code but can affect real-time timing. For performance-critical applications, assign DMA streams to different memory regions (e.g. DMA reads from flash, DMA2 reads from SRAM) to avoid matrix contention.
Q: How does the debugger (JTAG/SWD) access memory when the CPU is running?
The debug port (SWD or JTAG) connects to an AHB Access Port (AHB-AP) inside the Cortex-M’s CoreSight debug architecture. The AHB-AP is a bus master on the System bus. When the debugger wants to read a memory address, it initiates an AHB transaction via the AHB-AP. The bus matrix arbitrates between the CPU and the AHB-AP. This is why you can inspect memory in STM32CubeIDE while the program is running — the debugger gets bus cycles when the CPU is not using the System bus.

Next: STM32 Peripheral Register Programming

Now that the bus architecture is clear, the next lecture applies it directly — working through the full bit-field structure of GPIO, RCC, and NVIC registers, and showing how to configure peripherals from scratch without HAL.

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *