How to Implement DMA Transfers for Efficient Peripheral Communication

Configuring DMA channels on a Cortex-M microcontroller for SPI and UART peripheral transfers

At 115200 baud, a UART fires an interrupt roughly every 87 microseconds. Each ISR takes maybe 40 cycles to context-switch into, grab the byte, stuff it into a buffer, and return. Sounds cheap.

But your control loop runs at 1 kHz. Your SPI sensor dumps 512 bytes every 10 ms. The I2C display wants attention too. Suddenly 30% of your CPU time is spent shuffling bytes from point A to point B. Your real-time deadline starts slipping. Your jitter gets ugly.

DMA exists to solve exactly this: let dedicated hardware move the data while the CPU does the actual thinking. One interrupt per block instead of one per byte. The CPU sets up the transfer, walks away, and gets tapped on the shoulder when it’s done.

This guide gives you the vendor-agnostic mental model for DMA on Cortex-M. We’ll use STM32 as an occasional concrete example, but the architecture transfers to NXP, Nordic, TI, or anything else with a DMA controller. Once you have the model, any vendor’s reference manual becomes readable.

Polling vs. Interrupts vs. DMA: A Progression of CPU Freedom

Think of these three paradigms as a spectrum of how much the CPU has to babysit data movement.

┌──────────────┬────────────────┬─────────────────┬───────────────────┐
│              │   Polling      │   Interrupt     │   DMA             │
├──────────────┼────────────────┼─────────────────┼───────────────────┤
│ CPU during   │ Busy-waiting   │ Free between    │ Free for entire   │
│ transfer     │ (100% used)    │ bytes (ISR/byte)│ block transfer    │
├──────────────┼────────────────┼─────────────────┼───────────────────┤
│ Interrupt    │ None           │ 1 per byte      │ 1 per block       │
│ frequency    │                │ (high)          │ (low)             │
├──────────────┼────────────────┼─────────────────┼───────────────────┤
│ Complexity   │ Low            │ Medium          │ Medium-High       │
├──────────────┼────────────────┼─────────────────┼───────────────────┤
│ Best for     │ Simple, slow   │ Moderate rates, │ High throughput,  │
│              │ peripherals    │ small transfers │ continuous stream  │
└──────────────┴────────────────┴─────────────────┴───────────────────┘

The key insight: DMA doesn’t replace interrupts. It reduces their frequency. Instead of 128 ISRs for 128 bytes, you get 1 or 2. Your ISR goes from “grab a byte” to “process a buffer.”

When is DMA overkill? If you’re reading a single temperature sensor over I2C once per second, an interrupt-driven transfer is perfectly fine. DMA earns its keep when you’re streaming: high-baud UART, fast SPI, continuous ADC sampling, or any time byte-by-byte handling threatens your timing budget.

The Building Blocks of DMA on Cortex-M

    ┌─────────┐         ┌──────────────────┐
    │  CPU    │◄───────►│                  │
    │ Cortex-M│         │   Bus Matrix     │
    └─────────┘         │  (Arbitration)   │
                        │                  │
    ┌─────────┐         │                  │        ┌──────────┐
    │  DMA    │◄───────►│                  │◄──────►│  SRAM    │
    │Controller│  req    │                  │        └──────────┘
    └────┬────┘         └──────┬───────────┘
         │                     │
         │ DMA requests        │
    ┌────┴─────────────────────┴────┐
    │     Peripheral Bus (APB)      │
    ├──────┬──────┬──────┬──────────┤
    │ UART │ SPI  │ I2C  │  ADC     │
    └──────┴──────┴──────┴──────────┘

The DMA controller and the CPU both talk to memory through the bus matrix. An arbiter decides who gets bus access on any given cycle. This means DMA isn’t free; it’s cheaper. The CPU might stall for a cycle when the DMA controller hogs the bus. In practice the contention is usually negligible, but on very tight loops it can matter.

The DMA controller has multiple channels (or “streams” on STM32F4/F7). Each channel can be wired to specific peripherals. Which peripheral maps to which channel is vendor-defined and fixed in silicon, so it’s the one part you’ll always need to look up in the datasheet.

When a peripheral has data ready (say, a UART RX byte arrives), it raises a DMA request to its assigned channel. The DMA controller reads from the peripheral’s data register and writes to your RAM buffer, all without the CPU lifting a finger. When the configured number of transfers finishes, the DMA controller fires an interrupt to let the CPU know.

Priority arbitration handles contention between channels. If your UART RX and SPI RX both request DMA simultaneously, the controller uses the priority you’ve configured (typically 4 levels) to decide who goes first. Equal priorities resolve by channel number.

Anatomy of a DMA Transfer: The Five Configuration Fields

Regardless of vendor, every DMA setup requires the same 5 pieces of information. Once you internalize this pattern, you can configure DMA on any Cortex-M chip.

// Conceptual DMA setup, not vendor-specific code
dma_config.channel      = PERIPHERAL_TO_DMA_CHANNEL;  // vendor-defined mapping
dma_config.direction    = PERIPHERAL_TO_MEMORY;
dma_config.src_addr     = &UART->DATA_REG;            // peripheral data register
dma_config.dst_addr     = &rx_buffer[0];               // RAM buffer
dma_config.transfer_len = 128;                         // number of transfers
dma_config.data_width   = BYTE;                        // 8-bit peripheral
dma_config.mode         = CIRCULAR;
dma_config.priority     = HIGH;

dma_enable_interrupt(HALF_TRANSFER | TRANSFER_COMPLETE);
dma_start(&dma_config);

Let’s walk through each piece.

1. Channel/stream selection. You look this up in the reference manual. On STM32, for example, USART2_RX might be DMA1, Stream 5, Channel 4. On an NXP LPC, it’s a different numbering scheme, but the concept is identical.

2. Direction. Three options: peripheral-to-memory (sensor reads, UART RX), memory-to-peripheral (UART TX, DAC output), or memory-to-memory (copying buffers without CPU involvement). Pick the wrong one and you’ll read garbage or write into a peripheral’s control registers.

3. Source and destination addresses. For peripheral-to-memory, the source is the peripheral’s data register (a fixed address), the destination is your RAM buffer (incrementing). The DMA controller needs to know which side increments and which stays fixed.

4. Transfer count and data width. Count is how many data units to move. Width matters: if you’re reading a 16-bit ADC but configured for byte-width, you’ll get half the data and corrupt the other half. Match the data width to the peripheral’s register size.

5. Mode. Normal (one-shot) means the DMA stops after the configured count. Circular means it wraps back to the start and keeps going. Pick the wrong mode and you either stop receiving or overwrite data you haven’t processed.

Two Core Patterns That Cover 90% of Use Cases

One-Shot (Normal Mode): SPI Flash Read

You need to read 256 bytes from an SPI flash chip. Here’s the flow:

  1. Configure DMA: peripheral-to-memory, source is SPI data register, destination is your buffer, count is 256, normal mode.
  2. Start the DMA. Start the SPI transaction.
  3. The CPU is now free. Run your control loop, service other tasks, go to sleep.
  4. 256 bytes later, the transfer-complete interrupt fires.
  5. In the ISR, set a flag or post to a queue. Process the buffer in your main loop.

The critical rule: don’t touch the destination buffer while the DMA is writing to it. The compiler doesn’t know about DMA. It might reorder reads, cache values in registers, or optimize away what it thinks are dead writes. We’ll come back to this in the pitfalls section.

One-shot mode requires re-arming. After the transfer completes, the DMA channel is done. If you need another transfer, you must reconfigure (or at least re-enable) it.

Circular + Double Buffer: UART RX Streaming

This is the workhorse pattern for continuous data. You’re receiving UART data at high speed and can’t afford to miss a byte.

        rx_buffer[0 .. 127]
    ┌───────────────┬───────────────┐
    │   HALF A      │   HALF B      │
    │  [0 .. 63]    │  [64 .. 127]  │
    └───────┬───────┴───────┬───────┘
            │               │
            ▼               ▼
     HT interrupt     TC interrupt
     (half-transfer)  (transfer-complete)
            │               │
            ▼               ▼
     CPU processes A   CPU processes B
     while DMA fills B while DMA fills A
            │               │
            └───────────────┘
              Continuous loop

Set up a 128-byte buffer in circular mode. Enable both the half-transfer (HT) and transfer-complete (TC) interrupts. When HT fires, DMA has filled bytes 0–63 and is writing into 64–127. You safely process the first half. When TC fires, the buffer wraps; DMA starts filling 0–63 again while you process the second half.

This ping-pong pattern means you never miss data and the CPU only wakes up twice per 128 bytes. Compare that to 128 ISRs in the interrupt-driven approach.

If you’re building on a platform like Hubble’s and need to manage BLE advertising packets efficiently alongside UART or SPI traffic, this same double-buffer pattern applies. The Hubble terrestrial SDK’s advertising packet documentation shows how packet data is structured; pairing that with DMA-driven peripheral I/O keeps your timing predictable.

The Pitfalls That Will Waste Your Weekend

Cache coherency (Cortex-M7 and up)

The Cortex-M7 has a data cache. DMA writes directly to SRAM, bypassing it. If the CPU reads from cached memory, it sees stale data. You have two options: place your DMA buffers in a non-cacheable memory region (via the MPU), or call SCB_InvalidateDCache_by_Addr() before reading. The MPU approach tends to win out because it’s less error-prone.

This doesn’t affect Cortex-M0/M3/M4 since they don’t have data caches.

Alignment and data width mismatches

Say you’ve configured DMA for word-width (32-bit) transfers, but your buffer starts at an odd address. On some platforms, that’s a hard fault. On others, silent data corruption. Always align your DMA buffers to at least the transfer width. Most compilers offer __attribute__((aligned(4))) or equivalent.

Volatile and compiler reordering

Declaring your buffer as volatile isn’t the full answer. volatile prevents the compiler from optimizing away reads, but it doesn’t prevent instruction reordering by the CPU. On Cortex-M, you’ll sometimes need a data synchronization barrier (__DSB()) after checking a DMA completion flag to ensure the CPU sees the final memory state. The pattern is: check flag, barrier, then read buffer.

The first transfer works, the second one doesn’t

One-shot DMA transfers don’t restart automatically. If your code expects continuous data and you forgot to re-enable the channel in your transfer-complete ISR, everything after the first block vanishes. This is especially sneaky because your initial testing looks perfect. You only notice the problem once you’re streaming real traffic.

DMA and low-power modes

Gate the bus clock to save power and the DMA controller stops too. On many chips, you need to explicitly keep the DMA and peripheral clocks running during sleep modes. Check your MCU’s power management chapter; the DMA section alone won’t mention this.

Your Integration Checklist

  1. Profile first. Measure your ISR load with a logic analyzer or cycle counter. If you’re spending less than 5% of CPU on data movement, DMA probably isn’t worth the complexity.
  2. Check the vendor DMA channel map. Find your peripheral in the reference manual’s DMA request table. Confirm there’s no conflict with another peripheral you need.
  3. Allocate buffers correctly. DMA-accessible memory, aligned to transfer width, in a non-cacheable region if you’re on M7.
  4. Choose your mode. One-shot for discrete transactions (SPI flash reads, command/response). Circular for continuous streams (UART RX, ADC sampling).
  5. Wire up your ISRs. Transfer-complete is mandatory. Half-transfer if you’re doing the ping-pong pattern. Error interrupts (transfer error, FIFO overrun) should at least log and recover.
  6. Test under real conditions. Caches on. Optimization at -O2 or your release level. Full clock speed. DMA bugs love to hide behind debug builds and slow clock configs.
  7. Handle re-arming. For normal mode, explicitly restart or reconfigure the channel in your completion ISR.

If you’re building a custom device on Hubble’s network and want to see how peripheral communication fits into a full firmware stack, the Hubble device integration guide provides a solid reference for how these pieces connect at the system level.

Applying This to Any New MCU

Next time you open a new MCU’s reference manual, go straight to the DMA chapter. Find the request mapping table (which peripheral maps to which channel). Find the configuration register descriptions. You’ll recognize every field: channel, direction, addresses, count, mode. You already know what each one means.


Hubble Network connects your devices from anywhere—no gateways, no line-of-sight—over standard Bluetooth. See how it works →