How to Use STM32WBA LPBAM for Background Sensor Reads Without Waking the CPU

A BLE sensor node sampling an accelerometer at 1 Hz spends maybe 1.8 ms of CPU time per second doing actual work. The other 998 ms it’s asleep. Sounds great, until you measure average current and find the wake-sleep transitions are eating most of your coin-cell budget.
You can shave the ISR. You can crank up the clock to finish faster. You can move to a lower Stop mode. None of it changes the fundamental cost: the M33 has to wake up just to babysit an I²C transaction the DMA could’ve done by itself.
This is where STM32WBA LPBAM earns its keep. The trick: keep the Cortex-M33 in Stop2 across hundreds of sensor reads in a row, with the GPDMA and I²C3 quietly handling everything in autonomous mode. Below is a working recipe (CubeMX setup, code, expected numbers) plus an honest accounting of where LPBAM stops being worth the fiddliness.
Assumes you’ve already picked a Stop mode and know your way around CubeMX. This article is part of the STM32WBA Development pillar and slots in after the “Choosing Between Stop0, Stop1, and Stop2” piece.
What LPBAM Actually Is
LPBAM (Low-Power Background Autonomous Mode) ties three things together. There’s the GPDMA, a subset of peripherals that ST has wired up to run without a clock from the CPU, and linked-list descriptors placed in SRAM2. Each descriptor tells the DMA what to do next when the previous transaction finishes.
On the STM32WBA, the LPBAM-capable peripherals are I²C3, SPI3, LPUART1, ADC4, and LPTIM1/3. GPDMA1 channels 0 through 7 can operate in autonomous mode. Everything else (I²C1, SPI1, USART1, the main timers) needs the CPU domain alive.
Peripherals talk to peripherals without the CPU. A trigger source fires, the GPDMA walks a linked list, each node describes one transfer, and the result lands in SRAM2. The M33 is asleep the whole time.
If you’ve worked with Nordic’s PPI/DPPI on nRF52/nRF54, the idea is similar. ST’s AN5870 is the authoritative reference and worth reading once you’ve finished here.
WBA-specific gotcha: the 2.4 GHz radio shares some power-domain plumbing with the rest of the chip. Your sensor pipeline and BLE link aren’t fully independent in every Stop2 sub-state. Verify on the bench, especially around radio events.
Example Project: 1 Hz LSM6DSO Read in Stop2
Concrete scope:
- LSM6DSO accelerometer on I²C3
- LPTIM1 generates a 1 Hz trigger from LSI
- GPDMA1 executes: write the register pointer (0x28, OUTX_L_A), then read 6 bytes, then store into a ring buffer in SRAM2
- CPU wakes only when the buffer hits N samples, or on a BLE event
Data flow:
┌─────────┐ trigger ┌──────────┐ I2C ┌─────────┐
│ LPTIM1 ├────────────►│ GPDMA ├────────►│ I2C3 │
│ (1Hz) │ │ (LL desc)│ │ (auto) │
└─────────┘ └────┬─────┘ └────┬────┘
│ │
▼ ▼
┌─────────┐ ┌──────────┐
│ SRAM2 │◄────────┤ LSM6DSO │
│ buffer │ └──────────┘
└─────────┘
CPU (M33): Stop2, wakes every N samplesCode targets STM32CubeWBA 1.4.x; LPBAM API signatures have shifted between SDK releases, so check yours.
CubeMX Configuration Walkthrough
Open the .ioc for your Nucleo-WBA55CG (or custom board) and work through these in order.
1. I²C3. Enable it, set the speed (400 kHz Fast Mode works for LSM6DSO), and in the peripheral properties panel mark it as Autonomous Mode. This is the single most-missed step. Without it the peripheral clock dies in Stop2 and your DMA transactions never complete.
2. LPTIM1. Clock from LSI (32 kHz nominal). Enable autonomous mode here too. Set period for 1 Hz (so 32768 counts with prescaler 1, or adjust as needed). Configure the update event as a DMA trigger source.
3. GPDMA1. Pick a channel (channel 0 is fine). Enable Linked-List Mode. Crucially, place the descriptor array in SRAM2. SRAM1 may be powered down in Stop2 depending on retention config, and a descriptor in unpowered RAM is the kind of bug that wastes a full afternoon. In your linker script, define a .lpbam_section mapped to SRAM2 and decorate the descriptor array with __attribute__((section(".lpbam_section"))).
4. Power. Stop2 entry, SRAM2 retention ON, ICACHE off during Stop2, GPIO retention configured for the I²C3 SCL/SDA pins. The I²C pins must keep their pull-ups across Stop2 or the bus will glitch on the first autonomous transaction after entry. This bites people who copied a U5 example without checking pin retention.
5. NVIC. Enable the GPDMA channel interrupt, but only for the buffer-full event (last node in the queue). Per-transfer interrupts defeat the whole point: every interrupt wakes the M33.
6. Generate code. CubeMX will create lpbam_*.c files. Leave the auto-generated init alone and write your queue configuration in user code sections.
Three traps to call out:
- Descriptors not in SRAM2. Silent failure, GPDMA reads zeros, sensor never responds.
- Calling
HAL_I2C_Mem_Readfrom your old code path. The HAL function blocks and wakes the CPU. You need the LPBAM advanced-layer API (ADV_LPBAM_I2C_MM_Read_SetDataQ). - I²C pin retention forgotten. Bus locks up after first Stop2 exit.
Building the Linked-List Queue
The advanced-layer API does most of the heavy lifting. Skeleton:
#include "stm32_adv_lpbam_i2c.h"
__attribute__((section(".lpbam_section")))
static DMA_QListTypeDef SensorReadQ;
__attribute__((section(".lpbam_section")))
static uint8_t sample_buffer[16][6]; // 16 samples, 6 bytes each
static volatile uint8_t sample_idx = 0;
void LPBAM_BuildSensorQueue(void)
{
ADV_LPBAM_I2C_MM_FullAdvConf_t cfg = {0};
cfg.AddressingMode = I2C_ADDRESSINGMODE_7BIT;
cfg.DevAddress = LSM6DSO_I2C_ADDR << 1;
cfg.MemAddress = LSM6DSO_OUTX_L_A;
cfg.MemAddSize = I2C_MEMADD_SIZE_8BIT;
cfg.Size = 6;
cfg.pData = sample_buffer[0];
cfg.WakeupIT = LPBAM_I2C_IT_NONE; // no wake mid-queue
ADV_LPBAM_I2C_MM_Read_SetFullQ(I2C3, &SensorReadQ, &cfg);
// Link LPTIM1 update event as the trigger source
HAL_DMAEx_List_LinkQ(&hgpdma1_ch0, &SensorReadQ);
HAL_DMAEx_List_Start(&hgpdma1_ch0);
}
void EnterLowPower(void)
{
HAL_PWREx_EnterSTOP2Mode(PWR_STOPENTRY_WFI);
// Execution proceeds here only on buffer-full or BLE event.
}The buffer-full ISR is short: bump a flag, notify your BLE task that a batch of samples is ready, re-arm the queue head pointer back to sample_buffer[0], and return. The M33 goes back into Stop2 within microseconds.
For deeper ring-buffer patterns (double-buffered DMA, two queues alternating), see ST’s LPBAM examples folder in the WBA Cube package: github.com/STMicroelectronics/STM32CubeWBA.
Benchmark: LPBAM vs. CPU-Wake Baseline
The numbers below are modeled from datasheet currents and the timing budget for a 6-byte read at 400 kHz: Stop2 typical IDD, M33 active IDD at 16 MHz, and I²C3 active current. They match the order of magnitude reported in AN5870. Run your own Power Profiler Kit II measurement before quoting these to a hardware lead.
Setup: Nucleo-WBA55CG, VDD = 3.0 V, LSI as LPTIM source, 1 Hz sampling, 6-byte burst read, radio off.
Scenario | Avg Current | Notes
----------------------------|-------------|------------------
CPU wake + HAL_I2C_Mem_Read | ~22 µA | M33 active ~1.8ms/s
LPBAM in Stop2 | ~3.5 µA | CPU never wakes
LPBAM, wake every 16 samples| ~3.7 µA | realistic BLE nodeState-timing intuition:
CPU-WAKE BASELINE (per 1s window)
CPU: ▁▁▁▁▁▁▁▁█▁▁▁▁▁▁▁▁▁▁█▁▁▁▁▁▁▁▁▁▁█▁
I2C: ░░░░░░░░█░░░░░░░░░░█░░░░░░░░░░█░
LPBAM IN STOP2
CPU: ▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁
I2C: ░░░░░░░░█░░░░░░░░░░█░░░░░░░░░░█░Roughly 6x reduction on the sensor path. On a 230 mAh CR2032, that’s the difference between months and years of field life, assuming the radio doesn’t dominate your total budget. Whether you realize that gain depends on your BLE duty cycle, which is its own accounting problem.
For BLE sensor nodes that report aggregated readings (asset trackers, environmental beacons, infrequent telemetry), the savings are real. We’ve written separately about how Bluetooth-only devices can reach satellite scale through aggregation patterns, and LPBAM is one of the building blocks that makes the math work.
When to Skip LPBAM
Don’t reach for it if:
- Your sensor logic needs a branch mid-transaction (read status, then decide what to read next). The CPU path is simpler.
- The peripheral isn’t on the autonomous list. I²C1 and SPI1 aren’t LPBAM-capable on WBA, period.
- Sample rate is above ~1 kHz, where the CPU is awake for processing anyway and you’ve optimized a non-bottleneck.
- You’re still in early prototyping. Debug pain (no breakpoints inside a running queue, descriptor bugs that fail silently) isn’t worth it until you have a power budget to defend.
Aligning LPBAM With Your BLE Connection Interval
Three pieces: autonomous peripherals, GPDMA linked lists in SRAM2, and disciplined Stop2 entry. Get those right and the pipeline runs without M33 involvement for hundreds of cycles at a time.
The next step for most BLE designs is aligning LPBAM wake events with the BLE connection interval, so you batch sensor data and radio activity into the same wake window. That’s the follow-on article in this pillar. If you’re cross-shopping silicon, the U5 vs. WBA comparison piece covers the tradeoffs between higher peripheral count (U5) and integrated radio (WBA).
Hubble Network enables BLE-to-cloud connectivity at global scale, so devices designed around tight power budgets like this stay reachable without gateway infrastructure. See how it works →