How to Use Keyword Spotting on a $2 MCU for Voice-Activated BLE Devices

Running wake-word detection on a low-cost microcontroller to trigger BLE devices by voice

A CR2032 holds about 220 mAh. Divide that by a year of runtime and you get 25 µA of headroom for everything: your mic, your DSP, your neural net, your BLE radio. Naive always-on audio + neural inference burns through that in a weekend.

And yet “Hey [Product]” voice activation is exactly the kind of feature that sells a wearable, remote, or smart tag. Can you do keyword spotting on the MCU you already have, or do you need to bolt on a DSP or NPU?

The honest answer: yes, you can run keyword spotting on an MCU like the nRF52840 (Cortex-M4F, 64 MHz, 256 KB RAM, integrated PDM, around $2 in volume) and still hit 6 to 12 months on a coin cell. But only if you treat this as a power-architecture problem first, and an ML problem second. Here’s how.

The Two-Stage Listener Architecture

The inference itself isn’t the budget killer. The always-on PDM mic and feature extraction loop is.

Stage the work instead. A cheap voice activity detector (VAD) runs at very low duty. It wakes the real keyword spotter only when there’s energy in the right band. The KWS model wakes BLE only on a confirmed match.

   ┌─────────┐    ┌──────────┐    ┌──────────┐    ┌─────────┐
   │  PDM    │───▶│   VAD    │───▶│   KWS    │───▶│  BLE    │
   │  Mic    │    │ (~5 µA)  │    │ (~200µA  │    │ Action  │
   │ 16 kHz  │    │ always-on│    │  burst)  │    │         │
   └─────────┘    └──────────┘    └──────────┘    └─────────┘
        │              │                │
        └─ DMA ────────┘                │
                  Sleep between windows─┘

The VAD can be as simple as a windowed energy threshold over a couple of mel bins, or a tiny NN (a few hundred parameters). Either way, target sub-10 µA average. The KWS model is allowed to be expensive in bursts, because it runs rarely.

This staging matters because wake word embedded systems live or die on duty cycle. A 40 ms inference burst at 3 mA averages to 5 µA if it fires once a minute. The same burst running every 100 ms averages to 1.2 mA, which is 50× your budget.

The Feature Pipeline

Sample at 16 kHz mono through the PDM peripheral with EasyDMA. The CPU stays asleep between DMA half/full interrupts. That’s the single biggest power lever in the whole design.

Frame the audio in 30 ms windows with 20 ms stride. A 1 second clip becomes 49 frames. For each frame, run an FFT, apply a mel filterbank, and take the log.

Log-mel beats MFCC for this. The extra DCT in MFCC costs cycles and gives you almost nothing in accuracy for small-vocabulary KWS. Use 10 mel coefficients. CMSIS-DSP’s arm_rfft_fast_f32 plus a hand-rolled filterbank gets you there in well under a millisecond per frame on an M4F.

Quantize the feature output to int8 before it hits the model. Your rolling feature buffer is around 2 KB. Keep it in a static array, not on the heap.

The pattern: PDM DMA half-full interrupt fires, the ISR kicks a worker that extracts features for the just-filled buffer, then WFI. The CPU spends most of its life asleep between these wakeups.

Choosing and Training the Model

For tinyml voice workloads on a Cortex-M4, the depthwise-separable CNN (DS-CNN) from Google’s “Hello Edge” paper remains the best accuracy-per-kilobyte target I know of. For 1 or 2 keywords plus “unknown” and “silence” classes, you can land:

  • Under 50 KB of int8 weights
  • Under 30 KB tensor arena
  • Roughly 5 to 10 MMACs per inference
  • 92 to 95% accuracy on Google Speech Commands

Alternatives worth knowing: a tiny GRU gets smaller but typically loses 1 to 3 points of accuracy. A plain 1D-CNN is simpler still and fine if you only need 1 keyword. Tiny Transformers are interesting in papers but rarely win at this scale once you account for the attention memory footprint.

Train on Google Speech Commands v0.02. Augment heavily: room impulse responses, background noise, and (critically) recordings of your product’s own mechanical and electrical noise. Hard negatives win field accuracy. Skip them and your demo unit fails its first dinner party.

Post-training int8 quantization should cost you less than 2% accuracy. If it costs more, you’ve got a calibration problem, not a model problem.

Convert the trained model with xxd -i model.tflite > model_data.cc and link it as a const array in flash.

TFLM Integration on nRF52

Pull TFLM in as a Zephyr module or build it as a static library. Be ruthless about ops: only register what the model actually uses. For DS-CNN that’s DEPTHWISE_CONV_2D, CONV_2D, FULLY_CONNECTED, SOFTMAX, and RESHAPE. Every op you don’t register saves flash.

Link the CMSIS-NN kernels. This isn’t optional. The reference kernels are 3 to 10× slower on Cortex-M, and that ratio shows up directly in your power budget.

static tflite::MicroMutableOpResolver<5> resolver;
resolver.AddDepthwiseConv2D();
resolver.AddConv2D();
resolver.AddFullyConnected();
resolver.AddSoftmax();
resolver.AddReshape();

tflite::MicroInterpreter interp(model, resolver,
                                arena, kArenaSize);
interp.AllocateTensors();

Size the arena empirically. Start at 16 KB, call interp.arena_used_bytes() after AllocateTensors(), then trim to actual + a small margin. Guessing leads to silent allocation failures at runtime.

The inference call site belongs in a dedicated thread or work item that the VAD wakes. Don’t run inference in an ISR. Don’t share the arena with anything else.

Power Budgeting in Practice

Numbers from a Power Profiler Kit II on a reference nRF52840 build, your mileage will vary:

Subsystem          Avg Current   % Budget
─────────────────  ───────────   ────────
PDM mic + DMA          18 µA       38%
VAD (every 30ms)        8 µA       17%
KWS inference burst     5 µA       11%
BLE adv (1 Hz)         12 µA       25%
MCU sleep + leakage     4 µA        9%
─────────────────  ───────────   ────────
TOTAL                  47 µA      100%

At 47 µA average against 220 mAh, you get roughly 6 months of life. Want a year? Trim the BLE advertising rate, or drop to 8 kHz sampling. The 8 kHz move costs 2 to 4 points of accuracy but cuts mic and feature power by about 30%.

The tradeoff knobs, in order of impact:

  1. VAD threshold. A tighter threshold means fewer false KWS wakes, which dominates burst-phase power.
  2. KWS confirmation window. Requiring two consecutive positive frames before triggering BLE cuts false positives without much accuracy hit.
  3. Sample rate. 16 kHz is the standard, but 8 kHz works for short keywords.
  4. PDM clock and decimation. The PDM peripheral’s clock rate trades quality for current. Profile it on your actual board.

For voice activated BLE designs, the BLE side of the budget often surprises people. A 1 Hz connectable advertiser can easily out-cost your entire audio pipeline. Spend most of the device’s life in non-connectable beacon mode, and escalate only on a confirmed wake word. Hubble’s terrestrial transmission guidance covers the advertising-side tradeoffs in more detail if you’re connecting these devices to a wide-area network.

Pitfalls and Field Notes

A few things that have cost real teams real weeks:

  • False wakes from TV and radio voices. Train with hard negatives sampled from broadcast audio.
  • DC bias and PDM clock jitter. What works in a Python simulation can fall apart on hardware. Capture raw PDM on the actual board early.
  • Flash wear from logging. Tempting to log every inference for debugging. Don’t. Use a RAM ring buffer and dump it only on a trigger.
  • BLE radio coupling into mic traces. PCB layout matters. Keep PDM lines short, ground them well, and check spectrogram noise floor when the radio is transmitting.
  • Bootloader and OTA budget. Your KWS model is going to want updates. Plan dual-bank flash from the start. The device integration guides cover provisioning patterns that play well with OTA flows.

When to Step Up to a Bigger Chip

The $2 Cortex-M4 tier covers 1 to maybe 5 keywords with reasonable accuracy. Once you need:

  • More than 10 keywords or short command recognition
  • Speaker identification
  • A model that won’t fit in 80 KB or needs more than 20 MMACs/s sustained

look at nRF5340 (dual-core M33), Ambiq Apollo4 (sub-threshold magic, very low µA), or chips with dedicated NPUs like the Syntiant NDP family or Alif Ensemble. The cost step is real (roughly 2 to 4× BOM), but so is the capability.

Start With the µA Table, Not the Model

If you’re scoping a voice-activated BLE product today, start with the power budget, not the model. Sketch the µA table for your duty cycle before you train anything. Pick the chip that fits the budget with margin. Then build the two-stage listener, prove the VAD power on bench hardware, and only then start tuning model accuracy.

The teams that ship this successfully treat KWS as a systems problem. The model is maybe 20% of the work. The other 80%: feature pipeline, DMA, sleep states, BLE coexistence, and field-realistic training data. That’s where the coin-cell math actually gets won.


Hubble Network keeps always-on BLE devices discoverable at global scale without burning your power budget on connection overhead. See how it works →