Your BLE MCU Can Run Edge AI — Until It Thermal-Throttles and Misses Advertising Windows

The demo looked great. You flashed the model onto an eval board, wired it up on the bench, and watched inference run alongside a live BLE connection for 20 minutes without a hiccup. Latency numbers matched the datasheet. Ship it.
Then the sealed prototype came back from field testing, and the phone app couldn’t find the device half the time. Connections dropped for no obvious reason. Nobody could reproduce it at their desk.
The eval board left something out. It ran cold, open to the air, on a bench supply. Your product runs hot, sealed, on a coin cell that sags under load. That gap between peak and sustained is where edge AI on a BLE MCU quietly falls apart.
This piece walks the causal chain from continuous inference to broken radio timing, shows you how to measure each link, and covers what you can do about it, including the point where the honest answer is to stop doing it on-die.
Part 1 — The Throttle-to-Jitter Chain
Why sustained inference is a different animal
TOPS ratings and single-shot latency get measured under the friendliest conditions possible: a cold die, an open board, a stiff bench supply, and one inference at a time with idle gaps in between. MLPerf Tiny and vendor enablement pages report those peak figures honestly. But peak isn’t what your product does.
Continuous inference means continuous current draw and continuous heat, with no idle window for the die to shed it. Seal that MCU in a plastic enclosure with no airflow and the heat has nowhere to go. The accelerator that sipped power in a 50 ms burst on the bench now runs back-to-back, and the thermal picture changes completely. Sustained inference power is a different metric from peak, and almost nobody publishes it.
How heat becomes a clock problem
Modern MCUs with ML accelerators (Cortex-M55 plus Ethos-U class parts, dual-core nRF54-class devices, and similar) all share two budgets: a thermal envelope and a supply budget. DVFS and thermal governors exist to keep the chip inside both.
When die temperature climbs past a threshold, the governor throttles the core clock to cut power. A core that was running at, say, 128 MHz drops to keep the junction temperature in check. Verify the exact thresholds and step-downs against your own part’s reference manual, because they vary a lot across families.
A second trigger often bites first on a coin cell. A CR2032 has real internal resistance, and a heavy inference burst pulls a current spike that drops the terminal voltage. If that droop crosses a brownout or DVFS threshold, you get the same clock throttle, no heat required. Two independent triggers, one symptom.
How a slower clock breaks BLE timing
BLE advertising and connection events are hard real-time. The link layer has to wake, service the radio, and transmit or receive inside a window defined by the Bluetooth Core Spec, and those windows don’t wait for you.
A throttled core has fewer cycles per unit time. An inference pass that hogs the CPU can starve the link layer outright. Either way, the stack shows up late to its own scheduled event.
Miss it, and you get advertising jitter: the actual advertising interval drifts off its nominal value, hurting discoverability. During a connection, a missed connection event burns a chance to exchange data, and enough consecutive misses trips the supervision timeout and drops the link.
Sustained inference
│ (continuous current + heat, no idle recovery)
▼
Die temp / supply stress rises
│ (DVFS + thermal governor responds)
▼
Core clock throttled
│ (fewer cycles for the BLE link layer)
▼
Radio event serviced late / missed
│
▼
Advertising jitter → missed connection events → timeoutsWhy it’s intermittent and hard to catch
Thermal buildup takes time. A 15-minute bench session on a cool, open board may never reach the throttle threshold. The same firmware in a sealed enclosure at 40°C ambient crosses it after an hour.
Battery sag compounds it. A fresh cell holds voltage. A cell near end-of-charge has higher internal resistance and droops harder under the same load. The bug prefers hot rooms, sealed cases, and old batteries, exactly the combination you never test at your desk. It surfaces in the field and nowhere else.
Part 2 — Diagnosing edge AI BLE MCU timing failures
What to measure, and how
You can’t fix what you can’t see on one timeline. Capture these four signals together and line them up:
| Signal | Tool | What to watch for |
|---|---|---|
| Die temperature | On-chip temp sensor, or a thermal camera on an open/decapped sample | Steady climb toward the governor threshold |
| Instantaneous current | Power profiler (Nordic PPK2, Joulescope) or a shunt + scope | Inference bursts, and voltage droop on a coin cell |
| Core clock frequency | Governor logs or DVFS/debug registers | A step-down that lines up with the temp rise |
| Advertising interval | Logic analyzer or BLE sniffer on the PHY | Interval drifting off nominal, missed events |
The sequence you’re hunting for shows up in order: temperature rises, clock steps down, advertising interval starts drifting. See that and you’ve confirmed the chain rather than chasing a ghost. If current droop trips it first, you’ll see the voltage sag and clock step without much heat.
Watch the supply rail and the clock at the same time. The two triggers look identical downstream, and you want to know which one you’re fighting.
Reproducing field conditions
Bench-at-room-temperature is the whole problem. Recreate what the device actually sees:
- Test inside the real, sealed enclosure, not on an open board.
- Bring the chamber (or the room) to rated maximum ambient.
- Run from a coin cell near end-of-charge, or a source that mimics its internal resistance, not a stiff bench supply.
- Run inference continuously, at production duty, for longer than you think you need to.
- Log all four signals for the full soak, not a 5-minute snapshot.
Let it cook until the die reaches thermal steady state. That’s when the throttle shows up, and that’s the condition your customer lives in.
Part 3 — Mitigation Levers
Duty cycling the inference
The simplest lever is to stop running inference all the time. Batch it, or gate it behind a trigger, and let the die cool between runs. You trade latency for thermal recovery.
If your application tolerates a result every few seconds instead of continuously, duty cycling can drop average power enough to keep the governor out of it. Watch your edge AI battery budget while you tune this, because the duty cycle you pick sets both the thermal profile and the coin-cell drain.
Scheduling inference around radio events
Even at low duty, an inference pass that lands on top of a scheduled advertising or connection event will still cause a miss. Fix it with cooperative scheduling.
Give the BLE stack strict priority, and structure inference so a pass never straddles a radio window. Break long inferences into chunks that yield the CPU before each scheduled event, and let the link layer preempt cleanly. On an RTOS, that means the radio’s timing-critical tasks sit above the inference task in priority, and inference is written to be interruptible. Get this wrong and you’ll see jitter even when the chip is stone cold, because the problem there is CPU contention, not heat.
Power gating and model shrinking
Gate the accelerator’s power domain between runs so it isn’t leaking or idling hot. Every microwatt you save is heat you don’t have to dump.
On the model side, quantize (int8 instead of float) and prune to cut both the energy per inference and the time the core spends busy. A smaller model dwells less and heats less. These help, but the returns diminish. Past a point you’re shaving milliwatts off a workload that’s fundamentally too heavy for the envelope.
The hard ceiling
Every lever above moves the problem around. If the workload is genuinely sustained and heavy, no amount of scheduling or quantizing gets you under the envelope. You’re rationing cycles between two things that both want them all.
The honest architectural move is to split the work.
ON-DIE (light, latency-critical) | OFFLOAD (heavy, sustained)
inference wake-word / trigger / gate | full model / vision / batch
power microwatts, bursty | handled off-device
radio keeps timing guarantees | ─Keep the light, latency-critical inference on the MCU: a wake-word, a trigger, a gate that decides whether anything interesting is happening. Push the heavy, sustained model somewhere with power and thermal headroom. When you offload heavy inference to an edge tier, the device stops fighting itself, and the radio gets its timing budget back. The heavy model runs where there’s a real power supply and a heatsink, and the MCU does the one thing it has to do on schedule.
Design for the radio first
On a BLE MCU, timing is the whole job. The link layer has hard deadlines it can’t negotiate, and every cycle you spend on inference is a cycle the radio might need.
Budget your thermal and power headroom for the radio before you spend any of it on on-device ML. Decide early, while you can still change silicon, what has to run on-die and what belongs off it. Do that before you commit to a part, and you won’t be reproducing an intermittent connection bug in a hot room six months from now.
Hubble Network keeps your MCU’s radio on schedule by moving sustained inference off the device, so timing budgets stay intact under thermal load. See how it works →