The Hybrid MCU Battery Trap: When Splitting BLE Work Across Cores Actually Costs You More Current

You split the work. Sensor DSP went to the application core, the BLE stack stayed on the network core. Two cores, each doing less, each awake for shorter bursts. Battery life should go up.
Then you hooked up a power profiler and watched average current climb by 15% over the single-core build you were trying to beat.
That’s not a measurement error, and it isn’t a bug in your sleep config. It’s the predictable outcome of an accounting mistake almost everyone makes on their first hybrid MCU BLE design: optimizing active current while ignoring the retention and IPC tail that a second active domain drags behind it.
The vendor slides say the second core is free power savings. Sometimes it is. Often it’s a bill you didn’t read. Here’s a framework to tell which one you’re building before you commit firmware architecture.
Why Average Current Is the Only Number That Matters
Your battery doesn’t care how fast a core runs a burst. It cares how much charge leaves the cell per second, averaged over everything.
Battery life ≈ Capacity(mAh) / I_avg(mA)
I_avg = Σ(I_state × t_state) / t_totalThat second line is the whole game. Average current is the time-weighted sum of every state your system sits in: active bursts, radio TX/RX, idle-waiting, deep sleep with RAM retention, all of it.
Two cores active in parallel for 400 µs can beat one core active sequentially for 900 µs, because the parallel version reaches deep sleep sooner. That’s the argument the marketing framing rests on, and it’s genuinely true in the right conditions.
The trap is that engineers stop the math at I_active. They see the shorter active window and declare victory, never adding the current spent waking a second domain, keeping a second RAM domain retained, and burning idle cycles while one core waits on the other.
Average current optimization lives or dies in that tail, not in the burst. A design that halves active time but adds a steady 8 µA retention floor and 20 wake events per second can easily lose. If you’re tuning a BLE battery budget, this is the layer past connection-interval math, where hybrid designs quietly overspend. The terrestrial transmission guidance covers the layer beneath this one.
The Hidden Costs of Splitting
The “free core” story omits four line items. I’ll use nRF5340 specifics where they help, but flag what’s universal to any dual-core part (STM32H7, Ambiq Apollo, and friends behave the same way in principle, only the numbers move).
IPC wake events. Every time the app core hands data to the network core, both cores have to be awake for the handoff. On the nRF5340 that’s the IPC peripheral plus a shared-memory region. Each message is a paired wake: one core writes, one reads and acknowledges. If your protocol needs a round trip, you pay twice. MCU peripheral offload discussions almost always skip this cost.
Shared RAM retention. When one core sleeps but must preserve state, you retain its RAM. Split the work and you now retain two working sets across two power domains instead of one. Retained RAM isn’t free. It’s a small but constant leakage current that runs 24/7. A larger retained footprint raises your idle floor, exactly the number that dominates average current in a low-duty-cycle device.
Wake-synchronization idle. One core finishes early and spins or idles waiting for the other to hand off or acknowledge. That idle time is invisible in a naive budget because nobody schedules a line item for “waiting.” It still draws current.
Duplicated clock domains. Two active cores can mean two high-frequency clock trees running when a single-core design would run one. Peripheral clocks don’t gate themselves out of politeness.
| Cost source | Single-core | Split (naive) | Notes |
|----------------------|-------------|---------------|--------------------------|
| Active bursts | Sequential | Parallel | Split may shorten wall t |
| IPC wake events | None | Per handoff | Both cores wake |
| Shared RAM retention | 1 domain | 2 domains | Higher idle floor |
| Wake-sync idle | N/A | Spin/wait | Hidden average-current |None of these are large on their own. Added up across a connection interval, repeated thousands of times a day, they’re the difference between hitting budget and missing it.
The Decision Framework
Before you split, run the design through three questions, in order.
1. Duty cycle overlap. Do both workloads need to be active at the same time, or can they interleave?
If your sensor processing and your BLE activity naturally happen at different moments, splitting has room to help, because each core can sleep while the other works. If they overlap heavily (app logic that must run during every connection event), you’ve got two cores awake simultaneously anyway. You’ve bought parallelism you can’t sleep through. Low overlap favors splitting; high overlap usually kills it.
2. IPC frequency. How often must the cores talk per connection interval?
Every handoff is a paired wake. A design that exchanges one message per second is cheap. A design where the app core reacts to every connection event, feeds data back per packet, and coordinates timing is chatty. High-frequency IPC erodes the savings from shorter active bursts, and past some threshold it reverses them outright. Count your messages per interval before you assume the split is free.
3. Independent deep-sleep feasibility. Can each core reach its deepest sleep state without waiting on the other?
This is the one people miss. If the app core can’t drop into System OFF because it’s holding state the network core will need in 5 ms, then the network core’s duty cycle is setting the app core’s sleep floor. One domain drags the other up. You wanted two independent sleepers; you built a pair handcuffed together.
Split BLE work across cores?
|
+------------+------------+
| Do workloads overlap |
| in time? |
+------------+------------+
no | | yes
v v
+-------------+ +------------------+
| IPC freq | | Likely NO gain — |
| per event? | | overlap keeps |
+------+------+ | both awake |
low | high +------------------+
v \
+-----------+ \--> +-------------------+
| Each core | | IPC tail erodes |
| deep-sleep| | savings — measure |
| alone? | +-------------------+
+-----+-----+
yes | no
v \
+---------+ \--> +------------------+
| SPLIT | | Drag: one core |
| WINS | | raises other's |
+---------+ | floor. Reconsider|
+------------------+The heuristic in one line: splitting pays off when the workloads are asynchronous, low-IPC, and independently sleepable. Miss any one of those and you’re probably spending current, not saving it.
Worked Example: nRF5340
Two scenarios, same part. Treat the numbers as conceptual illustrations, not spec guarantees. Verify against the current nRF5340 datasheet and Nordic’s online power profiler before you budget against them.
Where splitting wins. A device wakes every few seconds to run a chunk of sensor DSP, then advertises or maintains a loose connection. The DSP runs on the app core as a self-contained burst, hands a small result to the network core once, then the app core drops to System OFF. The network core handles steady BLE on its own. IPC is one message per cycle. Both cores sleep deep between events. Here the parallelism shortens wall-clock active time and the retention/IPC tail stays small, so I_avg comes out lower than the single-core equivalent, plausibly by a meaningful margin.
Where splitting loses. App logic that must run on every connection event: parse the incoming packet, decide a response, feed it back before the next event. Now you’re doing paired IPC wakes every connection interval, the app core can never reach its deepest state (it’s always about to be needed again), and both cores share overlapping awake windows. The retention floor doubles, wake-sync idle piles up, and average current lands above the single-core build. The split cost you.
Same silicon. The difference is entirely in the three answers.
When to Un-Split
If the framework says you’re dragging, the fix is often to stop splitting.
Consolidate everything onto the network core and put the app core into its deepest sleep, powered down and out of the average entirely. On the nRF5340 you can run a surprising amount of application logic alongside the BLE stack on the network core when the app workload is light. A core that’s fully dark contributes almost nothing to I_avg, which beats a core that’s half-awake shuttling messages.
The reverse works too when the app workload dominates and BLE is intermittent: keep the app core home, wake the network core only for radio windows. One dark core often costs less than two busy ones.
If you’re building the firmware to test this, the Hubble Device SDK and its reference apps give you a working BLE stack to profile against.
Building This Into Your Requirements
The second core is a tool, not a discount. It saves current only when it lets both domains sleep deeper or reach sleep faster, and it costs current whenever it keeps one domain awake waiting on the other.
Put the three-question gate in your design review before firmware architecture is locked: check duty cycle overlap, count IPC frequency per interval, and confirm each core can deep-sleep independently. Then stop trusting the math and measure. Capture real average current with a power profiler across a full connection interval, including the retention tail, not just the active burst.
The burst is the part everyone sees. The tail is the part that drains the battery.
Hubble Network lets you reach the satellite backhaul from the same low-power BLE stack you’re already profiling, so radio windows stay short and cores stay asleep. See how it works →