How to Implement a State Machine for BLE Device Firmware

Designing a state machine to manage BLE connection and pairing flows in firmware

You’ve written this code. A bool is_connected, a bool is_bonded, a bool advertising_paused, maybe a pairing_in_progress for good measure. Then a callback fires mid-pairing, the central drops, and your firmware ends up in a state that no flag combination actually describes. The device hangs, or worse, it re-advertises while the BT stack still thinks it’s connected.

This is flag soup, and almost every BLE peripheral firmware starts there. The fix is modeling the thing you already have: a finite state machine. The BLE spec practically draws it for you.

By the end of this article you’ll have a working BLE state machine scaffold for nRF52 with Zephyr, covering advertising, connection, pairing, and bonding, with a dispatcher you can drop into your project this afternoon.

Defining the States

A BLE peripheral has maybe 7 states worth naming. Write them down, give them an enum, stop pretending they’re implicit.

typedef enum {
    ST_IDLE,
    ST_ADVERTISING,
    ST_CONNECTED_UNBONDED,
    ST_PAIRING,
    ST_CONNECTED_BONDED,
    ST_DISCONNECTING,
    ST_ERROR,
} sm_state_t;

Here’s the lifecycle in ASCII so you can paste it into a comment block and stop arguing with yourself in code review:

   ┌──────┐  start_adv   ┌─────────────┐  central_connects   ┌────────────────────┐
   │ IDLE ├─────────────►│ ADVERTISING ├────────────────────►│ CONNECTED_UNBONDED │
   └──────┘              └─────┬───────┘                     └──────┬─────────────┘
       ▲                       │ adv_timeout                        │ pair_req
       │                       ▼                                    ▼
       │                    ┌──────┐                            ┌─────────┐
       │                    │ IDLE │                            │ PAIRING │
       │                    └──────┘                            └────┬────┘
       │                                                             │ bonded
       │                                                             ▼
       │  disconnect      ┌───────────────┐   keys_loaded   ┌──────────────────┐
       └──────────────────┤ DISCONNECTING │◄────────────────┤ CONNECTED_BONDED │
                          └───────────────┘                 └──────────────────┘

Two things to notice. First, CONNECTED_UNBONDED and CONNECTED_BONDED are different states, not flags on a “connected” state. The legal GATT operations differ between them, and that distinction belongs in the state rather than a guard scattered across 12 characteristic write handlers. Second, DISCONNECTING is a real state. Calling bt_conn_disconnect() doesn’t drop you to IDLE instantly. Ignoring this is where most “phantom advertising” bugs come from.

Defining the Events

Events arrive from three places: the BT stack (callbacks), timers (advertising timeout, pairing timeout), and the application itself (user pressed a button, sensor wants to push a notification).

typedef enum {
    EV_APP_START_ADV,
    EV_BT_CONNECTED,
    EV_BT_DISCONNECTED,
    EV_BT_PAIR_REQ,
    EV_BT_SECURITY_CHANGED,
    EV_BT_BOND_LOADED,
    EV_TIMER_ADV_TIMEOUT,
    EV_APP_REQUEST_DISCONNECT,
    EV_ERROR,
} sm_event_t;

typedef struct {
    sm_event_t type;
    void *ctx;  /* optional payload, e.g. bt_conn* */
} sm_msg_t;

The rule: callbacks post events, they don’t do work. The Zephyr BT thread calls your connected() callback; that callback’s only job is to drop a message into the queue and return. The SM thread does the actual work, in its own context, where it’s safe to call bt_le_adv_stop() and friends.

The Dispatcher

You have two reasonable choices: a giant nested switch (switch(state) { case X: switch(event) {...} }) or a transition table. Switches are fine up to about 5 states. Past that, you’ll find yourself scrolling 400 lines to figure out what EV_BT_DISCONNECTED does in ST_PAIRING. A table puts every transition on one line.

typedef bool (*guard_fn)(const sm_msg_t *);
typedef void (*action_fn)(const sm_msg_t *);

typedef struct {
    sm_state_t from;
    sm_event_t evt;
    sm_state_t to;
    guard_fn   guard;   /* NULL = always allow */
    action_fn  action;  /* NULL = no side effect */
} transition_t;

static const transition_t TABLE[] = {
    { ST_IDLE,                EV_APP_START_ADV,       ST_ADVERTISING,        NULL,         act_start_adv },
    { ST_ADVERTISING,         EV_BT_CONNECTED,        ST_CONNECTED_UNBONDED, NULL,         act_on_connect },
    { ST_ADVERTISING,         EV_TIMER_ADV_TIMEOUT,   ST_IDLE,               NULL,         act_stop_adv },
    { ST_CONNECTED_UNBONDED,  EV_BT_PAIR_REQ,         ST_PAIRING,            NULL,         act_accept_pairing },
    { ST_CONNECTED_UNBONDED,  EV_BT_BOND_LOADED,      ST_CONNECTED_BONDED,   guard_keys_ok, NULL },
    { ST_PAIRING,             EV_BT_SECURITY_CHANGED, ST_CONNECTED_BONDED,   guard_encrypted, act_on_bonded },
    { ST_PAIRING,             EV_BT_DISCONNECTED,     ST_IDLE,               NULL,         act_cleanup },
    { ST_CONNECTED_BONDED,    EV_APP_REQUEST_DISCONNECT, ST_DISCONNECTING,   NULL,         act_disconnect },
    { ST_DISCONNECTING,       EV_BT_DISCONNECTED,     ST_IDLE,               NULL,         act_cleanup },
    /* catch-all */
    { ST_CONNECTED_UNBONDED,  EV_BT_DISCONNECTED,     ST_IDLE,               NULL,         act_cleanup },
    { ST_CONNECTED_BONDED,    EV_BT_DISCONNECTED,     ST_IDLE,               NULL,         act_cleanup },
};

The dispatcher itself is about 15 lines:

static sm_state_t state = ST_IDLE;

static void sm_dispatch(const sm_msg_t *msg) {
    for (size_t i = 0; i < ARRAY_SIZE(TABLE); i++) {
        const transition_t *t = &TABLE[i];
        if (t->from != state || t->evt != msg->type) continue;
        if (t->guard && !t->guard(msg))             continue;

        LOG_INF("[SM] %s --(%s)--> %s",
                state_name(state), event_name(msg->type), state_name(t->to));

        state = t->to;
        if (t->action) t->action(msg);
        return;
    }
    LOG_WRN("[SM] no transition: state=%s evt=%s",
            state_name(state), event_name(msg->type));
}

Every transition is data. Adding a new event handled in a new state is one row in the table.

Wiring It Into Zephyr

You need one thread, one message queue, and a handful of callbacks that do nothing but post messages.

K_MSGQ_DEFINE(sm_q, sizeof(sm_msg_t), 16, 4);
K_THREAD_STACK_DEFINE(sm_stack, 1024);
static struct k_thread sm_thread;

static void sm_thread_fn(void *a, void *b, void *c) {
    sm_msg_t msg;
    while (1) {
        k_msgq_get(&sm_q, &msg, K_FOREVER);
        sm_dispatch(&msg);
    }
}

static void on_connected(struct bt_conn *conn, uint8_t err) {
    sm_msg_t m = { .type = err ? EV_ERROR : EV_BT_CONNECTED, .ctx = conn };
    k_msgq_put(&sm_q, &m, K_NO_WAIT);
}

static void on_disconnected(struct bt_conn *conn, uint8_t reason) {
    sm_msg_t m = { .type = EV_BT_DISCONNECTED, .ctx = conn };
    k_msgq_put(&sm_q, &m, K_NO_WAIT);
}

static void on_security_changed(struct bt_conn *conn, bt_security_t level,
                                enum bt_security_err err) {
    sm_msg_t m = { .type = EV_BT_SECURITY_CHANGED, .ctx = conn };
    k_msgq_put(&sm_q, &m, K_NO_WAIT);
}

BT_CONN_CB_DEFINE(conn_callbacks) = {
    .connected        = on_connected,
    .disconnected     = on_disconnected,
    .security_changed = on_security_changed,
};

The flow looks like this:

   ┌────────────────┐    ┌──────────────┐    ┌──────────────────┐
   │ BT stack       │───►│ event queue  │───►│ SM thread        │
   │ (callbacks)    │    │  (k_msgq)    │    │ (dispatcher loop)│
   └────────────────┘    └──────────────┘    └──────────────────┘
          ▲                                          │
          │            bt_le_adv_start /             │
          └──────────  bt_conn_disconnect  ◄─────────┘
                       (action functions)

1 KB of stack is usually enough for the SM thread. Measure it under your worst-case action: settings load or key derivation are the usual suspects. Action functions are the only place that touches bt_le_adv_start, bt_conn_disconnect, settings storage, anything blocking or stack-heavy. Never call those from inside a BT callback. The host thread won’t thank you.

Bonding and Connection Params Belong in the State

The reason CONNECTED_BONDED is a state and not a flag: GATT operations on sensitive characteristics should only succeed when the link is encrypted with bonded keys. Put that check in a guard, not in every characteristic handler.

static bool guard_bonded_write(const sm_msg_t *msg) {
    return state == ST_CONNECTED_BONDED;
}

Connection parameter updates are a judgment call. If they’re brief and well-bounded, treat them as transient and don’t bother with a state. If they affect your power budget, going from 7.5 ms to 400 ms intervals before a firmware update is the canonical example, model it as a substate or a separate flag inside CONNECTED_BONDED. The point is to make the decision deliberately, not by accident in a callback.

Why Not One RTOS Task Per State?

You may have seen the “thread per state” pattern recommended on forums. For BLE on an nRF52, don’t.

Three reasons. RAM: each thread stack is 512 B to 1 KB minimum, and with 7 states that’s potentially 7 KB of stack for something a single thread handles in 1 KB. Synchronization becomes inter-thread handoffs, which means mutexes, condition variables, and a much larger surface area for races. And the BT host stack is itself effectively single-threaded; fighting that with parallel state-owners gains you nothing and costs you determinism.

A single SM thread is deterministic, debuggable, and matches the stack’s own model. Miro Samek’s Practical UML Statecharts makes the formal argument if you want it. The QP framework is excellent, but for a single BLE peripheral it’s heavier than the hand-rolled table above.

Debugging and Observability

Log every transition. One line, structured:

[SM] ADVERTISING --(BT_CONNECTED)--> CONNECTED_UNBONDED
[SM] CONNECTED_UNBONDED --(BT_PAIR_REQ)--> PAIRING
[SM] PAIRING --(BT_DISCONNECTED)--> IDLE

Add counters: how many times have you entered each state, and what’s the max time spent there? A device stuck in PAIRING for 30 seconds is a bug; a device stuck there for 30 minutes is a fleet-wide alert. Memfault has written well about this on the observability side, and it’s worth your time.

At fleet scale, this log stream is the difference between a product that works and a product you can actually operate. Telemetry showing “0.4% of devices spent >10s in DISCONNECTING last week” catches silicon-revision bugs before they become RMAs. If you’re moving toward connecting devices over a wide-area BLE network, the Hubble device SDK docs cover how this kind of architecture extends to a transmit-side BLE stack.

Ship This Week

  • Name your states. 7 is plenty for most BLE peripherals.
  • Define events from BT callbacks, timers, and app intents. Callbacks post events, nothing else.
  • Use a transition table once you pass 5 states. The dispatcher is 15 lines.
  • One SM thread, one k_msgq, action functions own all blocking BT calls.
  • Bonded vs. unbonded is a state, not a flag. Guards enforce sensitive operations.
  • Log every transition. Counters and dwell times pay for themselves the first time a device misbehaves in the field.

Operating thousands of devices reliably is a different problem from making one work, and it’s the one Hubble works on. If you’re building BLE products that need to phone home over a global network, take a look at what we’re doing at hubble.com.


Hubble Network lets BLE devices reach the cloud globally without provisioning gateways or managing connectivity infrastructure. See how it works →