How to Use STM32Cube.AI to Convert a TensorFlow Model for STM32

Converting a TensorFlow neural network to run on an STM32 microcontroller using STM32Cube.AI

A 250 KB neural network can run 10x faster on the STM32N6 than on an STM32H7, while burning a fraction of the power. The catch? You have to get it through STM32Cube.AI first, and the toolchain has opinions about your model that TensorFlow never warned you about.

This guide walks you through the full pipeline: a pretrained Keras keyword-spotting model, quantized to INT8, converted to C, deployed to an STM32N6570-DK, and benchmarked on the Neural-ART NPU. Allow yourself an afternoon. By the end you’ll know which errors to expect, how to read the analyzer report, and what kind of inference latency to plan for in your product.

[TF/Keras .h5] -> [TFLite INT8] -> [STM32Cube.AI] -> [C code] -> [STM32N6]
     train          quantize         analyze/gen       build       deploy

Why STM32Cube.AI and Why the N6

STM32Cube.AI (shipped as the X-CUBE-AI expansion pack) takes a trained TensorFlow, Keras, ONNX, or TFLite model and emits optimized C code that links against ST’s neural network runtime. It handles memory layout, operator fusion, and on the STM32N6 specifically, it compiles INT8 graphs down to the Neural-ART accelerator, ST’s NPU running alongside the 800 MHz Cortex-M55.

The N6 matters because most STM32 parts run inference on the CPU. The N6 has dedicated silicon (rated by ST at 600 GOPS) so models that would take 40+ ms on an M7 finish in single-digit milliseconds. That’s the difference between a feasibility study and a shippable product.

This tutorial uses STM32Cube.AI 10.0 (verify the current version on ST’s developer zone before you start, the CLI flags shift release to release).

A Quick Primer on TF Models and Quantization

STM32Cube.AI accepts Keras .h5, SavedModel directories, TFLite .tflite, and ONNX. For the N6 you want .tflite with full INT8 quantization, because that’s what the Neural-ART can actually accelerate. A float32 model will still convert and run, but it’ll fall back to the CPU and you’ll have paid for an NPU you aren’t using.

Quantization maps your float32 weights and activations to 8-bit integers. You get a 4x size reduction and (on the N6) a massive speedup, in exchange for a small accuracy hit, typically 0.5 to 2% for well-behaved models.

Two flavors:

  • Post-training quantization (PTQ): convert after the fact using a small representative dataset. Easy, usually good enough.
  • Quantization-aware training (QAT): simulate quantization during training. More work, better accuracy preservation. Use this if PTQ costs you more than 1% accuracy on a metric you care about.

We’ll use PTQ here.

Prerequisites

Hardware:

  • STM32N6570-DK discovery board
  • USB-C cable, ST-LINK works over the same port

Software:

  • STM32CubeIDE 1.16+
  • X-CUBE-AI 10.0 (install via CubeMX: Help → Manage embedded software packages → STMicroelectronics → X-CUBE-AI)
  • Python 3.9+, TensorFlow 2.15+
  • The stm32ai CLI. It ships inside the X-CUBE-AI install, usually under ~/STM32Cube/Repository/Packs/STMicroelectronics/X-CUBE-AI/<version>/Utilities/<os>/.

Sample model: a 12-class keyword-spotting CNN trained on the Google Speech Commands dataset. Input is a 49x10 MFCC, output is a 12-way softmax. Float model is ~380 KB. You can swap in your own; the workflow doesn’t change.

Step 1: Quantize the TensorFlow Model

Load the Keras model and convert it to INT8 TFLite. The representative dataset is the part everyone forgets, and skipping it silently produces a float model dressed up in TFLite clothing.

import tensorflow as tf
import numpy as np

model = tf.keras.models.load_model("kws_cnn.h5")

def representative_dataset():
    # 100-200 real samples from your validation set
    for sample in val_samples[:200]:
        yield [sample.astype(np.float32).reshape(1, 49, 10, 1)]

converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8

tflite_model = converter.convert()
with open("kws_int8.tflite", "wb") as f:
    f.write(tflite_model)

Two things to check:

  1. File size dropped roughly 4x (mine went from 380 KB to 98 KB).
  2. Accuracy on your validation set is within tolerance. Run the .tflite through tf.lite.Interpreter and compare top-1 against the float model. A 0.7% drop is fine. A 15% drop means PTQ isn’t working and you need QAT or a closer look at which layers are causing the damage.

Common pitfall: if you omit representative_dataset, the converter falls back to dynamic-range quantization (weights INT8, activations float). The file will be small but the N6 NPU won’t touch it.

Step 2: Analyze the Model with STM32Cube.AI

The CubeMX GUI works, but the CLI is reproducible and scriptable, which is what you want once you’re iterating.

stm32ai analyze \
  --model kws_int8.tflite \
  --target stm32n6 \
  --st-neural-art default@user_neuralart.json \
  --workspace ./workspace \
  --output ./output

The report shows:

  • MACCs: multiply-accumulate operations per inference. Useful for back-of-envelope timing.
  • ROM/RAM footprint: weights go to flash, activations to RAM. The N6 has 4.2 MB of contiguous SRAM, generous by MCU standards.
  • Per-layer breakdown: which ops mapped to the Neural-ART, which fell back to the M55. You want as few CPU fallbacks as possible.
  • Complexity warnings: layers flagged as expensive or unsupported.

If you see fallbacks, the analyzer tells you why. Usually it’s an op the Neural-ART doesn’t support, or an INT8 layer that didn’t quantize cleanly. ST publishes an op support matrix per release; check yours against it.

Step 3: Validate Numerical Equivalence

Code generation rewrites operators, fuses layers, and repacks memory, so confirm the deployed model still produces the same outputs as the original .tflite.

Desktop validation:

stm32ai validate --model kws_int8.tflite --target stm32n6 --mode host

On-target validation (board plugged in):

stm32ai validate --model kws_int8.tflite --target stm32n6 --mode target

The tool feeds inputs through both the reference model and the deployed network. It reports cross-accuracy, RMSE, and max error per output. For an INT8 KWS model I’d expect cross-accuracy >99% and max error <2 LSB. If it’s worse, something is off in the conversion (often a fused activation that didn’t round-trip).

Step 4: Generate Code and Integrate

stm32ai generate \
  --model kws_int8.tflite \
  --target stm32n6 \
  --st-neural-art default@user_neuralart.json \
  --name network \
  --output ./generated

You’ll get network.c, network.h, network_data.c, plus the Neural-ART compiled blob. Drop these into your CubeIDE project under Middlewares/ST/AI/. CubeMX can do this for you if you enabled X-CUBE-AI on the project.

Minimal inference call in main.c:

#include "network.h"
#include "network_data.h"

static ai_handle network = AI_HANDLE_NULL;
AI_ALIGNED(32) static ai_u8 activations[AI_NETWORK_DATA_ACTIVATIONS_SIZE];

AI_ALIGNED(32) static ai_i8 in_data[AI_NETWORK_IN_1_SIZE];
AI_ALIGNED(32) static ai_i8 out_data[AI_NETWORK_OUT_1_SIZE];

static ai_buffer ai_input[AI_NETWORK_IN_NUM]  = AI_NETWORK_IN;
static ai_buffer ai_output[AI_NETWORK_OUT_NUM] = AI_NETWORK_OUT;

void ai_init(void) {
    ai_network_create_and_init(&network, activations, NULL);
    ai_input[0].data  = AI_HANDLE_PTR(in_data);
    ai_output[0].data = AI_HANDLE_PTR(out_data);
}

void ai_run(void) {
    // populate in_data with INT8 MFCC features
    ai_network_run(network, ai_input, ai_output);
    // out_data now holds 12 INT8 class scores
}

Wire ai_run() to your audio frontend and dequantize the output if you want probabilities.

Real Benchmarks on the STM32N6

Numbers below are representative for the KWS model described above, measured on an STM32N6570-DK with code in internal flash and activations in internal SRAM. Yours will vary with model topology, clock settings, and memory placement. Always measure on your own hardware before committing to a spec sheet.

Inference Latency (ms, lower = better)
M55 @ 800 MHz |####################### 38.4
Neural-ART    |### 4.2
              +-----------------------------
              0    10    20    30    40
BackendLatencyRAMFlashNotes
M55 CPU only~38 ms84 KB102 KBINT8, CMSIS-NN kernels
Neural-ART NPU~4 ms68 KB102 KBINT8, NPU accelerated

A few honest caveats. Benchmark steady state, not the first inference, which includes one-time setup. Memory placement matters a lot: putting weights in external Octo-SPI flash can double latency for weight-heavy layers. NPU utilization depends on whether your ops map cleanly. One unsupported layer breaks everything. A single op in the middle of the graph forces a CPU round-trip that tanks the speedup.

ST publishes reference benchmarks for common topologies (MobileNet v2, ResNet, YOLO variants) in the Model Zoo. Use those as sanity checks against your own measurements.

Troubleshooting Common Conversion Errors

ErrorLikely CauseFix
Unsupported operator: XOp not in current backend support matrixSwap for an equivalent supported op, or refactor the layer
Tensor shape mismatchDynamic batch dimensionSet an explicit input shape before exporting from Keras
Quantization parameters missingRepresentative dataset omitted or too smallProvide 100–200 real samples in the converter
Model too large for targetWeights exceed internal flashPrune, quantize more aggressively, or place weights in external flash
NPU fallback to CPU on most layersA non-INT8 op breaks the graphForce full integer quantization (TFLITE_BUILTINS_INT8 only)
Validation cross-accuracy <95%PTQ damaged a sensitive layerTry QAT, or exclude that layer from quantization

If you hit something not on this list, the stm32ai CLI has a --verbosity 3 flag that dumps the full conversion log. Most errors trace back to either an op mismatch or a quantization gap.

Next Moves Once Inference Works

  • Browse the ST Model Zoo on GitHub for pretrained, pre-quantized reference models you can adapt (image classification, object detection, human activity recognition, audio event detection).
  • If you’re working with sensor data and don’t need a deep network, look at NanoEdge AI Studio, ST’s sibling tool for classical ML on MCUs. Often a tiny SVM or anomaly detector beats a CNN at a fraction of the cost.
  • For accuracy-critical applications, redo the pipeline with quantization-aware training instead of PTQ. Same conversion steps, better numerics.
  • Profile power, not just latency. The N6’s NPU is fast, but the energy-per-inference figure is what determines battery life, and it doesn’t always track latency linearly.

The toolchain is moving quickly. Pin your STM32Cube.AI version in your build system, note it in your release docs, and re-validate when you upgrade. The C ABI of generated code has changed across major releases before, and it’ll change again.


Hubble Network lets you push model updates and collect inference telemetry from STM32 devices in the field over Bluetooth, without provisioning gateways or cellular.