How to Use STM32Cube.AI to Convert a TensorFlow Model for STM32

A 250 KB neural network can run 10x faster on the STM32N6 than on an STM32H7, while burning a fraction of the power. The catch? You have to get it through STM32Cube.AI first, and the toolchain has opinions about your model that TensorFlow never warned you about.
This guide walks you through the full pipeline: a pretrained Keras keyword-spotting model, quantized to INT8, converted to C, deployed to an STM32N6570-DK, and benchmarked on the Neural-ART NPU. Allow yourself an afternoon. By the end you’ll know which errors to expect, how to read the analyzer report, and what kind of inference latency to plan for in your product.
[TF/Keras .h5] -> [TFLite INT8] -> [STM32Cube.AI] -> [C code] -> [STM32N6]
train quantize analyze/gen build deployWhy STM32Cube.AI and Why the N6
STM32Cube.AI (shipped as the X-CUBE-AI expansion pack) takes a trained TensorFlow, Keras, ONNX, or TFLite model and emits optimized C code that links against ST’s neural network runtime. It handles memory layout, operator fusion, and on the STM32N6 specifically, it compiles INT8 graphs down to the Neural-ART accelerator, ST’s NPU running alongside the 800 MHz Cortex-M55.
The N6 matters because most STM32 parts run inference on the CPU. The N6 has dedicated silicon (rated by ST at 600 GOPS) so models that would take 40+ ms on an M7 finish in single-digit milliseconds. That’s the difference between a feasibility study and a shippable product.
This tutorial uses STM32Cube.AI 10.0 (verify the current version on ST’s developer zone before you start, the CLI flags shift release to release).
A Quick Primer on TF Models and Quantization
STM32Cube.AI accepts Keras .h5, SavedModel directories, TFLite .tflite, and ONNX. For the N6 you want .tflite with full INT8 quantization, because that’s what the Neural-ART can actually accelerate. A float32 model will still convert and run, but it’ll fall back to the CPU and you’ll have paid for an NPU you aren’t using.
Quantization maps your float32 weights and activations to 8-bit integers. You get a 4x size reduction and (on the N6) a massive speedup, in exchange for a small accuracy hit, typically 0.5 to 2% for well-behaved models.
Two flavors:
- Post-training quantization (PTQ): convert after the fact using a small representative dataset. Easy, usually good enough.
- Quantization-aware training (QAT): simulate quantization during training. More work, better accuracy preservation. Use this if PTQ costs you more than 1% accuracy on a metric you care about.
We’ll use PTQ here.
Prerequisites
Hardware:
- STM32N6570-DK discovery board
- USB-C cable, ST-LINK works over the same port
Software:
- STM32CubeIDE 1.16+
- X-CUBE-AI 10.0 (install via CubeMX: Help → Manage embedded software packages → STMicroelectronics → X-CUBE-AI)
- Python 3.9+, TensorFlow 2.15+
- The
stm32aiCLI. It ships inside the X-CUBE-AI install, usually under~/STM32Cube/Repository/Packs/STMicroelectronics/X-CUBE-AI/<version>/Utilities/<os>/.
Sample model: a 12-class keyword-spotting CNN trained on the Google Speech Commands dataset. Input is a 49x10 MFCC, output is a 12-way softmax. Float model is ~380 KB. You can swap in your own; the workflow doesn’t change.
Step 1: Quantize the TensorFlow Model
Load the Keras model and convert it to INT8 TFLite. The representative dataset is the part everyone forgets, and skipping it silently produces a float model dressed up in TFLite clothing.
import tensorflow as tf
import numpy as np
model = tf.keras.models.load_model("kws_cnn.h5")
def representative_dataset():
# 100-200 real samples from your validation set
for sample in val_samples[:200]:
yield [sample.astype(np.float32).reshape(1, 49, 10, 1)]
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
tflite_model = converter.convert()
with open("kws_int8.tflite", "wb") as f:
f.write(tflite_model)Two things to check:
- File size dropped roughly 4x (mine went from 380 KB to 98 KB).
- Accuracy on your validation set is within tolerance. Run the
.tflitethroughtf.lite.Interpreterand compare top-1 against the float model. A 0.7% drop is fine. A 15% drop means PTQ isn’t working and you need QAT or a closer look at which layers are causing the damage.
Common pitfall: if you omit representative_dataset, the converter falls back to dynamic-range quantization (weights INT8, activations float). The file will be small but the N6 NPU won’t touch it.
Step 2: Analyze the Model with STM32Cube.AI
The CubeMX GUI works, but the CLI is reproducible and scriptable, which is what you want once you’re iterating.
stm32ai analyze \
--model kws_int8.tflite \
--target stm32n6 \
--st-neural-art default@user_neuralart.json \
--workspace ./workspace \
--output ./outputThe report shows:
- MACCs: multiply-accumulate operations per inference. Useful for back-of-envelope timing.
- ROM/RAM footprint: weights go to flash, activations to RAM. The N6 has 4.2 MB of contiguous SRAM, generous by MCU standards.
- Per-layer breakdown: which ops mapped to the Neural-ART, which fell back to the M55. You want as few CPU fallbacks as possible.
- Complexity warnings: layers flagged as expensive or unsupported.
If you see fallbacks, the analyzer tells you why. Usually it’s an op the Neural-ART doesn’t support, or an INT8 layer that didn’t quantize cleanly. ST publishes an op support matrix per release; check yours against it.
Step 3: Validate Numerical Equivalence
Code generation rewrites operators, fuses layers, and repacks memory, so confirm the deployed model still produces the same outputs as the original .tflite.
Desktop validation:
stm32ai validate --model kws_int8.tflite --target stm32n6 --mode hostOn-target validation (board plugged in):
stm32ai validate --model kws_int8.tflite --target stm32n6 --mode targetThe tool feeds inputs through both the reference model and the deployed network. It reports cross-accuracy, RMSE, and max error per output. For an INT8 KWS model I’d expect cross-accuracy >99% and max error <2 LSB. If it’s worse, something is off in the conversion (often a fused activation that didn’t round-trip).
Step 4: Generate Code and Integrate
stm32ai generate \
--model kws_int8.tflite \
--target stm32n6 \
--st-neural-art default@user_neuralart.json \
--name network \
--output ./generatedYou’ll get network.c, network.h, network_data.c, plus the Neural-ART compiled blob. Drop these into your CubeIDE project under Middlewares/ST/AI/. CubeMX can do this for you if you enabled X-CUBE-AI on the project.
Minimal inference call in main.c:
#include "network.h"
#include "network_data.h"
static ai_handle network = AI_HANDLE_NULL;
AI_ALIGNED(32) static ai_u8 activations[AI_NETWORK_DATA_ACTIVATIONS_SIZE];
AI_ALIGNED(32) static ai_i8 in_data[AI_NETWORK_IN_1_SIZE];
AI_ALIGNED(32) static ai_i8 out_data[AI_NETWORK_OUT_1_SIZE];
static ai_buffer ai_input[AI_NETWORK_IN_NUM] = AI_NETWORK_IN;
static ai_buffer ai_output[AI_NETWORK_OUT_NUM] = AI_NETWORK_OUT;
void ai_init(void) {
ai_network_create_and_init(&network, activations, NULL);
ai_input[0].data = AI_HANDLE_PTR(in_data);
ai_output[0].data = AI_HANDLE_PTR(out_data);
}
void ai_run(void) {
// populate in_data with INT8 MFCC features
ai_network_run(network, ai_input, ai_output);
// out_data now holds 12 INT8 class scores
}Wire ai_run() to your audio frontend and dequantize the output if you want probabilities.
Real Benchmarks on the STM32N6
Numbers below are representative for the KWS model described above, measured on an STM32N6570-DK with code in internal flash and activations in internal SRAM. Yours will vary with model topology, clock settings, and memory placement. Always measure on your own hardware before committing to a spec sheet.
Inference Latency (ms, lower = better)
M55 @ 800 MHz |####################### 38.4
Neural-ART |### 4.2
+-----------------------------
0 10 20 30 40| Backend | Latency | RAM | Flash | Notes |
|---|---|---|---|---|
| M55 CPU only | ~38 ms | 84 KB | 102 KB | INT8, CMSIS-NN kernels |
| Neural-ART NPU | ~4 ms | 68 KB | 102 KB | INT8, NPU accelerated |
A few honest caveats. Benchmark steady state, not the first inference, which includes one-time setup. Memory placement matters a lot: putting weights in external Octo-SPI flash can double latency for weight-heavy layers. NPU utilization depends on whether your ops map cleanly. One unsupported layer breaks everything. A single op in the middle of the graph forces a CPU round-trip that tanks the speedup.
ST publishes reference benchmarks for common topologies (MobileNet v2, ResNet, YOLO variants) in the Model Zoo. Use those as sanity checks against your own measurements.
Troubleshooting Common Conversion Errors
| Error | Likely Cause | Fix |
|---|---|---|
Unsupported operator: X | Op not in current backend support matrix | Swap for an equivalent supported op, or refactor the layer |
Tensor shape mismatch | Dynamic batch dimension | Set an explicit input shape before exporting from Keras |
Quantization parameters missing | Representative dataset omitted or too small | Provide 100–200 real samples in the converter |
Model too large for target | Weights exceed internal flash | Prune, quantize more aggressively, or place weights in external flash |
| NPU fallback to CPU on most layers | A non-INT8 op breaks the graph | Force full integer quantization (TFLITE_BUILTINS_INT8 only) |
| Validation cross-accuracy <95% | PTQ damaged a sensitive layer | Try QAT, or exclude that layer from quantization |
If you hit something not on this list, the stm32ai CLI has a --verbosity 3 flag that dumps the full conversion log. Most errors trace back to either an op mismatch or a quantization gap.
Next Moves Once Inference Works
- Browse the ST Model Zoo on GitHub for pretrained, pre-quantized reference models you can adapt (image classification, object detection, human activity recognition, audio event detection).
- If you’re working with sensor data and don’t need a deep network, look at NanoEdge AI Studio, ST’s sibling tool for classical ML on MCUs. Often a tiny SVM or anomaly detector beats a CNN at a fraction of the cost.
- For accuracy-critical applications, redo the pipeline with quantization-aware training instead of PTQ. Same conversion steps, better numerics.
- Profile power, not just latency. The N6’s NPU is fast, but the energy-per-inference figure is what determines battery life, and it doesn’t always track latency linearly.
The toolchain is moving quickly. Pin your STM32Cube.AI version in your build system, note it in your release docs, and re-validate when you upgrade. The C ABI of generated code has changed across major releases before, and it’ll change again.
Hubble Network lets you push model updates and collect inference telemetry from STM32 devices in the field over Bluetooth, without provisioning gateways or cellular.