Skip to content

Latest commit

 

History

History
215 lines (148 loc) · 5.87 KB

File metadata and controls

215 lines (148 loc) · 5.87 KB

Deploying a Model on a Microcontroller — Full Practical Guide

Language: English · فارسی

This document explains the Deploy stage from a .tflite file to a blinking LED.


Big picture

models/xxx_int8.tflite
        │
        ▼  python/04_tflite_to_c_array.py
firmware/.../model_data.h   (byte array in Flash)
        │
        ▼  Arduino IDE / PlatformIO / ESP-IDF
Microcontroller → TFLM Interpreter → Invoke() → decision

Prerequisites on PC (free)

  1. Arduino IDE 2.x or PlatformIO
  2. Suitable USB cable for the board
  3. Serial driver (for some CH340/CP2102 boards)

Until you have a board, complete steps 1–4 fully on the PC and keep the firmware ready.


Path 1 — Full control with TFLM (best for learning)

Step 1: Build an INT8 model

python python\02_train_mnist_tiny_cnn.py
python python\03_quantize_to_tflite.py --model models\mnist_tiny_cnn.keras

Step 2: Convert to a C header

python python\04_tflite_to_c_array.py `
  --tflite models\mnist_tiny_cnn_int8.tflite `
  --out firmware\hello_sine_arduino\model_data.h `
  --var g_model

Step 3: Firmware skeleton

Conceptual code (simplified):

#include "tensorflow/lite/micro/micro_interpreter.h"
#include "tensorflow/lite/micro/micro_mutable_op_resolver.h"
#include "model_data.h"

constexpr int kArenaSize = 80 * 1024;
uint8_t tensor_arena[kArenaSize];

void setup() {
  const tflite::Model* model = tflite::GetModel(g_model);
  static tflite::MicroMutableOpResolver<6> resolver;
  // resolver.AddConv2D(); resolver.AddFullyConnected(); ...
  static tflite::MicroInterpreter interpreter(
      model, resolver, tensor_arena, kArenaSize);
  interpreter.AllocateTensors();
}

void loop() {
  // 1) sensor → buffer
  // 2) same preprocessing as training
  // 3) copy into interpreter.input(0)
  // 4) interpreter.Invoke()
  // 5) read interpreter.output(0) → LED / serial
}

Ready templates: firmware/hello_sine_arduino/ and firmware/esp32_tflm_template/

Step 4: Arena size

  • Too small → Allocate/Invoke fails
  • Too large → fills all RAM

Practical method: start at 60KB, raise until stable, then trim with arena_used_bytes().

Step 5: Only the operators you need

For production use MicroMutableOpResolver.
Inspect model ops with a TFLite analysis tool or Netron:

https://netron.app


Path 2 — Edge Impulse (fastest to board)

  1. Go to https://studio.edgeimpulse.com and create a free account
  2. Create project → data type (Audio / Accel / Image)
  3. Upload data or capture from the board with the CLI
  4. Impulse design: Processing block + Learning block
  5. Generate features → Train
  6. Model testing
  7. Deployment → Arduino library → Build
  8. In Arduino IDE: Sketch → Include Library → Add .ZIP Library
  9. Upload the library’s ready example

Official docs:

https://docs.edgeimpulse.com/hardware/deployments/run-arduino-2-0

Pros: Preprocessing + model + SDK in one package.
Cons: You see less “under the hood” — for deep learning, also walk Path 1.


Path 3 — Board by board

ESP32 / ESP32-S3 (best first purchase)

  1. Arduino IDE → Preferences → Additional Boards URL:
https://raw.githubusercontent.com/espressif/arduino-esp32/gh-pages/package_esp32_index.json
  1. Boards Manager → install esp32
  2. Board: ESP32S3 Dev Module (or your board model)
  3. If you have vision: PSRAM = Enabled
  4. TFLM library / official Espressif example or Edge Impulse

Production notes:

  • Wi‑Fi together with heavy inference can stress RAM/CPU
  • Use I2S for audio
  • ESP-NN improves speed

Arduino Nano 33 BLE Sense

  • Onboard sensors: IMU, microphone, …
  • Official Arduino TensorFlow Lite library (match version to the board)
  • Edge Impulse examples for Nano are ready

Raspberry Pi Pico (RP2040 / RP2350)

  • Cheap and popular
  • TFLM is supported; ecosystem is a bit more hands-on than Arduino Sense
  • Great for IMU and educational projects

STM32

  • STM32Cube.AI for model conversion
  • CMSIS-NN performs well on Cortex-M
  • More industrial path

Checklist before Upload

  • Model is Full INT8
  • Firmware input shape = training shape
  • Input scale (0–1, −1–1, int8 scale) matches
  • Arena is large enough
  • Serial opens correctly at 115200 or 125200
  • You have a golden test (a known sample that must yield class X)

On-device preprocessing pattern (image example)

Training: 96×96 image, grayscale, normalize 0..1 then quantize to int8.

On MCU:

  1. Capture camera frame
  2. Resize to 96×96
  3. RGB→Gray (if needed)
  4. Convert to int8 with the same model scale/zero_point
  5. memcpy into input->data.int8
  6. Invoke()
  7. argmax on output

If any of these differ, accuracy dies even if the model is excellent.


Debugging when “it doesn’t work”

Symptom Likely cause
AllocateTensors fails Arena too small / model too large
Output always one class Bad input / wrong scale
Board resets Stack overflow / RAM exhaustion
Good on PC, bad on board Mismatched preprocess / sensor noise
TFLM compile error Incompatible library version

Golden tactic: first send a precomputed feature vector from the PC over serial and only Invoke on the MCU. If that works, the problem is sensor/preprocess — not the model.


What to do without a board?

  1. Complete the pipeline through .tflite and model_data.h
  2. Measure inference on PC with the TFLite Interpreter
  3. Build a project in Edge Impulse and test in Studio
  4. Write firmware and comment where the sensor hooks in

When the board arrives, just plug in the cable and Upload.


Next step: 05-Roadmap-0-to-100.md