Skip to content
 
 

Repository files navigation

MicroFlow

A robust and efficient TinyML inference engine


MicroFlow is a robust and efficient TinyML inference engine designed for deploying machine learning models on embedded systems. It was developed by Matteo Carnelos as part of his master's thesis project at the University of Padova in collaboration with Grepit AB.

MicroFlow uses a compiler-based approach, resulting in the following engine structure:

graph LR
  subgraph host[Host]
    model(Neural Network Model) --> compiler(MicroFlow Compiler)
  end
  subgraph target[Target]
    code(Generated Source Code) --- weights[(Weights)]
    code --- runtime(MicroFlow Runtime)
  end
  compiler --> code
  compiler --> weights
Loading

MicroFlow consists of two primary components: the compiler, represented by the microflow-macros crate, and the runtime, represented by the microflow crate. The compiler, which runs prior to the Rust compiler, is responsible for parsing and pre-processing the model. It generates the necessary source code to enable inference on the model. On the other hand, the runtime is a [no_std] component designed to run on the target MCU. It encompasses the implementation of operators, activation functions, and quantization procedures.

Usage

MicroFlow utilizes Rust Procedural Macros as its user interface. By applying the model macro to a struct and providing the model's path, the MicroFlow compiler generates a predict() method. This method can be called to perform inference on the given model. Currently, MicroFlow only supports models in the TensorFlow Lite format (.tflite).

Here is a minimal example showcasing the usage of MicroFlow:

use microflow::model;

#[model("path/to/model.tflite")]
struct MyModel;

fn main() {
    let prediction = MyModel::predict(input_data);
}

Documentation

Examples

The examples provided with MicroFlow can be found in the examples folder. To run an example on a target board, cd into the board directory for the example (e.g. examples/arduino-uno) and run the command:

cargo run --example <example-name>

Otherwise, to run the example locally, just run the above command in the root directory.

Note

For board examples, you might need to install additional tools and configure the runner to make the example work for your setup.

Benchmarks

cargo bench runs the criterion suites; cargo run --release --example latency prints min/avg/max model latency over 20k runs:

cargo bench --bench conv1d   # Conv1D vs the reshape trick + node model latency
cargo run --release --example latency

The conv1d suite compares the dedicated 1-D kernel against the same work expressed through the generic Conv2D path (the "reshape trick") on the two conv layers of a real node model, plus the end-to-end predict() latency of the bundled node models. Set MICROFLOW_CONV2D_ONLY=1 (with its own CARGO_TARGET_DIR — the env var is read at macro-expansion time and cargo does not fingerprint it) to rebuild a whole model onto the trick path and compare the model_* rows across runs. See ../NOTES.md, week 6.

Supported Operators

Currently, MicroFlow supports the following operators and activation functions:

Operator Quantized Tensor Type
FullyConnected ✓ Tensor2D
Conv2D ✓ Tensor4D
Conv1D ✓ Tensor4D
DepthwiseConv2D ✓ Tensor4D
AveragePool2D ✓ Tensor4D
Transpose ✓ Tensor2D, Tensor4D
Reshape ✓ Tensor2D, Tensor4D
ExpandDims folded shape only
Shape/StridedSlice/Pack folded the Flatten chain

Conv1D covers the way Keras Conv1D layers serialize to TFLite: a CONV_2D over a (1, 1, T, C) tensor (see docs/conv1d-spec.md). The "folded" operators produce no code: the compiler folds them into virtual reshapes at compile time, so rank-3 Keras Conv1D graphs build through #[model] with a 2-D user-facing buffer.

Activation Function Quantized
ReLU ✓
ReLU6 ✓
Softmax ✓

These operators and activation functions cover common building blocks for neural networks and enable efficient inference with reduced memory and computational requirements. However, MicroFlow's development roadmap includes plans for implementing additional operators and activation functions to expand the range of supported models.

Tested Models and MCUs

The examples folder contains the code used to test MicroFlow on different MCUs, including:

  • ESP32 (32-bit Xtensa)
  • ATSAMV71 (32-bit Cortex-M7F)
  • nRF52840 (32-bit Cortex-M4F)
  • LM3S6965 (32-bit Cortex-M3)
  • ATmega328 (8-bit AVR)

The models ued to test the inference engines can be found in the models directory. These models include:

  • A sine predictor
  • A speech command recognizer (TinyConv)
  • A person detector (MobileNet v1)

Contributing

Contributors are welcome. For major changes, please open an issue first to discuss what you would like to change. Please make sure to update tests as appropriate.

Citation

The MicroFlow paper has been published in Elsevier's Internet of Things journal and can be cited as follows:

@article{CARNELOS2025101498,
  title = {MicroFlow: An Efficient Rust-Based Inference Engine for TinyML},
  journal = {Internet of Things},
  volume = {30},
  pages = {101498},
  year = {2025},
  issn = {2542-6605},
  doi = {https://doi.org/10.1016/j.iot.2025.101498},
  url = {https://www.sciencedirect.com/science/article/pii/S2542660525000113},
  author = {Matteo Carnelos and Francesco Pasti and Nicola Bellotto},
  keywords = {TinyML, Rust, Neural networks, Embedded systems, IoT}
}

License

Licensed under either of

at your option.

Copyright © 2025, Matteo Carnelos

About

A robust and efficient TinyML inference engine.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages