MicroFlow is a robust and efficient TinyML inference engine designed for deploying machine learning models on embedded systems. It was developed by Matteo Carnelos as part of his master's thesis project at the University of Padova in collaboration with Grepit AB.
MicroFlow uses a compiler-based approach, resulting in the following engine structure:
graph LR
subgraph host[Host]
model(Neural Network Model) --> compiler(MicroFlow Compiler)
end
subgraph target[Target]
code(Generated Source Code) --- weights[(Weights)]
code --- runtime(MicroFlow Runtime)
end
compiler --> code
compiler --> weights
MicroFlow consists of two primary components: the compiler, represented by the microflow-macros crate, and the runtime, represented by the microflow crate.
The compiler, which runs prior to the Rust compiler, is responsible for parsing and pre-processing the model.
It generates the necessary source code to enable inference on the model.
On the other hand, the runtime is a [no_std] component designed to run on the target MCU.
It encompasses the implementation of operators, activation functions, and quantization procedures.
MicroFlow utilizes Rust Procedural Macros as its user interface.
By applying the model macro to a struct and providing the model's path, the MicroFlow compiler generates a predict() method.
This method can be called to perform inference on the given model.
Currently, MicroFlow only supports models in the TensorFlow Lite format (.tflite).
Here is a minimal example showcasing the usage of MicroFlow:
use microflow::model;
#[model("path/to/model.tflite")]
struct MyModel;
fn main() {
let prediction = MyModel::predict(input_data);
}The examples provided with MicroFlow can be found in the examples folder.
To run an example on a target board, cd into the board directory for the example (e.g. examples/arduino-uno) and run the command:
cargo run --example <example-name>Otherwise, to run the example locally, just run the above command in the root directory.
Note
For board examples, you might need to install additional tools and configure the runner to make the example work for your setup.
cargo bench runs the criterion suites; cargo run --release --example latency
prints min/avg/max model latency over 20k runs:
cargo bench --bench conv1d # Conv1D vs the reshape trick + node model latency
cargo run --release --example latencyThe conv1d suite compares the dedicated 1-D kernel against the same work
expressed through the generic Conv2D path (the "reshape trick") on the
two conv layers of a real node model, plus the end-to-end predict()
latency of the bundled node models. Set MICROFLOW_CONV2D_ONLY=1 (with its
own CARGO_TARGET_DIR — the env var is read at macro-expansion time and
cargo does not fingerprint it) to rebuild a whole model onto the trick path
and compare the model_* rows across runs. See ../NOTES.md, week 6.
Currently, MicroFlow supports the following operators and activation functions:
| Operator | Quantized | Tensor Type |
|---|---|---|
FullyConnected |
✓ | Tensor2D |
Conv2D |
✓ | Tensor4D |
Conv1D |
✓ | Tensor4D |
DepthwiseConv2D |
✓ | Tensor4D |
AveragePool2D |
✓ | Tensor4D |
Transpose |
✓ | Tensor2D, Tensor4D |
Reshape |
✓ | Tensor2D, Tensor4D |
ExpandDims |
folded | shape only |
Shape/StridedSlice/Pack |
folded | the Flatten chain |
Conv1D covers the way Keras Conv1D layers serialize to TFLite: a
CONV_2D over a (1, 1, T, C) tensor (see docs/conv1d-spec.md). The
"folded" operators produce no code: the compiler folds them into virtual
reshapes at compile time, so rank-3 Keras Conv1D graphs build through
#[model] with a 2-D user-facing buffer.
| Activation Function | Quantized |
|---|---|
ReLU |
✓ |
ReLU6 |
✓ |
Softmax |
✓ |
These operators and activation functions cover common building blocks for neural networks and enable efficient inference with reduced memory and computational requirements. However, MicroFlow's development roadmap includes plans for implementing additional operators and activation functions to expand the range of supported models.
The examples folder contains the code used to test MicroFlow on different MCUs, including:
- ESP32 (32-bit Xtensa)
- ATSAMV71 (32-bit Cortex-M7F)
- nRF52840 (32-bit Cortex-M4F)
- LM3S6965 (32-bit Cortex-M3)
- ATmega328 (8-bit AVR)
The models ued to test the inference engines can be found in the models directory.
These models include:
- A sine predictor
- A speech command recognizer (TinyConv)
- A person detector (MobileNet v1)
Contributors are welcome. For major changes, please open an issue first to discuss what you would like to change. Please make sure to update tests as appropriate.
The MicroFlow paper has been published in Elsevier's Internet of Things journal and can be cited as follows:
@article{CARNELOS2025101498,
title = {MicroFlow: An Efficient Rust-Based Inference Engine for TinyML},
journal = {Internet of Things},
volume = {30},
pages = {101498},
year = {2025},
issn = {2542-6605},
doi = {https://doi.org/10.1016/j.iot.2025.101498},
url = {https://www.sciencedirect.com/science/article/pii/S2542660525000113},
author = {Matteo Carnelos and Francesco Pasti and Nicola Bellotto},
keywords = {TinyML, Rust, Neural networks, Embedded systems, IoT}
}Licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0)
- MIT license (LICENSE-MIT or http://opensource.org/licenses/MIT)
at your option.
Copyright © 2025, Matteo Carnelos
