Fitting a Neural Network into 128 KiB: Quantization and Embedded AI on Deepgreen

by Théo Taburet

How do you run a neural network on a microcontroller with only 128 KiB of RAM?

In the Deepgreen project, the goal is to automatically detect the early signatures of a fault in a rotating machine, using a vibration sensor. The detector is a small neural network, and the computing platform is a microcontroller with 128 KiB of RAM. To make the two fit together, we quantized the network: its numbers, represented as 32-bit floating-point values, become 8-bit integers accompanied by a scale. For example, the value 0.37 becomes the integer 47 with a scale of 1/128. The network uses four times less memory, and the processor performs integer calculations, faster and with lower energy consumption than floating-point computation.

Quantization has a second advantage, less well known than the reduction in memory usage. When the scales are powers of two, the entire network can be computed using integer multiplications, 32-bit additions, and bit shifts. No floating-point rounding is involved, and the program produces exactly the same result, bit for bit, on a desktop computer and on the embedded board. We can therefore verify the embedded code identically before loading it onto the MCU, and compare it against its reference with no tolerance: any discrepancy causes the export to fail.

The challenge comes down to one question: how many distinct values are needed to represent a signal? With 8 bits, there are 256. The scale therefore has to cover the extreme values without losing the resolution of the more common ones. For the weights, the learned parameters of the network, one scale per channel is sufficient. For the activations, the values computed at each pass of the signal, everything depends on calibration: choosing the ranges from a representative data stream.

There are two main approaches. Post-training quantization (PTQ) simply calibrates the network as it was trained. Quantization-aware training (QAT) retrains the network with simulated rounding inside the training loop. The latter is more expensive, and it can never compensate for a poorly chosen input range.

In Deepgreen, the network is an autoencoder: it learns to reconstruct the signal from a healthy machine, and the reconstruction error is used as an anomaly score. A machine's alarm threshold is the worst score observed during its healthy period, so the most extreme signal windows determine everything. Two choices follow from this. Activations remain at 16 bits, preserving both the dynamic range of these extreme windows and the resolution of ordinary windows. The weights are reduced to 8 bits, hence the name of the variant, w8a16. The calibration ranges are also computed over the entire training stream, so that no extreme window is clipped — all without retraining: calibration alone is sufficient.

The result is ultimately judged by the alarms. The quantized network detects the same early fault signatures as the floating-point network and produces the same false alarms on the same measurements. The detector validated in floating point therefore remains valid in int8, without requiring a new qualification. On the embedded board, the program reproduces the simulation bit for bit and processes each signal window in a fraction of a second, with no measurable loss in performance.