MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers
arXiv:2010.11267
Abstract
Executing machine learning workloads locally on resource constrained microcontrollers (MCUs) promises to drastically expand the application space of IoT. However, so-called TinyML presents severe technical challenges, as deep neural network inference demands a large compute and memory budget. To address this challenge, neural architecture search (NAS) promises to help design accurate ML models that meet the tight MCU memory, latency and energy constraints. A key component of NAS algorithms is their latency/energy model, i.e., the mapping from a given neural network architecture to its inference latency/energy on an MCU. In this paper, we observe an intriguing property of NAS search spaces for MCU model design: on average, model latency varies linearly with model operation (op) count under a uniform prior over models in the search space. Exploiting this insight, we employ differentiable NAS (DNAS) to search for models with low memory usage and low op count, where op count is treated as a viable proxy to latency. Experimental results validate our methodology, yielding our MicroNet models, which we deploy on MCUs using Tensorflow Lite Micro, a standard open-source NN inference runtime widely used in the TinyML community. MicroNets demonstrate state-of-the-art results for all three TinyMLperf industry-standard benchmark tasks: visual wake words, audio keyword spotting, and anomaly detection. Models and training scripts can be found at github.com/ARM-software/ML-zoo.
10 pages, 8 figures, 3 tables
References in corpus (16)
- Distilling the Knowledge in a Neural Network
- MCUNet: Tiny Deep Learning on IoT Devices
- Benchmarking TinyML Systems: Challenges and Direction
- TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems
- TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids
- Visual Wake Words Dataset
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- Neural Architecture Search For Keyword Spotting
- FixyNN: Efficient Hardware for Mobile Computer Vision via Transfer Learning
- FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- Not All Ops Are Created Equal!
- Deep Dense and Convolutional Autoencoders for Unsupervised Anomaly Detection in Machine Condition Sounds
- Ternary Hybrid Neural-Tree Networks for Highly Constrained IoT Applications
- Small-footprint Keyword Spotting with Graph Convolutional Network
- Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices
Cited by in corpus (21)
- Machine Learning for Microcontroller-Class Hardware: A Review
- Tiny Machine Learning: Progress and Futures
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Intelligence at the Extreme Edge: A Survey on Reformable TinyML
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Real-time Neural Network Inference on Extremely Weak Devices: Agile Offloading with Explainable AI
- An Ultra-low Power TinyML System for Real-time Visual Processing at Edge
- Few-Shot Keyword Spotting in Any Language
- Machine Learning with Confidential Computing: A Systematization of Knowledge
- Comprehensive Mapping of Continuous/Switching Circuits in CCM and DCM to Machine Learning Domain using Homogeneous Graph Neural Networks
- Software Engineering Approaches for TinyML based IoT Embedded Vision: A Systematic Literature Review
- Colab NAS: Obtaining lightweight task-specific convolutional neural networks following Occam's razor
- Power-Performance Characterization of TinyML Systems
- Hardware Aware Training for Efficient Keyword Spotting on General Purpose and Specialized Hardware
- AnalogNets: ML-HW Co-Design of Noise-robust TinyML Models and Always-On Analog Compute-in-Memory Accelerator
- Running hardware-aware neural architecture search on embedded devices under 512MB of RAM
- TinyOL: TinyML with Online-Learning on Microcontrollers
- MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
- Active multi-fidelity Bayesian online changepoint detection
- DietCNN: Multiplication-free Inference for Quantized CNNs
- Doping: A technique for efficient compression of LSTM models using sparse structured additive matrices