Tiny Machine Learning: Progress and Futures
arXiv:2403.19076 · doi:10.1109/MCAS.2023.3302182
Abstract
Tiny Machine Learning (TinyML) is a new frontier of machine learning. By squeezing deep learning models into billions of IoT devices and microcontrollers (MCUs), we expand the scope of AI applications and enable ubiquitous intelligence. However, TinyML is challenging due to hardware constraints: the tiny memory resource makes it difficult to hold deep learning models designed for cloud and mobile platforms. There is also limited compiler and inference engine support for bare-metal devices. Therefore, we need to co-design the algorithm and system stack to enable TinyML. In this review, we will first discuss the definition, challenges, and applications of TinyML. We then survey the recent progress in TinyML and deep learning on MCUs. Next, we will introduce MCUNet, showing how we can achieve ImageNet-scale AI applications on IoT devices with system-algorithm co-design. We will further extend the solution from inference to training and introduce tiny on-device training techniques. Finally, we present future directions in this area. Today's large model might be tomorrow's tiny model. The scope of TinyML should evolve and adapt over time.
arXiv admin note: text overlap with arXiv:2206.15472
References in corpus (42)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- TensorFlow: A system for large-scale machine learning
- Neural Architecture Search with Reinforcement Learning
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Federated Learning: Strategies for Improving Communication Efficiency
- PaLM: Scaling Language Modeling with Pathways
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- EfficientNetV2: Smaller Models and Faster Training
- Compressing Deep Convolutional Networks using Vector Quantization
- An Analysis of Deep Neural Network Models for Practical Applications
- Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks
- Trained Ternary Quantization
- PACT: Parameterized Clipping Activation for Quantized Neural Networks
- Training Deep Nets with Sublinear Memory Cost
- Large Batch Training of Convolutional Networks
- MCUNet: Tiny Deep Learning on IoT Devices
- Learning to Optimize Tensor Programs
- MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers
- DORY: Automatic End-to-End Deployment of Real-World DNNs on Low-Cost IoT MCUs
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
- Highway and Residual Networks learn Unrolled Iterative Estimation
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
- On-Device Training Under 256KB Memory
- A Tiny CNN Architecture for Medical Face Mask Detection for Resource-Constrained Endpoints
- MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
- K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning
- LFFD: A Light and Fast Face Detector for Edge Devices
- EXTD: Extremely Tiny Face Detector via Iterative Filter Reuse
- Few-Shot Keyword Spotting in Any Language
- TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learning
- Rethinking Generalization in American Sign Language Prediction for Edge Devices with Extremely Low Memory Footprint
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- RNNPool: Efficient Non-linear Pooling for RAM Constrained Inference
- Enabling Large Neural Networks on Tiny Microcontrollers with Swapping
- POET: Training Neural Networks on Tiny Devices with Integrated Rematerialization and Paging
Cited by in corpus (8)
- On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
- Edge Intelligence for Wildlife Conservation: Real-Time Hornbill Call Classification Using TinyML
- Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
- Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
- Beyond ImageNet: Understanding Cross-Dataset Robustness of Lightweight Vision Models
- EmbBERT: Attention Under 2 MB Memory
- TinyML NLP Scheme for Semantic Wireless Sentiment Classification with Privacy Preservation
- From Tiny Machine Learning to Tiny Deep Learning: A Survey