A TinyML Platform for On-Device Continual Learning with Quantized Latent Replays
arXiv:2110.10486 · doi:10.1109/JETCAS.2021.3121554
Abstract
In the last few years, research and development on Deep Learning models and techniques for ultra-low-power devices in a word, TinyML has mainly focused on a train-then-deploy assumption, with static models that cannot be adapted to newly collected data without cloud-based data collection and fine-tuning. Latent Replay-based Continual Learning (CL) techniques[1] enable online, serverless adaptation in principle, but so farthey have still been too computation and memory-hungry for ultra-low-power TinyML devices, which are typically based on microcontrollers. In this work, we introduce a HW/SW platform for end-to-end CL based on a 10-core FP32-enabled parallel ultra-low-power (PULP) processor. We rethink the baseline Latent Replay CL algorithm, leveraging quantization of the frozen stage of the model and Latent Replays (LRs) to reduce their memory cost with minimal impact on accuracy. In particular, 8-bit compression of the LR memory proves to be almost lossless (-0.26% with 3000LR) compared to the full-precision baseline implementation, but requires 4x less memory, while 7-bit can also be used with an additional minimal accuracy degradation (up to 5%). We also introduce optimized primitives for forward and backward propagation on the PULP processor. Our results show that by combining these techniques, continual learning can be achieved in practice using less than 64MB of memory an amount compatible with embedding in TinyML devices. On an advanced 22nm prototype of our platform, called VEGA, the proposed solution performs onaverage 65x faster than a low-power STM32 L4 microcontroller, being 37x more energy efficient enough for a lifetime of 535h when learning a new mini-batch of data once every minute.
14 pages
References in corpus (6)
- Three scenarios for continual learning
- On Tiny Episodic Memories in Continual Learning
- DORY: Automatic End-to-End Deployment of Real-World DNNs on Low-Cost IoT MCUs
- Robust High-dimensional Memory-augmented Neural Networks
- Technical Report: NEMO DNN Quantization for Deployment Model
- Batch-level Experience Replay with Review for Continual Learning
Cited by in corpus (9)
- Machine Learning for Microcontroller-Class Hardware: A Review
- Intelligence at the Extreme Edge: A Survey on Reformable TinyML
- DARKSIDE: A Heterogeneous RISC-V Compute Cluster for Extreme-Edge On-Chip DNN Inference and Training
- On-device Learning of EEGNet-based Network For Wearable Motor Imagery Brain-Computer Interface
- FeTrIL: Feature Translation for Exemplar-Free Class-Incremental Learning
- QCore: Data-Efficient, On-Device Continual Calibration for Quantized Models -- Extended Version
- TinySV: Speaker Verification in TinyML with On-device Learning
- An Ultra-Low Power Wearable BMI System with Continual Learning Capabilities
- TActiLE: Tiny Active LEarning for wearable devices