DaCapo: Accelerating Continuous Learning in Autonomous Systems for Video Analytics
arXiv:2403.14353 · doi:10.1109/ISCA59077.2024.00093
Abstract
Deep neural network (DNN) video analytics is crucial for autonomous systems such as self-driving vehicles, unmanned aerial vehicles (UAVs), and security robots. However, real-world deployment faces challenges due to their limited computational resources and battery power. To tackle these challenges, continuous learning exploits a lightweight "student" model at deployment (inference), leverages a larger "teacher" model for labeling sampled data (labeling), and continuously retrains the student model to adapt to changing scenarios (retraining). This paper highlights the limitations in state-of-the-art continuous learning systems: (1) they focus on computations for retraining, while overlooking the compute needs for inference and labeling, (2) they rely on power-hungry GPUs, unsuitable for battery-operated autonomous systems, and (3) they are located on a remote centralized server, intended for multi-tenant scenarios, again unsuitable for autonomous systems due to privacy, network availability, and latency concerns. We propose a hardware-algorithm co-designed solution for continuous learning, DaCapo, that enables autonomous systems to perform concurrent executions of inference, labeling, and training in a performant and energy-efficient manner. DaCapo comprises (1) a spatially-partitionable and precision-flexible accelerator enabling parallel execution of kernels on sub-accelerators at their respective precisions, and (2) a spatiotemporal resource allocation algorithm that strategically navigates the resource-accuracy tradeoff space, facilitating optimal decisions for resource allocation to achieve maximal accuracy. Our evaluation shows that DaCapo achieves 6.5% and 5.5% higher accuracy than a state-of-the-art GPU-based continuous learning systems, Ekya and EOMU, respectively, while consuming 254x less power.
References in corpus (14)
- Overcoming catastrophic forgetting in neural networks
- BinaryConnect: Training Deep Neural Networks with binary weights during propagations
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- Mixed Precision Training
- Trained Ternary Quantization
- Fixed Point Quantization of Deep Convolutional Networks
- SCALE-Sim: Systolic CNN Accelerator Simulator
- Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks
- Incremental Learning in Deep Convolutional Neural Networks Using Partial Network Sharing
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Online Coreset Selection for Rehearsal-based Continual Learning
- Online Distillation with Continual Learning for Cyclic Domain Shifts
- Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems
- Towards Efficient Convolutional Neural Network for Embedded Hardware via Multi-Dimensional Pruning