BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online Policies
arXiv:2108.09376 · doi:10.1109/ICCV48922.2021.00511
Abstract
In this paper we propose BlockCopy, a scheme that accelerates pretrained frame-based CNNs to process video more efficiently, compared to standard frame-by-frame processing. To this end, a lightweight policy network determines important regions in an image, and operations are applied on selected regions only, using custom block-sparse convolutions. Features of non-selected regions are simply copied from the preceding frame, reducing the number of computations and latency. The execution policy is trained using reinforcement learning in an online fashion without requiring ground truth annotations. Our universal framework is demonstrated on dense prediction tasks such as pedestrian detection, instance segmentation and semantic segmentation, using both state of the art (Center and Scale Predictor, MGAN, SwiftNet) and standard baseline networks (Mask-RCNN, DeepLabV3+). BlockCopy achieves significant FLOPS savings and inference speedup with minimal impact on accuracy.
Accepted for International Conference on Computer Vision (ICCV 2021)
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Distilling the Knowledge in a Neural Network
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference