More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation
arXiv:1912.00869
Abstract
Current state-of-the-art models for video action recognition are mostly based on expensive 3D ConvNets. This results in a need for large GPU clusters to train and evaluate such architectures. To address this problem, we present a lightweight and memory-friendly architecture for action recognition that performs on par with or better than current architectures by using only a fraction of resources. The proposed architecture is based on a combination of a deep subnet operating on low-resolution frames with a compact subnet operating on high-resolution frames, allowing for high efficiency and accuracy at the same time. We demonstrate that our approach achieves a reduction by times in FLOPs and times in memory usage compared to the baseline. This enables training deeper models with more input frames under the same computational budget. To further obviate the need for large-scale 3D convolutions, a temporal aggregation module is proposed to model temporal dependencies in a video at very small additional computational costs. Our models achieve strong performance on several action recognition benchmarks including Kinetics, Something-Something and Moments-in-time. The code and models are available at https://github.com/IBM/bLVNet-TAM.
Accepted at NeurIPS 2019, codes and models are available at https://github.com/IBM/bLVNet-TAM
Cited by in corpus (29)
- Is Space-Time Attention All You Need for Video Understanding?
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Attention Bottlenecks for Multimodal Fusion
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- A Comprehensive Study of Deep Video Action Recognition
- Space-time Mixing Attention for Video Transformer
- An Image is Worth 16x16 Words, What is a Video Worth?
- TDN: Temporal Difference Networks for Efficient Action Recognition
- TAM: Temporal Adaptive Module for Video Recognition
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal Reasoning
- Knowing What, Where and When to Look: Efficient Video Action Modeling with Attention
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- GTA: Global Temporal Attention for Video Action Understanding
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- Video Is Graph: Structured Graph Module for Video Action Recognition
- Density-Guided Label Smoothing for Temporal Localization of Driving Actions
- Busy-Quiet Video Disentangling for Video Classification
- VA-RED: Video Adaptive Redundancy Reduction
- VidTr: Video Transformer Without Convolutions
- Representing Videos as Discriminative Sub-graphs for Action Recognition
- Dynamic Network Quantization for Efficient Video Inference
- Efficient Video Transformers with Spatial-Temporal Token Selection
- PolyViT: Co-training Vision Transformers on Images, Videos and Audio
- Recent Progress in Appearance-based Action Recognition
- SSAN: Separable Self-Attention Network for Video Representation Learning
- Efficient Modelling Across Time of Human Actions and Interactions
- ST-ABN: Visual Explanation Taking into Account Spatio-temporal Information for Video Recognition
- Boosting Video Representation Learning with Multi-Faceted Integration