Learnable pooling with Context Gating for video classification
arXiv:1706.06905
Abstract
Current methods for video analysis often extract frame-level features using pre-trained convolutional neural networks (CNNs). Such features are then aggregated over time e.g., by simple temporal averaging or more sophisticated recurrent neural networks such as long short-term memory (LSTM) or gated recurrent units (GRU). In this work we revise existing video representations and study alternative methods for temporal aggregation. We first explore clustering-based aggregation layers and propose a two-stream architecture aggregating audio and visual features. We then introduce a learnable non-linear unit, named Context Gating, aiming to model interdependencies among network activations. Our experimental results show the advantage of both improvements for the task of video classification. In particular, we evaluate our method on the large-scale multi-modal Youtube-8M v2 dataset and outperform all other methods in the Youtube 8M Large-Scale Video Understanding challenge.
Presented at Youtube 8M CVPR17 Workshop. Kaggle Winning model. Under review for TPAMI
References in corpus (8)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Language Modeling with Gated Convolutional Networks
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Temporal Modeling Approaches for Large-scale Youtube-8M Video Understanding
- The Monkeytyping Solution to the YouTube-8M Video Understanding Challenge
- Deep Learning Methods for Efficient Large Scale Video Labeling
- Aggregating Frame-level Features for Large-Scale Video Classification
Cited by in corpus (61)
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- Searching for Activation Functions
- Image Super-Resolution Using Very Deep Residual Channel Attention Networks
- SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos
- Use What You Have: Video Retrieval Using Representations From Collaborative Experts
- Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
- Efficient Facial Representations for Age, Gender and Identity Recognition in Organizing Photo Albums using Multi-output CNN
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- FLAVR: Flow-Agnostic Video Representations for Fast Frame Interpolation
- Long-Term Feature Banks for Detailed Video Understanding
- Preferences Prediction using a Gallery of Mobile Device based on Scene Recognition and Object Detection
- CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data
- Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019
- Investigation of Multimodal Features, Classifiers and Fusion Methods for Emotion Recognition
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- EnsembleNet: End-to-End Optimization of Multi-headed Models
- Heavily Augmented Sound Event Detection utilizing Weak Predictions
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Long Short-Term Transformer for Online Action Detection
- TransLoc3D : Point Cloud based Large-scale Place Recognition using Adaptive Receptive Fields
- Temporal Gaussian Mixture Layer for Videos
- Non-local NetVLAD Encoding for Video Classification
- FASTER Recurrent Networks for Efficient Video Classification
- EEV: A Large-Scale Dataset for Studying Evoked Expressions from Video
- End-to-End Video Classification with Knowledge Graphs
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Deep Multimodal Feature Encoding for Video Ordering
- SPIN: A High Speed, High Resolution Vision Dataset for Tracking and Action Recognition in Ping Pong
- Uncertainty Quantification for Deep Context-Aware Mobile Activity Recognition and Unknown Context Discovery
- Understanding and Training Deep Diagonal Circulant Neural Networks
- Temporal Query Networks for Fine-grained Video Understanding
- Learning Efficient Video Representation with Video Shuffle Networks
- Cross-Class Relevance Learning for Temporal Concept Localization
- Less is More: Sparse Sampling for Dense Reaction Predictions
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- Universal-to-Specific Framework for Complex Action Recognition
- DIANet: Dense-and-Implicit Attention Network
- Hallucinating Optical Flow Features for Video Classification
- Cycled Compositional Learning between Images and Text
- IntegralAction: Pose-driven Feature Integration for Robust Human Action Recognition in Videos
- BERT for Large-scale Video Segment Classification with Test-time Augmentation
- FOSNet: An End-to-End Trainable Deep Neural Network for Scene Recognition
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Coarse Temporal Attention Network (CTA-Net) for Driver's Activity Recognition
- PGT: A Progressive Method for Training Models on Long Videos
- Multi-attention Networks for Temporal Localization of Video-level Labels
- Deep Multimodal Learning: An Effective Method for Video Classification
- Higher-order Network for Action Recognition
- Learning to Localize Temporal Events in Large-scale Video Data
- RCT: Random Consistency Training for Semi-supervised Sound Event Detection
- End-to-end Language Identification using NetFV and NetVLAD
- Label Denoising with Large Ensembles of Heterogeneous Neural Networks
- Constrained-size Tensorflow Models for YouTube-8M Video Understanding Challenge
- Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
- Classifying Video based on Automatic Content Detection Overview
- Learning Joint Embedding for Cross-Modal Retrieval
- Smoothed Gaussian Mixture Models for Video Classification and Recommendation
- A Closed-Form Learned Pooling for Deep Classification Networks
- A Multimodal Framework for Video Ads Understanding
- Semi-Supervised NMF-CNN For Sound Event Detection