Attend and Interact: Higher-Order Object Interactions for Video Understanding
arXiv:1711.06330
Abstract
Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation or pairwise object relationships. Furthermore, learning interactions across multiple objects in hundreds of frames for video is computationally infeasible and performance may suffer since a large combinatorial space has to be modeled. In this paper, we propose to efficiently learn higher-order interactions between arbitrary subgroups of objects for fine-grained video understanding. We demonstrate that modeling object interactions significantly improves accuracy for both action recognition and video captioning, while saving more than 3-times the computation over traditional pairwise relationships. The proposed method is validated on two large-scale datasets: Kinetics and ActivityNet Captions. Our SINet and SINet-Caption achieve state-of-the-art performances on both datasets even though the videos are sampled at a maximum of 1 FPS. To the best of our knowledge, this is the first work modeling object interactions on open domain large-scale video datasets, and we additionally model higher-order object interactions which improves the performance with low computational costs.
CVPR 2018
References in corpus (26)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- YouTube-8M: A Large-Scale Video Classification Benchmark
- A simple neural network module for relational reasoning
- Deformable Convolutional Networks
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- Learnable pooling with Context Gating for video classification
- Scene Graph Generation by Iterative Message Passing
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- Detecting Visual Relationships with Deep Relational Networks
- CIDEr: Consensus-based Image Description Evaluation
- Visual Translation Embedding Network for Visual Relation Detection
- Detecting and Recognizing Human-Object Interactions
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning
- ActivityNet Challenge 2017 Summary
- Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Weakly Supervised Dense Video Captioning
- Learning to Detect Human-Object Interactions
- Modeling Relationships in Referential Expressions with Compositional Modular Networks
- Semantic Compositional Networks for Visual Captioning
- Video Captioning with Transferred Semantic Attributes
- Top-down Visual Saliency Guided by Captions