Temporal Relational Reasoning in Videos
arXiv:1711.08496
Abstract
Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species. In this paper, we introduce an effective and interpretable network module, the Temporal Relation Network (TRN), designed to learn and reason about temporal dependencies between video frames at multiple time scales. We evaluate TRN-equipped networks on activity recognition tasks using three recent video datasets - Something-Something, Jester, and Charades - which fundamentally depend on temporal relational reasoning. Our results demonstrate that the proposed TRN gives convolutional neural networks a remarkable capacity to discover temporal relations in videos. Through only sparsely sampled video frames, TRN-equipped networks can accurately predict human-object interactions in the Something-Something dataset and identify various human gestures on the Jester dataset with very competitive performance. TRN-equipped networks also outperform two-stream networks and 3D convolution networks in recognizing daily activities in the Charades dataset. Further analyses show that the models learn intuitive and interpretable visual common sense knowledge in videos.
camera-ready version for ECCV'18
References in corpus (7)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- A simple neural network module for relational reasoning
- Learning to Poke by Poking: Experiential Learning of Intuitive Physics
- Time-Contrastive Networks: Self-Supervised Learning from Video
- On the effectiveness of task granularity for transfer learning
Cited by in corpus (22)
- TSM: Temporal Shift Module for Efficient Video Understanding
- Audiovisual SlowFast Networks for Video Recognition
- SCAN: Self-and-Collaborative Attention Network for Video Person Re-identification
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- AVA-AVD: Audio-Visual Speaker Diarization in the Wild
- RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data
- Temporal Pyramid Network for Action Recognition
- Learning Discriminative Motion Features Through Detection
- VirtualHome: Simulating Household Activities via Programs
- Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model
- FASTER Recurrent Networks for Efficient Video Classification
- IF-TTN: Information Fused Temporal Transformation Network for Video Action Recognition
- DenseImage Network: Video Spatial-Temporal Evolution Encoding and Understanding
- BAR: Bayesian Activity Recognition using variational inference
- Motion2Vec: Semi-Supervised Representation Learning from Surgical Videos
- Motion Feature Network: Fixed Motion Filter for Action Recognition
- Universal-to-Specific Framework for Complex Action Recognition
- Video Time: Properties, Encoders and Evaluation
- A Universal Model for Cross Modality Mapping by Relational Reasoning
- Channel-Temporal Attention for First-Person Video Domain Adaptation
- Recurrent Residual Module for Fast Inference in Videos