Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
arXiv:1507.05738
Abstract
Every moment counts in action recognition. A comprehensive understanding of human activity in video requires labeling every frame according to the actions occurring, placing multiple labels densely over a video sequence. To study this problem we extend the existing THUMOS dataset and introduce MultiTHUMOS, a new dataset of dense labels over unconstrained internet videos. Modeling multiple, dense labels benefits from temporal relations within and across classes. We define a novel variant of long short-term memory (LSTM) deep networks for modeling these temporal relations via multiple input and output connections. We show that this model improves action labeling accuracy and further enables deeper understanding tasks ranging from structured retrieval to action prediction.
To appear in IJCV
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Beyond Short Snippets: Deep Networks for Video Classification
- Describing Videos by Exploiting Temporal Structure
- Finding Action Tubes
- Temporal Pyramid Pooling Based Convolutional Neural Networks for Action Recognition
- Initialization Strategies of Spatio-Temporal Convolutional Neural Networks
Cited by in corpus (25)
- Learning to Diagnose with LSTM Recurrent Neural Networks
- Action Recognition using Visual Attention
- SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos
- A Survey on Content-Aware Video Analysis for Sports
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Generating Visual Explanations
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- TricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation
- Crowdsourcing in Computer Vision
- Predictive-Corrective Networks for Action Detection
- End-to-end Learning of Action Detection from Frame Glimpses in Videos
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- RED: Reinforced Encoder-Decoder Networks for Action Anticipation
- Action Recognition with Joint Attention on Multi-Level Deep Features
- Asynchronous Temporal Fields for Action Recognition
- AMTnet: Action-Micro-Tube Regression by End-to-end Trainable Deep Architecture
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- Detecting the Moment of Completion: Temporal Models for Localising Action Completion
- Action Classification and Highlighting in Videos
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Online Action Detection
- Online Detection of Action Start in Untrimmed, Streaming Videos
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- Exploring Frame Segmentation Networks for Temporal Action Localization