The THUMOS Challenge on Action Recognition for Videos "in the Wild"
arXiv:1604.06182 · doi:10.1016/j.cviu.2016.10.018
Abstract
Automatically recognizing and localizing wide ranges of human actions has crucial importance for video understanding. Towards this goal, the THUMOS challenge was introduced in 2013 to serve as a benchmark for action recognition. Until then, video action recognition, including THUMOS challenge, had focused primarily on the classification of pre-segmented (i.e., trimmed) videos, which is an artificial task. In THUMOS 2014, we elevated action recognition to a more practical level by introducing temporally untrimmed videos. These also include `background videos' which share similar scenes and backgrounds as action videos, but are devoid of the specific actions. The three editions of the challenge organized in 2013--2015 have made THUMOS a common benchmark for action classification and detection and the annual challenge is widely attended by teams from around the world. In this paper we describe the THUMOS benchmark in detail and give an overview of data collection and annotation procedures. We present the evaluation protocols used to quantify results in the two THUMOS tasks of action classification and temporal detection. We also present results of submissions to the THUMOS 2015 challenge and review the participating approaches. Additionally, we include a comprehensive empirical study evaluating the differences in action recognition between trimmed and untrimmed videos, and how well methods trained on trimmed videos generalize to untrimmed videos. We conclude by proposing several directions and improvements for future THUMOS challenges.
Preprint submitted to Computer Vision and Image Understanding
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs
Cited by in corpus (57)
- A Review on Deep Learning Techniques for Video Prediction
- End-to-end Temporal Action Detection with Transformer
- Dynamic Sampling Networks for Efficient Action Recognition in Videos
- Deep Learning for Video Classification and Captioning
- Continuous Human Action Recognition for Human-Machine Interaction: A Review
- Zero-Shot Action Recognition in Videos: A Survey
- ACM-Net: Action Context Modeling Network for Weakly-Supervised Temporal Action Localization
- Video Action Understanding
- Precise Temporal Action Localization by Evolving Temporal Proposals
- Slow Motion Matters: A Slow Motion Enhanced Network for Weakly Supervised Temporal Action Localization
- 3C-Net: Category Count and Center Loss for Weakly-Supervised Action Localization
- Video-based Human Action Recognition using Deep Learning: A Review
- Deep Motion Prior for Weakly-Supervised Temporal Action Localization
- Weakly-Supervised Action Localization by Generative Attention Modeling
- Actor-agnostic Multi-label Action Recognition with Multi-modal Query
- Unsupervised learning from videos using temporal coherency deep networks
- VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph Convolutions
- Moviescope: Large-scale Analysis of Movies using Multiple Modalities
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Long Short-Term Transformer for Online Action Detection
- SF-Net: Single-Frame Supervision for Temporal Action Localization
- Multi-Scale Video Frame-Synthesis Network with Transitive Consistency Loss
- Action Unit Memory Network for Weakly Supervised Temporal Action Localization
- Relational Action Forecasting
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Classifying action correctness in physical rehabilitation exercises
- XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses
- Diagnosing Error in Temporal Action Detectors
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- Cross-Class Relevance Learning for Temporal Concept Localization
- Scale Matters: Temporal Scale Aggregation Network for Precise Action Localization in Untrimmed Videos
- Multi-modal Representation Learning for Video Advertisement Content Structuring
- Weakly Supervised Temporal Action Localization with Segment-Level Labels
- NAS-TC: Neural Architecture Search on Temporal Convolutions for Complex Action Recognition
- A Graph-based Interactive Reasoning for Human-Object Interaction Detection
- CMSN: Continuous Multi-stage Network and Variable Margin Cosine Loss for Temporal Action Proposal Generation
- Attention-Oriented Action Recognition for Real-Time Human-Robot Interaction
- W-TALC: Weakly-supervised Temporal Activity Localization and Classification
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Localizing the Common Action Among a Few Videos
- AIM 2019 Challenge on Video Temporal Super-Resolution: Methods and Results
- CoLA: Weakly-Supervised Temporal Action Localization with Snippet Contrastive Learning
- Deep Point-wise Prediction for Action Temporal Proposal
- Feature-Supervised Action Modality Transfer
- Efficient Modelling Across Time of Human Actions and Interactions
- TraMNet - Transition Matrix Network for Efficient Action Tube Proposals
- Bridging the gap between Human Action Recognition and Online Action Detection
- Fingerspelling Detection in American Sign Language
- Few-Shot Transformation of Common Actions into Time and Space
- ADNet: Temporal Anomaly Detection in Surveillance Videos
- Layout-induced Video Representation for Recognizing Agent-in-Place Actions
- SegCodeNet: Color-Coded Segmentation Masks for Activity Detection from Wearable Cameras
- Online Spatiotemporal Action Detection and Prediction via Causal Representations
- Exploring Frame Segmentation Networks for Temporal Action Localization
- DDLSTM: Dual-Domain LSTM for Cross-Dataset Action Recognition