UntrimmedNets for Weakly Supervised Action Recognition and Detection
arXiv:1703.03329
Abstract
Current action recognition methods heavily rely on trimmed videos for model training. However, it is expensive and time-consuming to acquire a large-scale trimmed video dataset. This paper presents a new weakly supervised architecture, called UntrimmedNet, which is able to directly learn action recognition models from untrimmed videos without the requirement of temporal annotations of action instances. Our UntrimmedNet couples two important components, the classification module and the selection module, to learn the action models and reason about the temporal duration of action instances, respectively. These two components are implemented with feed-forward networks, and UntrimmedNet is therefore an end-to-end trainable architecture. We exploit the learned models for action recognition (WSR) and detection (WSD) on the untrimmed video datasets of THUMOS14 and ActivityNet. Although our UntrimmedNet only employs weak supervision, our method achieves performance superior or comparable to that of those strongly supervised approaches on these two datasets.
camera-ready version to appear in CVPR2017
References in corpus (7)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Recurrent Models of Visual Attention
- Weakly Supervised Action Labeling in Videos Under Ordering Constraints
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
Cited by in corpus (8)
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
- Multi-granularity Generator for Temporal Action Proposal
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification
- Context-Aware RCNN: A Baseline for Action Detection in Videos
- Decoupling Localization and Classification in Single Shot Temporal Action Detection
- Person Re-identification Attack on Wearable Sensing