Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
arXiv:1504.00983 · doi:10.1145/2733373.2806226
Abstract
We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos.
Camera ready version for ACM Multimedia 2015
References in corpus (5)
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Unsupervised Learning of Video Representations using LSTMs
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
Cited by in corpus (13)
- A Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation
- Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
- Omni-sourced Webly-supervised Learning for Video Recognition
- Weak Supervision and Referring Attention for Temporal-Textual Association Learning
- ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization
- Action Unit Memory Network for Weakly Supervised Temporal Action Localization
- VRFP: On-the-fly Video Retrieval using Web Images and Fast Fisher Vector Products
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- Attention Transfer from Web Images for Video Recognition
- Foreground-Action Consistency Network for Weakly Supervised Temporal Action Localization
- Efficient Action Detection in Untrimmed Videos via Multi-Task Learning
- Temporal Bilinear Encoding Network of Audio-Visual Features at Low Sampling Rates
- CLTA: Contents and Length-based Temporal Attention for Few-shot Action Recognition