Objects2action: Classifying and localizing actions without any video example
arXiv:1510.06939
Abstract
The goal of this paper is to recognize actions in video without the need for examples. Different from traditional zero-shot approaches we do not demand the design and specification of attribute classifiers and class-to-attribute mappings to allow for transfer from seen classes to unseen classes. Our key contribution is objects2action, a semantic word embedding that is spanned by a skip-gram model of thousands of object categories. Action labels are assigned to an object encoding of unseen video based on a convex combination of action and object affinities. Our semantic embedding has three main characteristics to accommodate for the specifics of actions. First, we propose a mechanism to exploit multiple-word descriptions of actions and objects. Second, we incorporate the automated selection of the most responsive objects per action. And finally, we demonstrate how to extend our zero-shot approach to the spatio-temporal localization of actions in video. Experiments on four action datasets demonstrate the potential of our approach.
References in corpus (4)
Cited by in corpus (15)
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Zero-Shot Action Recognition in Videos: A Survey
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- Towards Universal Representation for Unseen Action Recognition
- Probabilistic Semantic Retrieval for Surveillance Videos with Activity Graphs
- Localizing Actions from Video Labels and Pseudo-Annotations
- Attention Transfer from Web Images for Video Recognition
- Efficient Action Detection in Untrimmed Videos via Multi-Task Learning
- Video Stream Retrieval of Unseen Queries using Semantic Memory
- Few-Shot Adaptation for Multimedia Semantic Indexing
- Knowledge Guided Learning: Towards Open Domain Egocentric Action Recognition with Zero Supervision
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Automated Image Captioning for Rapid Prototyping and Resource Constrained Environments
- Spot On: Action Localization from Pointly-Supervised Proposals
- Searching Scenes by Abstracting Things