Zero-Shot Action Recognition in Videos: A Survey
arXiv:1909.06423 · doi:10.1016/j.neucom.2021.01.036
Abstract
Zero-Shot Action Recognition has attracted attention in the last years and many approaches have been proposed for recognition of objects, events and actions in images and videos. There is a demand for methods that can classify instances from classes that are not present in the training of models, especially in the complex problem of automatic video understanding, since collecting, annotating and labeling videos are difficult and laborious tasks. We have identified that there are many methods available in the literature, however, it is difficult to categorize which techniques can be considered state of the art. Despite the existence of some surveys about zero-shot action recognition in still images and experimental protocol, there is no work focused on videos. Therefore, we present a survey of the methods that comprise techniques to perform visual feature extraction and semantic feature extraction as well to learn the mapping between these features considering specifically zero-shot action recognition in videos. We also provide a complete description of datasets, experiments and protocols, presenting open issues and directions for future work, essential for the development of the computer vision research field.
Preprint
References in corpus (14)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- A Robust Real-Time Automatic License Plate Recognition Based on the YOLO Detector
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- A Short Note about Kinetics-600
- Introduction to the Bag of Features Paradigm for Image Classification and Retrieval
- Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
- Review of Action Recognition and Detection Methods
- Action2Vec: A Crossmodal Embedding Approach to Action Learning
- TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
- All About Knowledge Graphs for Actions
- ARID: A New Dataset for Recognizing Action in the Dark