Video-based Human Action Recognition using Deep Learning: A Review
arXiv:2208.03775
Abstract
Human action recognition is an important application domain in computer vision. Its primary aim is to accurately describe human actions and their interactions from a previously unseen data sequence acquired by sensors. The ability to recognize, understand, and predict complex human actions enables the construction of many important applications such as intelligent surveillance systems, human-computer interfaces, health care, security, and military applications. In recent years, deep learning has been given particular attention by the computer vision community. This paper presents an overview of the current state-of-the-art in action recognition using video analysis with deep learning techniques. We present the most important deep learning models for recognizing human actions, and analyze them to provide the current progress of deep learning algorithms applied to solve human action recognition problems in realistic videos highlighting their advantages and disadvantages. Based on the quantitative analysis using recognition accuracies reported in the literature, our study identifies state-of-the-art deep architectures in action recognition and then provides current trends and open problems for future works in this field.
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Generating Videos with Scene Dynamics
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Spatiotemporal Residual Networks for Video Action Recognition
- Towards Good Practices for Very Deep Two-Stream ConvNets
- CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016
- R-CNNs for Pose Estimation and Action Detection
- Advances in Human Action Recognition: A Survey
- Hierarchical Attention Network for Action Recognition in Videos
- Review of Action Recognition and Detection Methods
- Learning to Recognize 3D Human Action from A New Skeleton-based Representation Using Deep Convolutional Neural Networks
- Deep Convolutional Neural Networks for Action Recognition Using Depth Map Sequences
- Action Recognition with Joint Attention on Multi-Level Deep Features
- Action Recognition Based on Joint Trajectory Maps with Convolutional Neural Networks