Analyzing Human-Human Interactions: A Survey
arXiv:1808.00022 · doi:10.1016/j.cviu.2019.102799
Abstract
Many videos depict people, and it is their interactions that inform us of their activities, relation to one another and the cultural and social setting. With advances in human action recognition, researchers have begun to address the automated recognition of these human-human interactions from video. The main challenges stem from dealing with the considerable variation in recording setting, the appearance of the people depicted and the coordinated performance of their interaction. This survey provides a summary of these challenges and datasets to address these, followed by an in-depth discussion of relevant vision-based recognition and detection methods. We focus on recent, promising work based on deep learning and convolutional neural networks (CNNs). Finally, we outline directions to overcome the limitations of the current state-of-the-art to analyze and, eventually, understand social human actions.
References in corpus (24)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Neural Architecture Search with Reinforcement Learning
- How transferable are features in deep neural networks?
- The Kinetics Human Action Video Dataset
- Theoretical Models of Learning to Learn
- Spatiotemporal Residual Networks for Video Action Recognition
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- Distilling a Neural Network Into a Soft Decision Tree
- Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Advances in Human Action Recognition: A Survey
- An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action Recognition
- Exploiting deep residual networks for human action recognition from skeletal data
- Learning Feature Pyramids for Human Pose Estimation
- Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning
- Describing Common Human Visual Actions in Images
- Multi-Fiber Networks for Video Recognition
- Two-stream Flow-guided Convolutional Attention Networks for Action Recognition
- Predicting Human Interaction via Relative Attention Model
- Discovering Human Interactions in Videos with Limited Data Labeling
Cited by in corpus (6)
- Two-person Graph Convolutional Network for Skeleton-based Human Interaction Recognition
- Learn to cycle: Time-consistent feature discovery for action recognition
- Co-Located Human-Human Interaction Analysis using Nonverbal Cues: A Survey
- Spatio-Temporal FAST 3D Convolutions for Human Action Recognition
- Learning Class Regularized Features for Action Recognition
- Efficient Modelling Across Time of Human Actions and Interactions