Unstructured Human Activity Detection from RGBD Images
arXiv:1107.0169 · doi:10.1109/ICRA.2012.6224591
Abstract
Being able to detect and recognize human activities is essential for several applications, including personal assistive robotics. In this paper, we perform detection and recognition of unstructured human activity in unstructured environments. We use a RGBD sensor (Microsoft Kinect) as the input sensor, and compute a set of features based on human pose and motion, as well as based on image and pointcloud information. Our algorithm is based on a hierarchical maximum entropy Markov model (MEMM), which considers a person's activity as composed of a set of sub-activities. We infer the two-layered graph structure using a dynamic programming approach. We test our algorithm on detecting and recognizing twelve different activities performed by four people in different environments, such as a kitchen, a living room, an office, etc., and achieve good performance even when the person was not seen before in the training set.
2012 IEEE International Conference on Robotics and Automation (A preliminary version of this work was presented at AAAI workshop on Pattern, Activity and Intent Recognition, 2011)
References in corpus (1)
Cited by in corpus (43)
- Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks
- An Unsupervised Approach for Automatic Activity Recognition based on Hidden Markov Model Regression
- PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding
- A Deep Structured Model with Radius-Margin Bound for 3D Human Activity Recognition
- Two-Stream RNN/CNN for Action Recognition in 3D Videos
- Deep Learning for Computer Vision based Activity Recognition and Fall Detection of the Elderly: a Systematic Review
- Action-Attending Graphic Neural Network
- Contextually Guided Semantic Labeling and Search for 3D Point Clouds
- A discussion on the validation tests employed to compare human action recognition methods using the MSR Action3D dataset
- ChaLearn Looking at People: IsoGD and ConGD Large-scale RGB-D Gesture Recognition
- Spatio-Temporal Graph Convolution for Skeleton Based Action Recognition
- Joint Attention in Driver-Pedestrian Interaction: from Theory to Practice
- Space-Time Representation of People Based on 3D Skeletal Data: A Review
- Exploiting the ConvLSTM: Human Action Recognition using Raw Depth Video-Based Recurrent Neural Networks
- Simultaneous Joint and Object Trajectory Templates for Human Activity Recognition from 3-D Data
- Human Activity Learning using Object Affordances from RGB-D Videos
- Mining Mid-level Features for Action Recognition Based on Effective Skeleton Representation
- Action Recognition in the Frequency Domain
- A Survey of Visual Analysis of Human Motion and Its Applications
- Learning and Refining of Privileged Information-based RNNs for Action Recognition from Depth Sequences
- Robust 3D Action Recognition through Sampling Local Appearances and Global Distributions
- Learning Human Activities and Object Affordances from RGB-D Videos
- What's the point? Frame-wise Pointing Gesture Recognition with Latent-Dynamic Conditional Random Fields
- Synthetic Defocus and Look-Ahead Autofocus for Casual Videography
- Watch-n-Patch: Unsupervised Learning of Actions and Relations
- A-MAL: Automatic Movement Assessment Learning from Properly Performed Movements in 3D Skeleton Videos
- Kinematic-Layout-aware Random Forests for Depth-based Action Recognition
- Action4D: Real-time Action Recognition in the Crowd and Clutter
- Dynamic Probabilistic Network Based Human Action Recognition
- 3D Scene Grammar for Parsing RGB-D Pointclouds
- From Pose to Activity: Surveying Datasets and Introducing CONVERSE
- Simultaneous Feature and Body-Part Learning for Real-Time Robot Awareness of Human Behaviors
- Deep-Temporal LSTM for Daily Living Action Recognition
- Real-time Human Action Recognition Using Locally Aggregated Kinematic-Guided Skeletonlet and Supervised Hashing-by-Analysis Model
- ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos
- Understanding Human Context in 3D Scenes by Learning Spatial Affordances with Virtual Skeleton Models
- Human activity recognition from skeleton poses
- Simultaneous Learning from Human Pose and Object Cues for Real-Time Activity Recognition
- Complexity-Aware Assignment of Latent Values in Discriminative Models for Accurate Gesture Recognition
- Efficient Modelling Across Time of Human Actions and Interactions
- Skeleton-based Activity Recognition with Local Order Preserving Match of Linear Patches
- Watch-Bot: Unsupervised Learning for Reminding Humans of Forgotten Actions
- Unsupervised Temporal Segmentation of Repetitive Human Actions Based on Kinematic Modeling and Frequency Analysis