A Deep Structured Model with Radius-Margin Bound for 3D Human Activity Recognition
arXiv:1512.01642 · doi:10.1007/s11263-015-0876-z
Abstract
Understanding human activity is very challenging even with the recently developed 3D/depth sensors. To solve this problem, this work investigates a novel deep structured model, which adaptively decomposes an activity instance into temporal parts using the convolutional neural networks (CNNs). Our model advances the traditional deep learning approaches in two aspects. First, { we incorporate latent temporal structure into the deep model, accounting for large temporal variations of diverse human activities. In particular, we utilize the latent variables to decompose the input activity into a number of temporally segmented sub-activities, and accordingly feed them into the parts (i.e. sub-networks) of the deep architecture}. Second, we incorporate a radius-margin bound as a regularization term into our deep model, which effectively improves the generalization performance for classification. For model training, we propose a principled learning algorithm that iteratively (i) discovers the optimal latent variables (i.e. the ways of activity decomposition) for all training instances, (ii) { updates the classifiers} based on the generated features, and (iii) updates the parameters of multi-layer neural networks. In the experiments, our approach is validated on several complex scenarios for human activity recognition and demonstrates superior performances over other state-of-the-art approaches.
16 pages, 9 figures, to appear in International Journal of Computer Vision 2015
References in corpus (3)
Cited by in corpus (12)
- DISC: Deep Image Saliency Computing via Progressive Representation Learning
- Pose-conditioned Spatio-Temporal Attention for Human Action Recognition
- Content-Adaptive Sketch Portrait Generation by Decompositional Representation Learning
- Learning Deep Similarity Models with Focus Ranking for Fabric Image Retrieval
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning
- Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
- Boosting Image Super-Resolution Via Fusion of Complementary Information Captured by Multi-Modal Sensors
- Coarse Temporal Attention Network (CTA-Net) for Driver's Activity Recognition
- Neural Task Planning with And-Or Graph Representations
- Attention-Driven Body Pose Encoding for Human Activity Recognition
- Learning to Segment Object Candidates via Recursive Neural Networks