STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
arXiv:1910.12906 · doi:10.1609/aaai.v34i02.5490
Abstract
We present a novel classifier network called STEP, to classify perceived human emotion from gaits, based on a Spatial Temporal Graph Convolutional Network (ST-GCN) architecture. Given an RGB video of an individual walking, our formulation implicitly exploits the gait features to classify the emotional state of the human into one of four emotions: happy, sad, angry, or neutral. We use hundreds of annotated real-world gait videos and augment them with thousands of annotated synthetic gaits generated using a novel generative network called STEP-Gen, built on an ST-GCN based Conditional Variational Autoencoder (CVAE). We incorporate a novel push-pull regularization loss in the CVAE formulation of STEP-Gen to generate realistic gaits and improve the classification accuracy of STEP. We also release a novel dataset (E-Gait), which consists of human gaits annotated with perceived emotions along with thousands of synthetic gaits. In practice, STEP can learn the affective features and exhibits classification accuracy of 89% on E-Gait, which is 14 - 30% more accurate over prior methods.
11 pages, 7 figures, 1 table
References in corpus (4)
Cited by in corpus (13)
- Emotion Recognition from Multiple Modalities: Fundamentals and Methodologies
- Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
- Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
- Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
- Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
- PhysiQ: Off-site Quality Assessment of Exercise in Physical Therapy
- Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
- Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
- HighlightMe: Detecting Highlights from Human-Centric Videos
- VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
- Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization