Spatial-Temporal Recurrent Neural Network for Emotion Recognition
arXiv:1705.04515 · doi:10.1109/TCYB.2017.2788081
Abstract
Emotion analysis is a crucial problem to endow artifact machines with real intelligence in many large potential applications. As external appearances of human emotions, electroencephalogram (EEG) signals and video face signals are widely used to track and analyze human's affective information. According to their common characteristics of spatial-temporal volumes, in this paper we propose a novel deep learning framework named spatial-temporal recurrent neural network (STRNN) to unify the learning of two different signal sources into a spatial-temporal dependency model. In STRNN, to capture those spatially cooccurrent variations of human emotions, a multi-directional recurrent neural network (RNN) layer is employed to capture longrange contextual cues by traversing the spatial region of each time slice from multiple angles. Then a bi-directional temporal RNN layer is further used to learn discriminative temporal dependencies from the sequences concatenating spatial features of each time slice produced from the spatial RNN layer. To further select those salient regions of emotion representation, we impose sparse projection onto those hidden states of spatial and temporal domains, which actually also increases the model discriminant ability because of this global consideration. Consequently, such a two-layer RNN model builds spatial dependencies as well as temporal dependencies of the input signals. Experimental results on the public emotion datasets of EEG and facial expression demonstrate the proposed STRNN method is more competitive over those state-of-the-art methods.
References in corpus (1)
Cited by in corpus (28)
- EEG based Emotion Recognition: A Tutorial and Review
- TSception: Capturing Temporal Dynamics and Spatial Asymmetry from EEG for Emotion Recognition
- Deep Learning for Human Affect Recognition: Insights and New Developments
- EEG-ITNet: An Explainable Inception Temporal Convolutional Network for Motor Imagery Classification
- Deep Learning in EEG: Advance of the Last Ten-Year Critical Period
- Transformer-based Spatial-Temporal Feature Learning for EEG Decoding
- Source Aware Deep Learning Framework for Hand Kinematic Reconstruction using EEG Signal
- PARSE: Pairwise Alignment of Representations in Semi-Supervised EEG Learning for Emotion Recognition
- Disentangling Identity and Pose for Facial Expression Recognition
- EEG-Based Emotion Recognition Using Regularized Graph Neural Networks
- Spatio-Temporal EEG Representation Learning on Riemannian Manifold and Euclidean Space
- Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot Learning
- Deep Feature Mining via Attention-based BiLSTM-GCN for Human Motor Imagery Recognition
- Multi-Modal Deep Learning for Credit Rating Prediction Using Text and Numerical Data Streams
- Multi-hop Convolutions on Weighted Graphs
- A Survey on Deep Learning-based Non-Invasive Brain Signals:Recent Advances and New Frontiers
- Recurrent Embedding Aggregation Network for Video Face Recognition
- HetEmotionNet: Two-Stream Heterogeneous Graph Recurrent Neural Network for Multi-modal Emotion Recognition
- LaFurca: Iterative Refined Speech Separation Based on Context-Aware Dual-Path Parallel Bi-LSTM
- Distilling EEG Representations via Capsules for Affective Computing
- Topological EEG Nonlinear Dynamics Analysis for Emotion Recognition
- On the Pitfalls of Learning with Limited Data: A Facial Expression Recognition Case Study
- Community Resilience Optimization Subject to Power Flow Constraints in Cyber-Physical-Social Systems in Power Engineering
- MICACL: Multi-Instance Category-Aware Contrastive Learning for Long-Tailed Dynamic Facial Expression Recognition
- Deep Recurrent Semi-Supervised EEG Representation Learning for Emotion Recognition
- Convolutional Neural Network for emotion recognition to assist psychiatrists and psychologists during the COVID-19 pandemic: experts opinion
- Emotion Correlation Mining Through Deep Learning Models on Natural Language Text
- Self-attention aggregation network for video face representation and recognition