Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition
arXiv:2008.09412 · doi:10.1109/TIP.2021.3087348
Abstract
Gesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to fully explore synergies among spatio-temporal modalities effectively for gesture recognition. The problems are partially due to the fact that the existing manually designed network architectures have low efficiency in the joint learning of multi-modalities. In this paper, we propose the first neural architecture search (NAS)-based method for RGB-D gesture recognition. The proposed method includes two key components: 1) enhanced temporal representation via the proposed 3D Central Difference Convolution (3D-CDC) family, which is able to capture rich temporal context via aggregating temporal difference information; and 2) optimized backbones for multi-sampling-rate branches and lateral connections among varied modalities. The resultant multi-modal multi-rate network provides a new perspective to understand the relationship between RGB and depth modalities and their temporal dynamics. Comprehensive experiments are performed on three benchmark datasets (IsoGD, NvGesture, and EgoGesture), demonstrating the state-of-the-art performance in both single- and multi-modality settings.The code is available at https://github.com/ZitongYu/3DCDC-NAS
Submitted to IEEE Transactions on Image Processing
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Neural Architecture Search with Reinforcement Learning
- NAS-FAS: Static-Dynamic Central Difference Network Search for Face Anti-Spoofing
- AutoHR: A Strong End-to-end Baseline for Remote Heart Rate Measurement with Neural Searching
- Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition
- Scheduled Differentiable Architecture Search for Visual Recognition
Cited by in corpus (12)
- NAS-FAS: Static-Dynamic Central Difference Network Search for Face Anti-Spoofing
- A Central Difference Graph Convolutional Operator for Skeleton-Based Action Recognition
- Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
- Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
- StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition
- Multi-Task and Multi-Modal Learning for RGB Dynamic Gesture Recognition
- Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
- Pixel Difference Networks for Efficient Edge Detection
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference Transformer
- Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN
- Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
- iMiGUE: An Identity-free Video Dataset for Micro-Gesture Understanding and Emotion Analysis