32 citations · 45 across the 4 of their papers we have counts for
9 papers
VidTr: Video Transformer Without Convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu +6
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatio-temporal infor…
Multi-Label Activity Recognition using Activity-specific Features and Activity Correlations
Yanyi Zhang, Xinyu Li, Ivan Marsic
Multi-label activity recognition is designed for recognizing multiple activities that are performed simultaneously or sequentially in each video. Most recent activity recognition n…
RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement
Jalal Abdulbaqi, Yue Gu, Ivan Marsic
Most current speech enhancement models use spectrogram features that require an expensive transformation and result in phase information loss. Previous work has overcome these issu…
Tri-axial Self-Attention for Concurrent Activity Recognition
Yanyi Zhang, Xinyu Li, Kaixiang Huang +3
We present a system for concurrent activity recognition. To extract features associated with different activities, we propose a feature-to-activity attention that maps the extracte…
Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level Alignment
Yue Gu, Kangning Yang, Shiyu Fu +3
Multimodal affective computing, learning to recognize and interpret human affects and subjective information from multiple data sources, is still challenging because: (i) it is har…
Deep Multimodal Learning for Emotion Recognition in Spoken Language
Yue Gu, Shuhong Chen, Ivan Marsic
In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics.…