303 citations · 631 across the 26 of their papers we have counts for
14 papers · 1 filter
Lip-reading with Hierarchical Pyramidal Convolution and Self-Attention
Hang Chen, Jun Du, Yu Hu +3
In this paper, we propose a novel deep learning architecture to improving word-level lip-reading. On the one hand, we first introduce the multi-scale processing into the spatial fe…
Exploring Emotion Features and Fusion Strategies for Audio-Video Emotion Recognition
Hengshun Zhou, Debin Meng, Yuanyuan Zhang +4
The audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotio…
The Third DIHARD Diarization Challenge
Neville Ryant, Prachi Singh, Venkat Krishnamohan +6
DIHARD III was the third in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variability in recording equipment, noise condit…
Frequency Gating: Improved Convolutional Neural Networks for Speech Enhancement in the Time-Frequency Domain
Koen Oostermeijer, Qing Wang, Jun Du
One of the strengths of traditional convolutional neural networks (CNNs) is their inherent translational invariance. However, for the task of speech enhancement in the time-frequen…
Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
Desh Raj, Pavel Denisov, Zhuo Chen +11
Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in syst…
A Two-Stage Approach to Device-Robust Acoustic Scene Classification
Hu Hu, Chao-Han Huck Yang, Xianjun Xia +13
To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convol…