Showing eess.ASShow all
3 papers · 1 filter
eess.AS2023★ 1 cited
The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction
Shilong Wu, Chenxi Wang, Hang Chen +13
Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advan…
eess.AS2023★ 3 cited
Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
Haotian Wang, Yuxuan Xi, Hang Chen +11
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as rob…
eess.AS2023
The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge
Ruoyu Wang, Maokui He, Jun Du +16
This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios…