activity
20172023
most citedOn Mean Absolute Error for Deep Neural Network Based Vector-to-Vector Regression

303 citations · 633 across the 29 of their papers we have counts for

collaborators
Showing 2023Show all

5 papers · 1 filter

eess.AS2023

The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge

Ruoyu Wang, Maokui He, Jun Du +16

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios…

cs.CL2023

Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder

Yusheng Dai, Hang Chen, Jun Du +4

In recent research, slight performance improvement is observed from automatic speech recognition systems to audio-visual speech recognition systems in the end-to-end framework with…

eess.AS20231 cited

Semi-supervised multi-channel speaker diarization with cross-channel attention

Shilong Wu, Jun Du, Maokui He +4

Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised…

eess.AS2023

Variance-Preserving-Based Interpolation Diffusion Models for Speech Enhancement

Zilu Guo, Jun Du, Chin-Hui Lee +2

The goal of this study is to implement diffusion models for speech enhancement (SE). The first step is to emphasize the theoretical foundation of variance-preserving (VP)-based int…

cs.MM20231 cited

The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition

Zhe Wang, Shilong Wu, Hang Chen +12

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research…