activity
20202026
most citedDevice-Robust Acoustic Scene Classification Based on Two-Stage Categorization and Data Augmentation

46 citations · 74 across the 11 of their papers we have counts for

collaborators

14 papers

eess.AS2026

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

Ya Jiang, Ruoyu Wang, Jingxuan Zhang +19

This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike con…

cs.SD2025

Exploring Speaker Diarization with Mixture of Experts

Gaobin Yang, Maokui He, Shutong Niu +3

In this paper, we propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates a…

eess.AS20241 cited

DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions

Shu-Tong Niu, Jun Du, Ruo-Yu Wang +4

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarizati…

cs.CV20241 cited

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

Haotian Wang, Yuzhe Weng, Yueyan Li +10

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In t…

eess.AS2024

The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge

Shutong Niu, Ruoyu Wang, Jun Du +17

This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conferenc…

cs.MM2024

Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization

Mao-Kui He, Jun Du, Shu-Tong Niu +2

In this paper, we propose a quality-aware end-to-end audio-visual neural speaker diarization framework, which comprises three key techniques. First, our audio-visual model takes bo…