activity
20242026
collaborators
Showing 2024Show all

5 papers · 1 filter

eess.AS2024

DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions

Shu-Tong Niu, Jun Du, Ruo-Yu Wang +4

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarizati…

cs.CV2024

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

Haotian Wang, Yuzhe Weng, Yueyan Li +10

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In t…

eess.AS2024

The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge

Shutong Niu, Ruoyu Wang, Jun Du +17

This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conferenc…

eess.AS2024

The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge

Ya Jiang, Hongbo Lan, Jun Du +2

In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a pri…

eess.AS2024

Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings

Ruoyu Wang, Shutong Niu, Gaobin Yang +4

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to…