4 papers
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
Ke Xue, Rongfei Fan, Kai Li +3
Diffusion models have recently set new benchmarks in Speech Enhancement (SE). However, most existing score-based models treat speech spectrograms merely as generic 2D images, apply…
Omni-directional attention mechanism based on Mamba for speech separation
Ke Xue, Chang Sun, Rongfei Fan +2
Mamba, a selective state-space model (SSM), has emerged as an efficient alternative to Transformers for speech modeling, enabling long-sequence processing with linear complexity. W…
From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation
Ke Xue, Rongfei Fan, Lixin +3
Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual informatio…
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
Ke Xue, Rongfei Fan, Shanping Yu +2
Target speaker extraction focuses on extracting a target speech signal from an environment with multiple speakers by leveraging an enrollment. Existing methods predominantly rely o…