collaborators

5 papers

cs.SD2026

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation

Mingda Lin, Lei Ding, Xinyue Zhou +6

While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion,…

eess.AS2026

Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming

Yongyi Deng, Hanchen Pei, Jianbo Ma +3

The minimum variance distortionless response (MVDR) beamformer is widely used for multichannel speech enhancement due to strong noise suppression while preserving target signals. I…

cs.SD2026

A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation

Hanchen Pei, Shujie Liu, Yanqing Liu +5

Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody,…

eess.AS2025

Spatial-Filter-Bank-Based Neural Method for Multichannel Speech Enhancement

Tianqin Zheng, Jilu Jin, Hanchen Pei +3

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approac…

eess.AS2025

LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention

Yaokai Zhang, Hanchen Pei, Wanqi Wang +1

Deep learning based end-to-end multi-channel speech enhancement methods have achieved impressive performance by leveraging sub-band, cross-band, and spatial information. However, t…