collaborators

10 papers

cs.SD2026

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation

Mingda Lin, Lei Ding, Xinyue Zhou +6

While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion,…

eess.AS2026

Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming

Yongyi Deng, Hanchen Pei, Jianbo Ma +3

The minimum variance distortionless response (MVDR) beamformer is widely used for multichannel speech enhancement due to strong noise suppression while preserving target signals. I…

eess.AS2026

A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic

Jing Yang, Shuqing Zhang, Yongyi Deng +5

Accurate phoneme recognition is pivotal for mispronunciation detection and diagnosis (MDD) in modern standard Arabic (MSA), yet remains constrained by data scarcity and the synthet…

eess.AS2026

Robust Online Overdetermined Independent Vector Analysis Based on Bilinear Decomposition

Kang Chen, Xianrui Wang, Yichen Yang +6

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (Over…

cs.SD2025

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

Ming Gao, Shilong Wu, Hang Chen +6

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted a…

cs.CL2025

Estimating LLM Uncertainty with Evidence

Huan Ma, Jingdong Chen, Joey Tianyi Zhou +2

Over the past few years, Large Language Models (LLMs) have developed rapidly and are widely applied in various domains. However, LLMs face the issue of hallucinations, generating r…