4 citations · 6 across the 11 of their papers we have counts for
4 papers · 1 filter
WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation
Mingda Lin, Lei Ding, Xinyue Zhou +6
While pre-trained models excel in specialized tasks, learning universal representations across diverse acoustic domains remains challenging. To address this, we propose WQ-Fusion,…
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
Ming Gao, Shilong Wu, Hang Chen +6
Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted a…
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
Jianyu Wang, Shanzheng Guan, Zhengqiao Zhao +2
Multichannel blind source separation (MBSS), which focuses on separating signals of interest from mixed observations, has been extensively studied in acoustic and speech processing…
Determined Blind Source Separation with Sinkhorn Divergence-based Optimal Allocation of the Source Power
Jianyu Wang, Shanzheng Guan, Nicolas Dobigeon +1
Blind source separation (BSS) refers to the process of recovering multiple source signals from observations recorded by an array of sensors. Common approaches to BSS, including ind…