4 papers · 1 filter
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
Junyu Wang, Ziyang Ma, Zhengding Luo +4
Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers…
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Junyu Wang, Tianrui Wang, Meng Ge +2
In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism…
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
Junyu Wang, Zizhen Lin, Tianrui Wang +3
In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mech…
MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
Zizhen Lin, Xiaoting Chen, Junyu Wang
Achieving a balance between lightweight design and high performance remains a challenging task for speech enhancement. In this paper, we introduce Multi-path Enhanced Taylor (MET)…