activity
20242026
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates

Zikang Huang, Meng Ge, Tianrui Wang +4

Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to…

cs.SD2025

Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

Junyu Wang, Ziyang Ma, Zhengding Luo +4

Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers…

cs.SD2025

ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning

Junyu Wang, Tianrui Wang, Meng Ge +2

In recent advancements in audio self-supervised representation learning, the standard Transformer architecture has emerged as the predominant approach, yet its attention mechanism…

cs.SD2025

Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module

Zhongjian Cui, Chenrui Cui, Tianrui Wang +6

The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an eff…

cs.SD2025

Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement

Junyu Wang, Zizhen Lin, Tianrui Wang +3

In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mech…