Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates
Zikang Huang, Meng Ge, Tianrui Wang +4
Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to…
cs.SD2025
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
Junyu Wang, Ziyang Ma, Zhengding Luo +4
Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers…