3 papers
cs.SD2026
From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
Yujian Ma, Jinqiu Sang, Ruizhe Li +2
Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Usin…
eess.AS2025
A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
Xikun Lu, Yujian Ma, Xianquan Jiang +2
Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutio…
eess.AS2025
Lightweight Implicit Neural Network for Binaural Audio Synthesis
Xikun Lu, Fang Liu, Weizhi Shi +1
High-fidelity binaural audio synthesis is crucial for immersive listening, but existing methods require extensive computational resources, limiting their edge-device application. T…