2 papers
cs.SD2025
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
Hualei Wang, Yiming Li, Shuo Ma +2
Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately…
cs.SD2024
Leveraging Language Model Capabilities for Sound Event Detection
Hualei Wang, Jianguo Mao, Zhifang Guo +3
Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality,…