1 paper · 1 filter
Luoyi Sun, Xiao Zhou, Zeqian Li +3
Large Audio-Language Models (ALMs) have recently demonstrated remarkable capabilities in holistic audio understanding, yet they remain unreliable for temporal grounding, i.e., the…