1 paper
Yuhang Dai, Xin Shu, Zengxi Li +3
Large audio language models (LALMs) can describe what is heard, but their ability to localize when queried content occurs remains less systematically evaluated. We present TAG-Benc…