1 paper
Yupei Li, Qiyang Sun, Mohamed Mady +4
Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern match…