1 paper
Shuo-Chun Lin, Hen-Hsen Huang
Multimodal Large Language Models (LLMs) have remarkable semantic audio understanding, yet they remain "spatially agnostic" due to their reliance on mono-channel audio representatio…