1 paper
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek +3
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained…