4 citations · 4 across the 1 of their papers we have counts for
1 paper · 1 filter
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek +3
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained…