4 papers
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek +3
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained…
Hearing Anywhere in Any Environment
Xiulong Liu, Anurag Kumar, Paul Calamia +7
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances…
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
Joanna Hong, Sanjeel Parekh, Honglie Chen +4
Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in perfor…
Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts
Shrestha Mohanty, Sarah Xuan, Jacob Jobraeel +3
We focus on enhancing comprehension in small-group recorded conversations, which serve as a medium to bring people together and provide a space for sharing personal stories and exp…