Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek +3
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained…
cs.SD2025
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
Joanna Hong, Sanjeel Parekh, Honglie Chen +4
Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in perfor…