7 papers
Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial rel…
One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA
Zhi Zheng, Ziqiao Meng, Hao Luan +2
External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing…
Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
Ye Tao, Lupeng Liu, Xuenan Xu +6
Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, re…
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
Zhaoyang Xu, Xusheng He, Wei Liu +2
Temporal-logic video question answering requires a model to reason about when actions occur relative to one another, such as before, after, until, since, overlap, and multi-event c…
GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions
Xiaomeng Chu, Jiajun Deng, Guoliang You +4
Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities…
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
Yue Dai, Soyeon Caren Han, Wei Liu
Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode. This work introduces a novel joint mul…