activity
20242026
collaborators

7 papers

eess.AS2026

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13

Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial rel…

cs.AI2026

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

Zhi Zheng, Ziqiao Meng, Hao Luan +2

External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing…

cs.SD2026

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

Ye Tao, Lupeng Liu, Xuenan Xu +6

Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, re…

cs.MM2026

TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering

Zhaoyang Xu, Xusheng He, Wei Liu +2

Temporal-logic video question answering requires a model to reason about when actions occur relative to one another, such as before, after, until, since, overlap, and multi-event c…

cs.RO2025

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

Xiaomeng Chu, Jiajun Deng, Guoliang You +4

Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities…

cs.CL2025

Graph-Based Multimodal Contrastive Learning for Chart Question Answering

Yue Dai, Soyeon Caren Han, Wei Liu

Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode. This work introduces a novel joint mul…