3 papers
cs.CV2025
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…
cs.RO2025
OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
Jilei Mao, Jiarui Guan, Yingjuan Tang +7
The visuomotor policy can easily overfit to its training datasets, such as fixed camera positions and backgrounds. This overfitting makes the policy perform well in the in-distribu…
cs.CV2025
Multi-Grained Compositional Visual Clue Learning for Image Intent Recognition
Yin Tang, Jiankai Li, Hongyu Yang +3
In an era where social media platforms abound, individuals frequently share images that offer insights into their intents and interests, impacting individual life quality and socie…