2 papers
cs.RO2026
Transferring the Intelligence of VLMs to Robotic Control
Meng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang +6
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself m…
cs.CV2025
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
Meng-Hao Guo, Xuanyu Chu, Qianrui Yang +12
The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini, and o3, with their capability to process and generate content across modalities s…