4 papers
Vero: An Open RL Recipe for General Visual Reasoning
Gabriel Sarch, Linrong Cai, Qunzhong Wang +3
What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest tha…
Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
Xueying Li, Feng Lyu, Hao Wu +3
Training-free Vision-Language Navigation (VLN) agents powered by foundation models can follow instructions and explore 3D environments. However, existing approaches rely on greedy…
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
Ji Zhang, Shihan Wu, Xu Luo +4
Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great pr…
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
Xiaoran Yin, Xu Luo, Hao Wu +2
The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challeng…