10 papers
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
Enshen Zhou, Yibo Li, Jingkun An +12
Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spa…
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
Hao Wang, Xiaobao Wei, Jingyang He +10
Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image d…
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
Mengzhen Liu, Enshen Zhou, Cheng Chi +6
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoi…
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
Yi Han, Enshen Zhou, Shanyu Rong +6
Vision-Language Models (VLMs) have shown remarkable capabilities in spatial reasoning, yet they remain fundamentally limited to qualitative precision and lack the computational pre…
RoboBrain 2.5: Depth in Sight, Time in Mind
Huajie Tan, Enshen Zhou, Zhiyu Li +32
We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on…