14 papers
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
Xiaohao Xu, Feng Xue, Xiang Li +5
A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces. Monocular depth est…
Latent Geometry Beyond Search: Amortizing Planning in World Models
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging. This ra…
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
Jun Wang, Xiaohao Xu, Xiaonan Huang
Safe human--robot collaboration requires more than visual description: a monitor must determine whether the robot body is safely separated, already colliding with the scene or a pe…
When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills
Yunfei Wang, Xiaohao Xu, Yang Li +1
Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator results shape the next populati…
The 3D Mirage: Probing and Taming 3D Hallucinations
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerability: they hallucinate illusory 3D…
Natural Selection via Foundation Models for Soft Robot Evolution
Changhe Chen, Xiaohao Xu, Xiangdong Wang +1
Designing soft robots is a complex and iterative process that demands cross-disciplinary expertise in materials science, mechanics, and control, often relying on intuition and exte…