15 papers
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
Xiaohao Xu, Feng Xue, Xiang Li +5
A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometrically valid surfaces. Monocular depth est…
Latent Geometry Beyond Search: Amortizing Planning in World Models
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging. This ra…
Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration
Jun Wang, Xiaohao Xu, Xiaonan Huang
Safe human--robot collaboration requires more than visual description: a monitor must determine whether the robot body is safely separated, already colliding with the scene or a pe…
The 3D Mirage: Probing and Taming 3D Hallucinations
Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerability: they hallucinate illusory 3D…
Bridging 3D Anomaly Localization and Repair via High-Quality Continuous Geometric Representation
Bozhong Zheng, Jinye Gan, Xiaohao Xu +5
3D point cloud anomaly detection is essential for robust vision systems but is challenged by pose variations and complex geometric anomalies. Existing patch-based methods often suf…
Natural Selection via Foundation Models for Soft Robot Evolution
Changhe Chen, Xiaohao Xu, Xiangdong Wang +1
Designing soft robots is a complex and iterative process that demands cross-disciplinary expertise in materials science, mechanics, and control, often relying on intuition and exte…