5 papers
OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation
Tianyi Zhang, Ziyang Gong, Zhenjie Yang +2
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align…
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI
Brain Team, Ziyang Gong, Haoming Gu +28
Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve f…
IGen: Scalable Data Generation for Robot Learning from Open-World Images
Chenghao Gu, Haolan Kang, Junchao Lin +10
The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-intensive and often limited to s…
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
Songze Li, Zun Wang, Gengze Zhou +8
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…
Robotic Visual Instruction
Yanbang Li, Ziyang Gong, Haoyang Li +4
Recently, natural language has been the primary medium for human-robot interaction. However, its inherent lack of spatial precision introduces challenges for robotic task definitio…