5 papers
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
Fan-Yun Sun, Shengguang Wu, Christian Jacobsen +13
Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of…
Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective
Xuning Yang, Clemens Eppner, Jonathan Tremblay +3
Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluati…
GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
Adithyavairavan Murali, Balakumar Sundaralingam, Yu-Wei Chao +7
Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize acro…
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
Ishika Singh, Ankit Goyal, Stan Birchfield +3
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware…
BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation
Van Nguyen Nguyen, Stephen Tyree, Andrew Guo +16
We present the evaluation methodology, datasets and results of the BOP Challenge 2024, the 6th in a series of public competitions organized to capture the state of the art in 6D ob…