7 papers
RoboVista: Evaluating Vision Language Models for Diverse Robot Applications
Shuangyu Xie, Kaiyuan Chen, Ziyang Chen +8
Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-…
MonoDuo: Using One Robot Arm to Learn Bimanual Policies
Sandeep Bajamahal, Lawrence Yunliang Chen, Toru Lin +3
Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-a…
OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
Guanhua Ji, Harsha Polavaram, Lawrence Yunliang Chen +5
Large and diverse datasets are needed for training generalist robot policies that have potential to control a variety of robot embodiments -- robot arm and gripper combinations --…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
Omni-Scan: Creating Visually-Accurate Digital Twin Object Models Using a Bimanual Robot with Handover and Gaussian Splat Merging
Tianshuang Qiu, Zehan Ma, Karim El-Refai +4
3D Gaussian Splats (3DGSs) are 3D object models derived from multi-view images. Such "digital twins" are useful for simulations, virtual reality, marketing, robot policy fine-tunin…
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Kaiyuan Chen, Shuangyu Xie, Zehan Ma +2
Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene unde…