15 papers
Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
Ben Zandonati, Tomás Lozano-Pérez, Leslie Pack Kaelbling
Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints. Modern robot imitation learning,…
Which Reconstruction Model Should a Robot Use? Routing Image-to-3D Models for Cost-Aware Robotic Manipulation
Akash Anand, Aditya Agarwal, Leslie Pack Kaelbling
Robotic manipulation tasks require 3D mesh reconstructions of varying quality: dexterous manipulation demands fine-grained surface detail, while collision-free planning tolerates c…
Open-World Task and Motion Planning via Vision-Language Model Generated Constraints
Nishanth Kumar, William Shen, Fabio Ramos +4
Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve comp…
TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
William Shen, Nishanth Kumar, Sahit Chintalapudi +8
We present TiPToP, a modular manipulation system that integrates pretrained foundation models with a GPU-accelerated Task and Motion Planner to solve tasks directly from RGB images…
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
Ashay Athalye, Nishanth Kumar, Tom Silver +4
Our aim is to learn to solve long-horizon decision-making problems in complex robotics domains given low-level skills and a handful of short-horizon demonstrations containing seque…
SceneComplete: Open-World 3D Scene Completion in Cluttered Real World Environments for Robot Manipulation
Aditya Agarwal, Gaurav Singh, Bipasha Sen +2
Careful robot manipulation in every-day cluttered environments requires an accurate understanding of the 3D scene, in order to grasp and place objects stably and reliably and to av…