DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
arXiv:2210.02438 · doi:10.1109/LRA.2023.3272516
Abstract
We introduce the first work to explore web-scale diffusion models for robotics. DALL-E-Bot enables a robot to rearrange objects in a scene, by first inferring a text description of those objects, then generating an image representing a natural, human-like arrangement of those objects, and finally physically arranging the objects according to that goal image. We show that this is possible zero-shot using DALL-E, without needing any further example arrangements, data collection, or training. DALL-E-Bot is fully autonomous and is not restricted to a pre-defined set of objects or scenes, thanks to DALL-E's web-scale pre-training. Encouraging real-world results, with both human studies and objective metrics, show that integrating web-scale diffusion models into robotics pipelines is a promising direction for scalable, unsupervised robot learning.
Webpage and videos: ( https://www.robot-learning.uk/dall-e-bot ) Published in IEEE Robotics and Automation Letters (RA-L)
References in corpus (12)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Large Language Models are Zero-Shot Reasoners
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- Rearrangement: A Challenge for Embodied AI
- CLIPort: What and Where Pathways for Robotic Manipulation
- Planning with Diffusion for Flexible Behavior Synthesis
- Learning Object Arrangements in 3D Scenes using Human Context
- Testing Relational Understanding in Text-Guided Image Generation
- Learning Universal Policies via Text-Guided Video Generation
- CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning
- Transformers are Adaptable Task Planners
- Where To Start? Transferring Simple Skills to Complex Environments
Cited by in corpus (5)
- Real-World Robot Applications of Foundation Models: A Review
- Language Models as Zero-Shot Trajectory Generators
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- LLM-GROP: Visually Grounded Robot Task and Motion Planning with Large Language Models
- Robust Robotic Exploration and Mapping Using Generative Occupancy Map Synthesis