Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
arXiv:2204.01691
Abstract
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment. We evaluate our method on a number of real-world robotic tasks, where we show the need for real-world grounding and that this approach is capable of completing long-horizon, abstract, natural language instructions on a mobile manipulator. The project's website and the video can be found at https://say-can.github.io/.
See website at https://say-can.github.io/ V1. Initial Upload. V2. Added PaLM results. Added study about new capabilities (drawer manipulation, chain of thought prompting, multilingual instructions). Added an ablation study of language model size. Added an open-source version of \algname on a simulated tabletop environment. Improved readability
Cited by in corpus (50)
- A Survey on Large Language Model based Autonomous Agents
- TidyBot: Personalized Robot Assistance with Large Language Models
- Understanding Large-Language Model (LLM)-powered Human-Robot Interaction
- Advances of Machine Learning in Materials Science: Ideas and Techniques
- ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application
- Large Language Models as Zero-Shot Human Models for Human-Robot Interaction
- Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
- Theory of Mind for Multi-Agent Collaboration via Large Language Models
- GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
- Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
- A Survey of Optimization-based Task and Motion Planning: From Classical To Learning Approaches
- On the Prospects of Incorporating Large Language Models (LLMs) in Automated Planning and Scheduling (APS)
- Deploying and Evaluating LLMs to Program Service Mobile Robots
- LLM-BT: Performing Robotic Adaptive Tasks based on Large Language Models and Behavior Trees
- Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- CoPAL: Corrective Planning of Robot Actions with Large Language Models
- Higher education assessment practice in the era of generative AI tools
- When Robots Get Chatty: Grounding Multimodal Human-Robot Conversation and Collaboration
- DRAGON: A Dialogue-Based Robot for Assistive Navigation with Visual Language Grounding
- Recent Advances of Deep Robotic Affordance Learning: A Reinforcement Learning Perspective
- On Grounded Planning for Embodied Tasks with Language Models
- REX: Designing User-centered Repair and Explanations to Address Robot Failures
- FASTNav: Fine-tuned Adaptive Small-language-models Trained for Multi-point Robot Navigation
- Exploring Large Language Models to Facilitate Variable Autonomy for Human-Robot Teaming
- LLM-Mediated Domain-Specific Voice Agents: The Case of TextileBot
- Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor Cues
- Panoptic Vision-Language Feature Fields
- Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes
- A Study of Situational Reasoning for Traffic Understanding
- TalkWithMachines: Enhancing Human-Robot Interaction for Interpretable Industrial Robotics Through Large/Vision Language Models
- Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
- BaRiFlex: A Robotic Gripper with Versatility and Collision Robustness for Robot Learning
- Hierarchical Path-planning from Speech Instructions with Spatial Concept-based Topometric Semantic Mapping
- Developmental Scaffolding with Large Language Models
- Bootstrapping Cognitive Agents with a Large Language Model
- Evaluating Language Model Agency through Negotiations
- DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
- PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation
- Camera Control at the Edge with Language Models for Scene Understanding
- Knowledge-enhanced Agents for Interactive Text Games
- HiddenTables & PyQTax: A Cooperative Game and Dataset For TableQA to Ensure Scale and Data Privacy Across a Myriad of Taxonomies
- Are Large Language Models Aligned with People's Social Intuitions for Human-Robot Interactions?
- LLM+Reasoning+Planning for Supporting Incomplete User Queries in Presence of APIs
- Interactive Task Encoding System for Learning-from-Observation
- Skill Generalization with Verbs
- Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
- KeyMPs: One-Shot Vision-Language Guided Motion Generation by Sequencing DMPs for Occlusion-Rich Tasks
- Remember what you did so you know what to do next
- SCOPE: Real-Time Natural Language Camera Agent at the Edge