Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
arXiv:2309.04316 · doi:10.3389/frobt.2024.1455375
Abstract
Natural-language dialog is key for intuitive human-robot interaction. It can be used not only to express humans' intents, but also to communicate instructions for improvement if a robot does not understand a command correctly. Of great importance is to endow robots with the ability to learn from such interaction experience in an incremental way to allow them to improve their behaviors or avoid mistakes in the future. In this paper, we propose a system to achieve incremental learning of complex behavior from natural interaction, and demonstrate its implementation on a humanoid robot. Building on recent advances, we present a system that deploys Large Language Models (LLMs) for high-level orchestration of the robot's behavior, based on the idea of enabling the LLM to generate Python statements in an interactive console to invoke both robot perception and action. The interaction loop is closed by feeding back human instructions, environment observations, and execution results to the LLM, thus informing the generation of the next statement. Specifically, we introduce incremental prompt learning, which enables the system to interactively learn from its mistakes. For that purpose, the LLM can call another LLM responsible for code-level improvements of the current interaction based on human feedback. The improved interaction is then saved in the robot's memory, and thus retrieved on similar requests. We integrate the system in the robot cognitive architecture of the humanoid robot ARMAR-6 and evaluate our methods both quantitatively (in simulation) and qualitatively (in simulation and real-world) by demonstrating generalized incrementally-learned knowledge.
This version (v3) adds further quantitative evaluation and many improvements. v2 was presented at the Workshop on Language and Robot Learning (LangRob) at the Conference on Robot Learning (CoRL) 2023. Supplementary video available at https://youtu.be/y5O2mRGtsLM
References in corpus (27)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Evaluating Large Language Models Trained on Code
- ReAct: Synergizing Reasoning and Acting in Language Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- TidyBot: Personalized Robot Assistance with Large Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- Augmented Language Models: a Survey
- ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application
- AgentBench: Evaluating LLMs as Agents
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- TALM: Tool Augmented Language Models
- Compositional Exemplars for In-context Learning
- Interactive Code Generation via Test-Driven User-Intent Formalization
- Interactive Natural Language Processing
- A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
- A Memory System of a Robot Cognitive Architecture and its Implementation in ArmarX
- Errors are Useful Prompts: Instruction Guided Task Programming with Verifier-Assisted Iterative Prompting
- Language Models Can Teach Themselves to Program Better
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
- If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents
- InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
- Interactive and Incremental Learning of Spatial Object Relations from Human Demonstrations
Cited by in corpus (4)
- To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions
- Visual Environment-Interactive Planning for Embodied Complex-Question Answering
- A Framework for Adapting Human-Robot Interaction to Diverse User Groups
- Episodic Memory Verbalization using Hierarchical Representations of Life-Long Robot Experience