Language Models as Zero-Shot Trajectory Generators
arXiv:2310.11604 · doi:10.1109/LRA.2024.3410155
Abstract
Large Language Models (LLMs) have recently shown promise as high-level planners for robots when given access to a selection of low-level skills. However, it is often assumed that LLMs do not possess sufficient knowledge to be used for the low-level trajectories themselves. In this work, we address this assumption thoroughly, and investigate if an LLM (GPT-4) can directly predict a dense sequence of end-effector poses for manipulation tasks, when given access to only object detection and segmentation vision models. We designed a single, task-agnostic prompt, without any in-context examples, motion primitives, or external trajectory optimisers. Then we studied how well it can perform across 30 real-world language-based tasks, such as "open the bottle cap" and "wipe the plate with the sponge", and we investigated which design choices in this prompt are the most important. Our conclusions raise the assumed limit of LLMs for robotics, and we reveal for the first time that LLMs do indeed possess an understanding of low-level robot control sufficient for a range of common tasks, and that they can additionally detect failures and then re-plan trajectories accordingly. Videos, prompts, and code are available at: https://www.robot-learning.uk/language-models-trajectory-generators.
Published in IEEE Robotics and Automation Letters (Volume: 9, Issue: 7, July 2024, Pages: 6728-6735); 10 pages, 12 figures
References in corpus (7)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Survey on Large Language Model based Autonomous Agents
- Evaluating Large Language Models Trained on Code
- Gemini: A Family of Highly Capable Multimodal Models
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
- Towards A Unified Agent with Foundation Models
- Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics
Cited by in corpus (7)
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- TalkWithMachines: Enhancing Human-Robot Interaction for Interpretable Industrial Robotics Through Large/Vision Language Models
- Uncrewed Vehicles in 6G Networks: A Unifying Treatment of Problems, Formulations, and Tools
- ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning
- Toward Generalist Neural Motion Planners for Robotic Manipulators: Challenges and Opportunities
- OpenNav: Open-World Navigation with Multimodal Large Language Models
- Action Tokenizer Matters in In-Context Imitation Learning