4 papers
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
Yifan Zhou, Sachin Grover, Mohamed El Mistiri +7
Reinforcement Learning (RL) traditionally relies on scalar reward signals, limiting its ability to leverage the rich semantic knowledge often available in real-world tasks. In cont…
SAS-Prompt: Large Language Models as Numerical Optimizers for Robot Self-Improvement
Heni Ben Amor, Laura Graesser, Atil Iscen +7
We demonstrate the ability of large language models (LLMs) to perform iterative self-improvement of robot policies. An important insight of this paper is that LLMs have a built-in…
Task Success is not Enough: Investigating the Use of Video-Language Models as Behavior Critics for Catching Undesirable Agent Behaviors
Lin Guan, Yifan Zhou, Denis Liu +3
Large-scale generative models are shown to be useful for sampling meaningful candidate solutions, yet they often overlook task constraints and user preferences. Their full power is…
Enabling Stateful Behaviors for Diffusion-based Policy Learning
Xiao Liu, Fabian Weigend, Yifan Zhou +1
While imitation learning provides a simple and effective framework for policy learning, acquiring consistent actions during robot execution remains a challenging task. Existing app…