1 citations · 1 across the 5 of their papers we have counts for
3 papers · 1 filter
Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning
Anthony GX-Chen, Ankit Anand, Gheorghe Comanici +7
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fin…
Code as Reward: Empowering Reinforcement Learning with VLMs
David Venuto, Sami Nur Islam, Martin Klissarov +3
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In t…
Policy composition in reinforcement learning via multi-objective policy optimization
Shruti Mishra, Ankit Anand, Jordan Hoffmann +4
We enable reinforcement learning agents to learn successful behavior policies by utilizing relevant pre-existing teacher policies. The teacher policies are introduced as objectives…