5 papers
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo +6
The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward sign…
Offline Action-Free Learning of Ex-BMDPs by Comparing Diverse Datasets
Alexander Levine, Peter Stone, Amy Zhang
While sequential decision-making environments often involve high-dimensional observations, not all features of these observations are relevant for control. In particular, the obser…
Learning a Fast Mixing Exogenous Block MDP using a Single Trajectory
Alexander Levine, Peter Stone, Amy Zhang
In order to train agents that can quickly adapt to new objectives or reward functions, efficient unsupervised representation learning in sequential decision-making environments can…
Proto Successor Measure: Representing the Behavior Space of an RL Agent
Siddhant Agarwal, Harshit Sikchi, Peter Stone +1
Having explored an environment, intelligent agents should be able to transfer their knowledge to most downstream tasks within that environment without additional interactions. Refe…
Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer
Adam Labiosa, Zhihan Wang, Siddhant Agarwal +10
Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) i…