collaborators

5 papers

cs.AI2025

RLZero: Direct Policy Inference from Language Without In-Domain Supervision

Harshit Sikchi, Siddhant Agarwal, Pranaya Jajoo +6

The reward hypothesis states that all goals and purposes can be understood as the maximization of a received scalar reward signal. However, in practice, defining such a reward sign…

cs.LG2025

Offline Action-Free Learning of Ex-BMDPs by Comparing Diverse Datasets

Alexander Levine, Peter Stone, Amy Zhang

While sequential decision-making environments often involve high-dimensional observations, not all features of these observations are relevant for control. In particular, the obser…

cs.LG2025

Learning a Fast Mixing Exogenous Block MDP using a Single Trajectory

Alexander Levine, Peter Stone, Amy Zhang

In order to train agents that can quickly adapt to new objectives or reward functions, efficient unsupervised representation learning in sequential decision-making environments can…

cs.LG2025

Proto Successor Measure: Representing the Behavior Space of an RL Agent

Siddhant Agarwal, Harshit Sikchi, Peter Stone +1

Having explored an environment, intelligent agents should be able to transfer their knowledge to most downstream tasks within that environment without additional interactions. Refe…

cs.RO2025

Reinforcement Learning Within the Classical Robotics Stack: A Case Study in Robot Soccer

Adam Labiosa, Zhihan Wang, Siddhant Agarwal +10

Robot decision-making in partially observable, real-time, dynamic, and multi-agent environments remains a difficult and unsolved challenge. Model-free reinforcement learning (RL) i…