most citedUnderstanding Self-Predictive Learning for Reinforcement Learning

1 citations · 2 across the 3 of their papers we have counts for

collaborators

10 papers

cs.LG2024

Averaging log-likelihoods in direct alignment

Nathan Grinsztajn, Yannis Flet-Berliac, Mohammad Gheshlaghi Azar +8

To better align Large Language Models (LLMs) with human judgment, Reinforcement Learning from Human Feedback (RLHF) learns a reward model and then optimizes it using regularized RL…

cs.LG2024

A/B testing under Interference with Partial Network Information

Shiv Shankar, Ritwik Sinha, Yash Chandak +2

A/B tests are often required to be conducted on subjects that might have social connections. For e.g., experiments on social media, or medical and social interventions to control t…

cs.LG20232 cited

Behavior Alignment via Reward Function Optimization

Dhawal Gupta, Yash Chandak, Scott M. Jordan +2

Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identif…

cs.LG2023

Coagent Networks: Generalized and Scaled

James E. Kostas, Scott M. Jordan, Yash Chandak +5

Coagent networks for reinforcement learning (RL) [Thomas and Barto, 2011] provide a powerful and flexible framework for deriving principled learning rules for arbitrary stochastic…

cs.LG20231 cited

Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition

Yash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo +4

Representation learning and exploration are among the key challenges for any deep reinforcement learning agent. In this work, we provide a singular value decomposition based method…

cs.LG2023

Asymptotically Unbiased Off-Policy Policy Evaluation when Reusing Old Data in Nonstationary Environments

Vincent Liu, Yash Chandak, Philip Thomas +1

In this work, we consider the off-policy policy evaluation problem for contextual bandits and finite horizon reinforcement learning in the nonstationary setting. Reusing old data i…