18 citations · 23 across the 3 of their papers we have counts for
3 papers
Reinforced Self-Training (ReST) for Language Modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan +11
Reinforcement learning from human feedback (RLHF) can improve the quality of large language model's (LLM) outputs by aligning them with human preferences. We propose a simple algor…
AlphaStar Unplugged: Large-Scale Offline Reinforcement Learning
Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan +21
StarCraft II is one of the most challenging simulated reinforcement learning environments; it is partially observable, stochastic, multi-agent, and mastering StarCraft II requires…
An Empirical Study of Implicit Regularization in Deep Offline RL
Caglar Gulcehre, Srivatsan Srinivasan, Jakub Sygnowski +5
Deep neural networks are the most commonly used function approximators in offline reinforcement learning. Prior works have shown that neural nets trained with TD-learning and gradi…