376 citations · 383 across the 3 of their papers we have counts for
3 papers
Model Selection in Batch Policy Optimization
Jonathan N. Lee, George Tucker, Ofir Nachum +1
We study the problem of model selection in batch policy optimization: given a fixed, partial-feedback dataset and model classes, learn a policy with performance that is competi…
DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma +3
Despite overparameterization, deep networks trained via supervised learning are easy to optimize and exhibit excellent generalization. One hypothesis to explain this is that overpa…
Regularizing Neural Networks by Penalizing Confident Output Distributions
Gabriel Pereyra, George Tucker, Jan Chorowski +2
We systematically explore regularizing neural networks by penalizing low entropy output distributions. We show that penalizing low entropy output distributions, which has been show…