2 papers
cs.LG2023
The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation
Mark Rowland, Yunhao Tang, Clare Lyle +3
We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorith…
cs.LG2021
DARTS without a Validation Set: Optimizing the Marginal Likelihood
Miroslav Fil, Binxin Ru, Clare Lyle +1
The success of neural architecture search (NAS) has historically been limited by excessive compute requirements. While modern weight-sharing NAS methods such as DARTS are able to f…