149 citations · 348 across the 18 of their papers we have counts for
16 papers · 1 filter
Off-policy Distributional Q(): Distributional RL without Importance Sampling
Yunhao Tang, Mark Rowland, Rémi Munos +2
We introduce off-policy distributional Q(), a new addition to the family of off-policy distributional evaluation algorithms. Off-policy distributional Q() does not apply impo…
Generalized Preference Optimization: A Unified Approach to Offline Alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng +7
Offline preference optimization allows fine-tuning large models directly from offline data, and has proved effective in recent alignment practices. We propose generalized preferenc…
A Novel Stochastic Gradient Descent Algorithm for Learning Principal Subspaces
Charline Le Lan, Joshua Greaves, Jesse Farebrother +4
Many machine learning problems encode their data as a matrix with a possibly very large number of rows and columns. In several applications like neuroscience, image compression or…
Understanding Self-Predictive Learning for Reinforcement Learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13
We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…
Understanding and Preventing Capacity Loss in Reinforcement Learning
Clare Lyle, Mark Rowland, Will Dabney
The reinforcement learning (RL) problem is rife with sources of non-stationarity, making it a notoriously difficult problem domain for the application of neural networks. We identi…
Marginalized Operators for Off-policy Reinforcement Learning
Yunhao Tang, Mark Rowland, Rémi Munos +1
In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi…