activity
20212024
most citedThe Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement Learning

3 citations · 11 across the 12 of their papers we have counts for

collaborators

12 papers

cs.LG20242 cited

Human Alignment of Large Language Models through Online Preference Optimisation

Daniele Calandriello, Daniel Guo, Remi Munos +10

Ensuring alignment of language models' outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensiv…

cs.LG2024

Off-policy Distributional Q(): Distributional RL without Importance Sampling

Yunhao Tang, Mark Rowland, Rémi Munos +2

We introduce off-policy distributional Q(), a new addition to the family of off-policy distributional evaluation algorithms. Off-policy distributional Q() does not apply impo…

cs.LG20241 cited

Learning Uncertainty-Aware Temporally-Extended Actions

Joongkyu Lee, Seung Joon Park, Yunhao Tang +1

In reinforcement learning, temporal abstraction in the action space, exemplified by action repetition, is a technique to facilitate policy learning through extended actions. Howeve…

cs.LG2023

DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm

Yunhao Tang, Tadashi Kozuno, Mark Rowland +4

Multi-step learning applies lookahead over multiple time steps and has proved valuable in policy evaluation settings. However, in the optimal control case, the impact of multi-step…

cs.LG2023

Towards a Better Understanding of Representation Dynamics under TD-learning

Yunhao Tang, Rémi Munos

TD-learning is a foundation reinforcement learning (RL) algorithm for value prediction. Critical to the accuracy of value predictions is the quality of state representations. In th…

cs.LG2023

The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation

Mark Rowland, Yunhao Tang, Clare Lyle +3

We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorith…