3 citations · 11 across the 12 of their papers we have counts for
12 papers
Human Alignment of Large Language Models through Online Preference Optimisation
Daniele Calandriello, Daniel Guo, Remi Munos +10
Ensuring alignment of language models' outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensiv…
Off-policy Distributional Q(): Distributional RL without Importance Sampling
Yunhao Tang, Mark Rowland, Rémi Munos +2
We introduce off-policy distributional Q(), a new addition to the family of off-policy distributional evaluation algorithms. Off-policy distributional Q() does not apply impo…
Learning Uncertainty-Aware Temporally-Extended Actions
Joongkyu Lee, Seung Joon Park, Yunhao Tang +1
In reinforcement learning, temporal abstraction in the action space, exemplified by action repetition, is a technique to facilitate policy learning through extended actions. Howeve…
DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm
Yunhao Tang, Tadashi Kozuno, Mark Rowland +4
Multi-step learning applies lookahead over multiple time steps and has proved valuable in policy evaluation settings. However, in the optimal control case, the impact of multi-step…
Towards a Better Understanding of Representation Dynamics under TD-learning
Yunhao Tang, Rémi Munos
TD-learning is a foundation reinforcement learning (RL) algorithm for value prediction. Critical to the accuracy of value predictions is the quality of state representations. In th…
The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation
Mark Rowland, Yunhao Tang, Clare Lyle +3
We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorith…