activity
20122024
most citedStatistical Linear Estimation with Penalized Estimators: an Application to Reinforcement Learning

17 citations · 34 across the 8 of their papers we have counts for

collaborators

8 papers

cs.LG2024

A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning

Khimya Khetarpal, Zhaohan Daniel Guo, Bernardo Avila Pires +7

Learning a good representation is a crucial challenge for Reinforcement Learning (RL) agents. Self-predictive learning provides means to jointly learn a latent representation and d…

cs.LG20242 cited

Human Alignment of Large Language Models through Online Preference Optimisation

Daniele Calandriello, Daniel Guo, Remi Munos +10

Ensuring alignment of language models' outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensiv…

cs.LG2024

Off-policy Distributional Q(): Distributional RL without Importance Sampling

Yunhao Tang, Mark Rowland, Rémi Munos +2

We introduce off-policy distributional Q(), a new addition to the family of off-policy distributional evaluation algorithms. Off-policy distributional Q() does not apply impo…

cs.LG2023

DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm

Yunhao Tang, Tadashi Kozuno, Mark Rowland +4

Multi-step learning applies lookahead over multiple time steps and has proved valuable in policy evaluation settings. However, in the optimal control case, the impact of multi-step…

cs.LG2023

Hierarchical Reinforcement Learning in Complex 3D Environments

Bernardo Avila Pires, Feryal Behbahani, Hubert Soyer +3

Hierarchical Reinforcement Learning (HRL) agents have the potential to demonstrate appealing capabilities such as planning and exploration with abstraction, transfer, and skill reu…

cs.LG20223 cited

The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement Learning

Yunhao Tang, Mark Rowland, Rémi Munos +3

We study the multi-step off-policy learning approach to distributional RL. Despite the apparent similarity between value-based RL and distributional RL, our study reveals intriguin…