activity
20182026
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations · 3.5k across the 13 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.LG2024

A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning

Khimya Khetarpal, Zhaohan Daniel Guo, Bernardo Avila Pires +7

Learning a good representation is a crucial challenge for Reinforcement Learning (RL) agents. Self-predictive learning provides means to jointly learn a latent representation and d…

cs.LG2024

Offline Regularised Reinforcement Learning for Large Language Models Alignment

Pierre Harvey Richemond, Yunhao Tang, Daniel Guo +15

The dominant framework for alignment of large language models (LLM), whether through reinforcement learning from human feedback or direct preference optimisation, is to learn from…

cs.LG2024

Understanding the performance gap between online and offline alignment algorithms

Yunhao Tang, Daniel Zhaohan Guo, Zeyu Zheng +8

Reinforcement learning from human feedback (RLHF) is the canonical framework for large language model alignment. However, rising popularity in offline alignment algorithms challeng…

cs.LG20242 cited

Human Alignment of Large Language Models through Online Preference Optimisation

Daniele Calandriello, Daniel Guo, Remi Munos +10

Ensuring alignment of language models' outputs with human preferences is critical to guarantee a useful, safe, and pleasant user experience. Thus, human alignment has been extensiv…

cs.LG2024

Off-policy Distributional Q(): Distributional RL without Importance Sampling

Yunhao Tang, Mark Rowland, Rémi Munos +2

We introduce off-policy distributional Q(), a new addition to the family of off-policy distributional evaluation algorithms. Off-policy distributional Q() does not apply impo…

cs.LG2024

Generalized Preference Optimization: A Unified Approach to Offline Alignment

Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng +7

Offline preference optimization allows fine-tuning large models directly from offline data, and has proved effective in recent alignment practices. We propose generalized preferenc…