activity
20172022
most citedBootstrap your own latent: A new approach to self-supervised Learning

3.4k citations · 3.5k across the 11 of their papers we have counts for

collaborators

12 papers

cs.LG20221 cited

Understanding Self-Predictive Learning for Reinforcement Learning

Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond +13

We study the learning dynamics of self-predictive learning for reinforcement learning, a family of algorithms that learn representations by minimizing the prediction error of their…

cs.LG20221 cited

Categorical SDEs with Simplex Diffusion

Pierre H. Richemond, Sander Dieleman, Arnaud Doucet

Diffusion models typically operate in the standard framework of generative modelling by producing continuously-valued datapoints. To this end, they rely on a progressive Gaussian s…

stat.ML2020

BYOL works even without batch statistics

Pierre H. Richemond, Jean-Bastien Grill, Florent Altché +8

Bootstrap Your Own Latent (BYOL) is a self-supervised learning approach for image representation. From an augmented view of an image, BYOL trains an online network to predict a tar…

cs.LG20203.4k cited

Bootstrap your own latent: A new approach to self-supervised Learning

Jean-Bastien Grill, Florian Strub, Florent Altché +11

We introduce Bootstrap Your Own Latent (BYOL), a new approach to self-supervised image representation learning. BYOL relies on two neural networks, referred to as online and target…

cs.LG20192 cited

Biologically inspired architectures for sample-efficient deep reinforcement learning

Pierre H. Richemond, Arinbjörn Kolbeinsson, Yike Guo

Deep reinforcement learning requires a heavy price in terms of sample efficiency and overparameterization in the neural networks used for function approximation. In this work, we u…

cs.LG20192 cited

Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control

Marta Sarrico, Kai Arulkumaran, Andrea Agostinelli +2

Deep networks have enabled reinforcement learning to scale to more complex and challenging domains, but these methods typically require large quantities of training data. An altern…