activity
20162022
most citedMONet: Unsupervised Scene Decomposition and Representation

193 citations · 746 across the 16 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG2022112 cited

Fine-tuning language models to find agreement among humans with diverse preferences

Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan +8

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static…

cs.LG202112 cited

Synthetic Returns for Long-Term Credit Assignment

David Raposo, Sam Ritter, Adam Santoro +5

Since the earliest days of reinforcement learning, the workhorse method for assigning credit to actions over time has been temporal-difference (TD) learning, which propagates credi…

cs.LG2021

Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents

Jane X. Wang, Michael King, Nicolas Porcel +14

There has been rapidly growing interest in meta-learning as a method for increasing the flexibility and sample efficiency of reinforcement learning. One problem in this area of res…

cs.LG2020

Rapid Task-Solving in Novel Environments

Sam Ritter, Ryan Faulkner, Laurent Sartran +3

We propose the challenge of rapid task-solving in novel environments (RTS), wherein an agent must solve a series of tasks as rapidly as possible in an unfamiliar environment. An ef…

cs.LG20207 cited

MEMO: A Deep Network for Flexible Combination of Episodic Memories

Andrea Banino, Adrià Puigdomènech Badia, Raphael Köster +7

Recent research developing neural network architectures with external memory have often used the benchmark bAbI question and answering dataset which provides a challenging number o…

cs.LG2019132 cited

Stabilizing Transformers for Reinforcement Learning

Emilio Parisotto, H. Francis Song, Jack W. Rae +10

Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown brea…