activity
20162020
most citedStabilizing Transformers for Reinforcement Learning

132 citations · 279 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG20201 cited

Do Transformers Need Deep Long-Range Memory

Jack W. Rae, Ali Razavi

Deep attention models have advanced the modelling of sequential data across many domains. For language modelling in particular, the Transformer-XL -- a Transformer augmented with a…

cs.LG201949 cited

Compressive Transformers for Long-Range Sequence Modelling

Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar +1

We present the Compressive Transformer, an attentive sequence model which compresses past memories for long-range sequence learning. We find the Compressive Transformer obtains sta…

cs.LG2019132 cited

Stabilizing Transformers for Reinforcement Learning

Emilio Parisotto, H. Francis Song, Jack W. Rae +10

Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown brea…

cs.AI201939 cited

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg +11

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy…

cs.LG201658 cited

Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes

Jack W Rae, Jonathan J Hunt, Tim Harley +5

Neural networks augmented with external memory have the ability to learn algorithmic solutions to complex tasks. These models appear promising for applications such as language mod…