49 citations · 49 across the 3 of their papers we have counts for
1 paper · 1 filter
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar +1
We present the Compressive Transformer, an attentive sequence model which compresses past memories for long-range sequence learning. We find the Compressive Transformer obtains sta…