6 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.LG2020★ 1 cited
Do Transformers Need Deep Long-Range Memory
Jack W. Rae, Ali Razavi
Deep attention models have advanced the modelling of sequential data across many domains. For language modelling in particular, the Transformer-XL -- a Transformer augmented with a…
cs.LG2019★ 6 cited
Meta-Learning Neural Bloom Filters
Jack W Rae, Sergey Bartunov, Timothy P Lillicrap
There has been a recent trend in training neural networks to replace data structures that have been crafted by hand, with an aim for faster execution, better accuracy, or greater c…