64 citations · 90 across the 3 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2022★ 64 cited
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, Jonathan Berant
Modeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video. While…
cs.LG2021
Value-aware Approximate Attention
Ankit Gupta, Jonathan Berant
Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…
cs.LG2020★ 21 cited
GMAT: Global Memory Augmentation for Transformers
Ankit Gupta, Jonathan Berant
Transformer-based models have become ubiquitous in natural language processing thanks to their large capacity, innate parallelism and high performance. The contextualizing componen…