64 citations · 90 across the 3 of their papers we have counts for
6 papers
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, Jonathan Berant
Modeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video. While…
Memory-efficient Transformers via Top- Attention
Ankit Gupta, Guy Dar, Shaya Goodman +2
Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…
Value-aware Approximate Attention
Ankit Gupta, Jonathan Berant
Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…
GMAT: Global Memory Augmentation for Transformers
Ankit Gupta, Jonathan Berant
Transformer-based models have become ubiquitous in natural language processing thanks to their large capacity, innate parallelism and high performance. The contextualizing componen…
Injecting Numerical Reasoning Skills into Language Models
Mor Geva, Ankit Gupta, Jonathan Berant
Large pre-trained language models (LMs) are known to encode substantial amounts of linguistic information. However, high-level reasoning skills, such as numerical reasoning, are di…
Break It Down: A Question Understanding Benchmark
Tomer Wolfson, Mor Geva, Ankit Gupta +4
Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decom…