activity
20202022
most citedDiagonal State Spaces are as Effective as Structured State Spaces

64 citations · 90 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG202264 cited

Diagonal State Spaces are as Effective as Structured State Spaces

Ankit Gupta, Albert Gu, Jonathan Berant

Modeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video. While…

cs.CL2021

Memory-efficient Transformers via Top- Attention

Ankit Gupta, Guy Dar, Shaya Goodman +2

Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…

cs.LG2021

Value-aware Approximate Attention

Ankit Gupta, Jonathan Berant

Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…

cs.LG202021 cited

GMAT: Global Memory Augmentation for Transformers

Ankit Gupta, Jonathan Berant

Transformer-based models have become ubiquitous in natural language processing thanks to their large capacity, innate parallelism and high performance. The contextualizing componen…

cs.CL2020

Injecting Numerical Reasoning Skills into Language Models

Mor Geva, Ankit Gupta, Jonathan Berant

Large pre-trained language models (LMs) are known to encode substantial amounts of linguistic information. However, high-level reasoning skills, such as numerical reasoning, are di…

cs.CL20205 cited

Break It Down: A Question Understanding Benchmark

Tomer Wolfson, Mor Geva, Ankit Gupta +4

Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decom…