90 citations · 422 across the 25 of their papers we have counts for
7 papers · 1 filter
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, Jonathan Berant
Modeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video. While…
Achieving Model Robustness through Discrete Adversarial Training
Maor Ivgi, Jonathan Berant
Discrete adversarial attacks are symbolic perturbations to a language input that preserve the output label but lead to a prediction error. While such attacks have been extensively…
Value-aware Approximate Attention
Ankit Gupta, Jonathan Berant
Following the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input le…
GMAT: Global Memory Augmentation for Transformers
Ankit Gupta, Jonathan Berant
Transformer-based models have become ubiquitous in natural language processing thanks to their large capacity, innate parallelism and high performance. The contextualizing componen…
White-to-Black: Efficient Distillation of Black-Box Adversarial Attacks
Yotam Gil, Yoav Chai, Or Gorodissky +1
Adversarial examples are important for understanding the behavior of neural models, and can improve their robustness through adversarial training. Recent work in natural language p…
Neural network gradient-based learning of black-box function interfaces
Alon Jacovi, Guy Hadash, Einat Kermany +4
Deep neural networks work well at approximating complicated functions when provided with data and trained by gradient descent methods. At the same time, there is a vast amount of e…