82 citations · 93 across the 2 of their papers we have counts for
1 paper · 1 filter
Manzil Zaheer, Guru Guruganesh, Avinava Dubey +8
Transformers-based models, such as BERT, have been one of the most successful deep learning models for NLP. Unfortunately, one of their core limitations is the quadratic dependency…