1 paper
Manzil Zaheer, Guru Guruganesh, Avinava Dubey +8
Transformers-based models, such as BERT, have been one of the most successful deep learning models for NLP. Unfortunately, one of their core limitations is the quadratic dependency…