1 paper
Duke Nguyen, Du Yin, Aditya Joshi +1
Linearization of attention using various kernel approximation and kernel learning techniques has shown promise. Past methods used a subset of combinations of component functions an…