1 paper
Hugo Pitorro, Pavlo Vasylenko, Marcos Treviso +1
Transformers are the current architecture of choice for NLP, but their attention layers do not scale well to long contexts. Recent works propose to replace attention with linear re…