1 paper
Carles Gelada, Jacob Buckman, Sean Zhang +1
We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the…