1 paper · 1 filter
Carles Gelada, Jacob Buckman, Sean Zhang +1
We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the…