2 papers
cs.LG2025
Scaling Context Requires Rethinking Attention
Carles Gelada, Jacob Buckman, Sean Zhang +1
We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the…
cs.LG2025
Conformal Transformations for Symmetric Power Transformers
Saurabh Kumar, Jacob Buckman, Carles Gelada +1
Transformers with linear attention offer significant computational advantages over softmax-based transformers but often suffer from degraded performance. The symmetric power (sympo…