1 paper
Tommaso Cerruti, Tim Rieder, George Rowlands +2
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper prese…