1 paper
Shanda Li, Chong You, Guru Guruganesh +7
Preventing the performance decay of Transformers on inputs longer than those used for training has been an important challenge in extending the context length of these models. Thou…