1 paper · 1 filter
Yimin Tang, Yurong Xu, Ning Yan +1
Transformers have a quadratic scaling of computational complexity with input size, which limits the input context window size of large language models (LLMs) in both training and i…