1 paper
Aryan Sood, Shantanu Acharya, Gaurav Kumar Nayak
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distributed blockwise methods such as…