1 paper · 1 filter
Xingjian Diao, Wenbo Li, Yashas Malur Saidutta +3
Long input sequences are central to document understanding and multi-step reasoning in Large Language Models, yet the quadratic cost of attention makes inference both memory-intens…