1 paper
Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron +1
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new token costs more than the previous one. For each additional tok…