5 citations · 5 across the 1 of their papers we have counts for
1 paper · 1 filter
Nathan Leroux, Paul-Philipp Manea, Chirag Sudarshan +4
Transformer networks, driven by self-attention, are central to Large Language Models. In generative Transformers, self-attention uses cache memory to store token projections, avoid…