1 paper
Timur Mudarisov, Mikhal Burtsev, Tatiana Petrova +1
We present a geometric framework for analysing multi-head attention in large language models (LLMs). Without altering the mechanism, we view standard attention through a top-N sele…