1 paper · 1 filter
Tomohiro Hayase, Benoît Collins, Ryo Karakida
Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective…