1 paper · 1 filter
Emanuele Troiani, Hugo Cui, Yatin Dandi +2
In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first…