1 paper
Heejune Sheen, Siyu Chen, Tianhao Wang +1
We study gradient flow on the exponential loss for a classification problem with a one-layer softmax attention model, where the key and query weight matrices are trained separately…