1 paper
Kasun Dewage, Marianna Pensky, Suranadi De Silva +1
We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outli…