1 paper · 1 filter
Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet
We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth…