1 paper
M. Sagitova, O. Duranthon, L. Zdeborová
Multi-head attention enables transformer models to represent multiple attention patterns simultaneously. Empirically, head specialization emerges in distinct stages during training…