1 citations · 1 across the 1 of their papers we have counts for
1 paper
Da Xiao, Qingye Meng, Shengping Li +1
Multi-Head Attention (MHA) is a key component of Transformer. In MHA, attention heads work independently, causing problems such as low-rank bottleneck of attention score matrices a…