1 paper
Jin Xu, Camille Couturier, Victor Rühle +2
Transformer blocks typically combine multi-head attention (MHA) for token mixing with gated MLPs for token-wise feature transformation, yet many choices in their parameterization r…