1 paper
Jia Cheng Hu, Roberto Cavicchioli, Giulia Berardinelli +1
Although the Transformer is currently the best-performing architecture in the homogeneous configuration (self-attention only) in Neural Machine Translation, many State-of-the-Art m…