1 paper
Timo Lohrenz, Björn Möller, Zhengyang Li +1
The powerful modeling capabilities of all-attention-based transformer architectures often cause overfitting and - for natural language processing tasks - lead to an implicitly lear…