1 paper
Omar Coser, Loredana Zollo, Paolo Soda +1
Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objec…