1 paper
Guoyin Wang, Chunyuan Li, Jianqiao Li +10
Neural language models are often trained with maximum likelihood estimation (MLE), where the next word is generated conditioned on the ground-truth word tokens. During testing, how…