24 citations · 61 across the 6 of their papers we have counts for
1 paper · 1 filter
Liqun Chen, Yizhe Zhang, Ruiyi Zhang +7
Sequence-to-sequence models are commonly trained via maximum likelihood estimation (MLE). However, standard MLE training considers a word-level objective, predicting the next word…