38 citations · 38 across the 1 of their papers we have counts for
1 paper
Yao Zhao, Misha Khalman, Rishabh Joshi +3
Conditional language models are predominantly trained with maximum likelihood estimation (MLE), giving probability mass to sparsely observed target sequences. While MLE trained mod…