1 paper · 1 filter
Tingkai Liu, Ari S. Benjamin, Anthony M. Zador
In the current work, we connect token-level uncertainty in causal language modeling to two types of training objectives: 1) masked maximum likelihood (MLE), 2) self-distillation. W…