1 paper
Xucong Hu, Jian-Qiao Zhu
Autoregressive language models are next-token predictors and have been criticized for only optimizing surface plausibility (i.e., local coherence) rather than maintaining correct l…