1 paper
Zhenghao Lin, Zhibin Gou, Yeyun Gong +8
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that "9l training". Our ini…