84 citations · 232 across the 12 of their papers we have counts for
1 paper · 1 filter
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière +2
Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens…