9 citations · 9 across the 1 of their papers we have counts for
1 paper
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière +2
Large language models such as GPT and Llama are trained with a next-token prediction loss. In this work, we suggest that training language models to predict multiple future tokens…