17 citations · 26 across the 14 of their papers we have counts for
1 paper · 1 filter
Jihoon Tack, Jack Lanchantin, Jane Yu +7
Next token prediction has been the standard training objective used in large language model pretraining. Representations are learned as a result of optimizing for token-level perpl…