1 paper · 1 filter
Ashutosh Sathe, Sunita Sarawagi
Maximizing the likelihood of the next token is an established, statistically sound objective for pre-training language models. In this paper we show that we can train better models…