5 citations · 16 across the 5 of their papers we have counts for
1 paper · 1 filter
Andrea Schioppa, Xavier Garcia, Orhan Firat
The recent rapid progress in pre-training Large Language Models has relied on using self-supervised language modeling objectives like next token prediction or span corruption. On t…