1 paper · 1 filter
Mingze Dong, Leda Wang, Yuval Kluger
Mask-based pretraining has become a cornerstone of modern large-scale models across language, vision, and recently biology. Despite its empirical success, its role and limits in le…