2 papers
cs.LG2024
SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection
Ke Ye, Heinrich Jiang, Afshin Rostamizadeh +6
Pre-training large language models is known to be extremely resource intensive and often times inefficient, under-utilizing the information encapsulated in the training text sequen…
cs.LG2023
Leveraging Importance Weights in Subset Selection
Gui Citovsky, Giulia DeSalvo, Sanjiv Kumar +3
We present a subset selection algorithm designed to work with arbitrary model families in a practical batch setting. In such a setting, an algorithm can sample examples one at a ti…