1 paper · 1 filter
Dian Li, Zekun Wang, Yaoru Wang +1
Modern LLM training breaks a core assumption behind offline batch samplers: the true training cost of a sample is only observable after preprocessing, augmentation, templating, tok…