Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Data-efficient pre-training by scaling synthetic megadocs
Konwoo Kim, Suhas Kotha, Yejin Choi +3
Synthetic data augmentation has emerged as a promising solution when pre-training is constrained by data rather than compute. We study how to design synthetic data algorithms that…
cs.LG2025
Pre-training under infinite compute
Konwoo Kim, Suhas Kotha, Percy Liang +1
Since compute grows much faster than web text available for language model pre-training, we ask how one should approach pre-training under fixed data and no compute constraints. We…
cs.LG2025
Eliciting Language Model Behaviors with Investigator Agents
Xiang Lisa Li, Neil Chowdhury, Daniel D. Johnson +4
Language models exhibit complex, diverse behaviors when prompted with free-form text, making it difficult to characterize the space of possible outputs. We study the problem of beh…