Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG
Seungju Han, Konwoo Kim, Chanwoo Park +5
Synthetic data augmentation helps language models learn new knowledge in data-constrained domains. However, naively scaling existing synthetic data methods by training on more synt…
cs.LG2026
Data-efficient pre-training by scaling synthetic megadocs
Konwoo Kim, Suhas Kotha, Yejin Choi +3
Synthetic data augmentation has emerged as a promising solution when pre-training is constrained by data rather than compute. We study how to design synthetic data algorithms that…
cs.LG2025
Pre-training under infinite compute
Konwoo Kim, Suhas Kotha, Percy Liang +1
Since compute grows much faster than web text available for language model pre-training, we ask how one should approach pre-training under fixed data and no compute constraints. We…