1 paper
Zhichen Dong, Zhixuan Liu, Yuyu Fan +3
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effe…