2 papers
cs.CL2026
From Rollouts to Recipes: Self-Contained Post-Training for LLMs
Yifei Li, Lingling Zhang, Muye Huang +3
Post-training large language models usually applies a single training recipe to all samples, even though the model's own rollouts reveal different sample-level learning states. We…
cs.AI2026
Self-Specialized Teachers for Domain Post-Training
Yifei Li, Rongman Xu, Lingling Zhang +5
Target-only post-training can improve performance in a specialized domain while degrading behaviors that a general-purpose base model acquired before adaptation. We study this prob…