Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training
Jianzhe Lin
Self-improvement can self-regress. In REINFORCE post-training for code, a model can quickly improve on its optimized metric and then collapse within the same training campaign. We…
cs.AI2026
Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines
Jianzhe Lin, Fei Wang, Xiaolin Li +2
Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime…
cs.AI2025
Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
Jianzhe Lin, Zeyu Pan, Yun Zhu +2
We introduce SuperIntelliAgent, an agentic learning framework that couples a trainable small diffusion model (the learner) with a frozen large language model (the verifier) to enab…