Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
On-Policy Replay for Continual Supervised Fine-Tuning
Yan Chen, Taojie Zhu, Meng Zhang +4
Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers from catastrophic forgetting…
cs.LG2026
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning
Taojie Zhu, Dongyang Xu, Ding Zou +4
Post-training paradigms for Large Language Models (LLMs), primarily Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), face a fundamental dilemma: SFT provides stability…