2 papers
cs.LG2026
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
Mahdi Sabbaghi, George Pappas, Adel Javanmard +1
Supervised fine-tuning (SFT) provides the standard approach for teaching LLMs new behaviors from offline expert demonstrations. However, standard SFT uniformly fits all samples --…
cs.LG2026
Robust Policy Optimization to Prevent Catastrophic Forgetting
Mahdi Sabbaghi, George Pappas, Adel Javanmard +1
Large language models are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet even small downstream updates can…