1 paper
Shuhao Li, Guodong Du, Anhao Zhao +3
Large language models have made strong reasoning gains through supervised fine-tuning, reinforcement learning, and on-policy distillation, yet these post-training methods are usual…