1 paper
Siyi Gu, Jialin Chen, Sophia Zhou +2
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-t…