1 paper
Yangsong Lan, Renkai Hu, HongKai Zheng +4
Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such traj…