1 paper
Yijia Luo, Yulin Song, Xingyao Zhang +5
Recent advancements in large language models (LLMs) have demonstrated remarkable reasoning capabilities through long chain-of-thought (CoT) reasoning. The R1 distillation scheme ha…