1 paper · 1 filter
Anhao Zhao, Ziyang Chen, Junlong Tong +6
Large reasoning models (LRMs) are commonly trained with reinforcement learning (RL) to explore long chain-of-thought reasoning, achieving strong performance at high computational c…