1 paper
Runxuan Liu, Xianhao Ou, Xinyan Ma +13
Long Chain-of-Thought (LCoT), achieved by Reinforcement Learning with Verifiable Rewards (RLVR), has proven effective in enhancing the reasoning capabilities of Large Language Mode…