1 paper · 1 filter
Liaoyaqi Wang, Chunsheng Zuo, William Jurayj +2
Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward of…