1 paper · 1 filter
Jiaao Yu, Shenwei Li, Mingjie Han +4
Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet…