1 paper · 1 filter
Congmin Zheng, Jiachen Zhu, Zhuoying Ou +8
Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final ans…