Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
Yu Liang, Liangxin Liu, Longzheng Wang +5
Generative reward models (GRMs) have emerged as a promising approach for aligning Large Language Models (LLMs) with human preferences by offering greater representational capacity…
cs.AI2026
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
Kai Qin, Liangxin Liu, Yu Liang +7
Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment quality of Large Language Models (…
cs.AI2025
Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games
Yunhao Liang, Yuan Qu, Jingyuan Yang +2
Coordinating multiple large language models (LLMs) to solve complex tasks collaboratively poses a fundamental trade-off between the computation costs and collective performance com…