Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
Kai Qin, Liangxin Liu, Yu Liang +7
Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment quality of Large Language Models (…
cs.AI2025
GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning
Jiaqi Wu, Qinlao Zhao, Zefeng Chen +4
Autonomous agents powered by large language models (LLMs) have shown impressive capabilities in tool manipulation for complex task-solving. However, existing paradigms such as ReAc…