Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
Yuhan Chen, Zhihua Tian, Mahavir Dabas +7
The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scor…
cs.AI2026
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
Jiacheng Liang, Yao Ma, Tharindu Kumarage +5
Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) ca…