Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective
Haichuan Wang, Tao Lin, Lingkai Kong +3
Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy.…
cs.AI2026
Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
Dingyi Yang, Junqi Zhao, Xue Li +2
Cognitive anthropology suggests that the distinction of human intelligence lies in the ability to infer other individuals' knowledge states and understand their intentions. In comp…