10 papers
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
Tong Che, Rui Wu
Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted}…
Constitutional Value Potentials: reading and steering internal priority margins in language models
Tong Che, Rui Wu
A constitution tells a language model what to value, but little tells us whether it does. Adherence is judged from outputs, and output evidence is most fragile on value conflicts,…
Reference Feature Atlases for Mechanistic Auditing of Language Models
Rui Wu, Tong Che
Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained…
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
Can Jin, Jiakang Li, Rui Wu +3
As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong ge…
Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
Bobo Li, Rui Wu, Zibo Ji +5
Large Language Model agents have rapidly evolved from static text generators into dynamic systems capable of executing complex autonomous workflows. To enhance reliability, multi-a…
From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering
Rui Wu, Ruixiang Tang
Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task. We systematically study this p…