collaborators

10 papers

cs.AI2026

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

Tong Che, Rui Wu

Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted}…

cs.LG2026

Constitutional Value Potentials: reading and steering internal priority margins in language models

Tong Che, Rui Wu

A constitution tells a language model what to value, but little tells us whether it does. Adherence is judged from outputs, and output evidence is most fragile on value conflicts,…

cs.AI2026

Reference Feature Atlases for Mechanistic Auditing of Language Models

Rui Wu, Tong Che

Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained…

cs.AI2026

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

Can Jin, Jiakang Li, Rui Wu +3

As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong ge…

cs.CL2026

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

Bobo Li, Rui Wu, Zibo Ji +5

Large Language Model agents have rapidly evolved from static text generators into dynamic systems capable of executing complex autonomous workflows. To enhance reliability, multi-a…

cs.LG2026

From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

Rui Wu, Ruixiang Tang

Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intended task. We systematically study this p…