Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
Tong Che, Rui Wu
Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted}…
cs.AI2026
Reference Feature Atlases for Mechanistic Auditing of Language Models
Rui Wu, Tong Che
Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained…
cs.AI2024
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
Di Zhang, Jianbo Wu, Jingdi Lei +9
This paper presents an advanced mathematical problem-solving framework, LLaMA-Berry, for enhancing the mathematical reasoning ability of Large Language Models (LLMs). The framework…