collaborators

13 papers

cs.LG2026

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Xuekang Wang, Zhuoyuan Hao, Shuo Hou +3

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in t…

cs.CL2026

WildReward: Learning Reward Models from In-the-Wild Human Interactions

Hao Peng, Yunjia Qi, Xiaozhi Wang +3

Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deplo…

cs.CL2026

On the Paradoxical Interference between Instruction-Following and Task Solving

Yunjia Qi, Hao Peng, Xintong Shi +5

Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counter…

cs.CL2025

Auxiliary Metrics Help Decoding Skill Neurons in the Wild

Yixiu Zhao, Xiaozhi Wang, Zijun Yao +2

Large language models (LLMs) exhibit remarkable capabilities across a wide range of tasks, yet their internal mechanisms remain largely opaque. In this paper, we introduce a simple…

cs.CL2025

Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons

Jianhui Chen, Xiaozhi Wang, Zijun Yao +3

Large language models (LLMs) excel in various capabilities but pose safety risks such as generating harmful content and misinformation, even after safety alignment. In this paper,…

cs.CL2025

LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder

Yi Jing, Zijun Yao, Hongzhu Guo +4

Large language models (LLMs) demonstrate exceptional performance on tasks requiring complex linguistic abilities, such as reference disambiguation and metaphor recognition/generati…