collaborators

13 papers

cs.LG2026

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

Fengyuan Liu, Yongliang Miao, Zirui He +3

Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framewo…

cs.CL2026

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

Jingyi He, Haiyan Zhao, Ruxue Shi +4

Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features…

cs.LG2026

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

Yongliang Miao, Fengyuan Liu, Wei Shi +4

Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert…

cs.LG2026

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models

Shuang Liu, Yuxuan Bo, Qiuyang Zhao +4

Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking. To evaluate reward-model robustness, we introduce RewardHackBench c…

cs.LG2026

Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization

Wenjie Sun, Jinning Yang, Shuai Zhang +1

The guidance of scaling laws has increased the resource demands of modern large language models (LLMs), yet it remains questionable whether these models utilize resources effective…

cs.CL2026

Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

Haiyan Zhao, Zirui He, Guanchu Wang +3

Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own ac…