collaborators

8 papers

cs.LG2026

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

Fengyuan Liu, Yongliang Miao, Zirui He +3

Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framewo…

cs.CL2026

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

Jingyi He, Haiyan Zhao, Ruxue Shi +4

Although Sparse Autoencoders (SAEs) have mitigated the opacity of large language models (LLMs) by decomposing dense representations into sparse features, explaining these features…

cs.LG2026

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

Yongliang Miao, Fengyuan Liu, Wei Shi +4

Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert…

cs.LG2026

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models

Shuang Liu, Yuxuan Bo, Qiuyang Zhao +4

Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking. To evaluate reward-model robustness, we introduce RewardHackBench c…

cs.CL2026

FinAnchor: Aligned Multi-Model Representations for Financial Prediction

Zirui He, Huopu Zhang, Yanguang Liu +2

Financial prediction from long documents involves significant challenges, as actionable signals are often sparse and obscured by noise, and the optimal LLM for generating embedding…

cs.CL2026

NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models

Weiqi Liu, Yongliang Miao, Haiyan Zhao +2

Neuron-level interpretation in large language models (LLMs) is fundamentally challenged by widespread polysemanticity, where individual neurons respond to multiple distinct semanti…