1 citations · 1 across the 11 of their papers we have counts for
5 papers · 1 filter
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity
Fengyuan Liu, Yongliang Miao, Zirui He +3
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framewo…
RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning
Yongliang Miao, Fengyuan Liu, Wei Shi +4
Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations, often treating a single expert…
HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
Shuang Liu, Yuxuan Bo, Qiuyang Zhao +4
Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking. To evaluate reward-model robustness, we introduce RewardHackBench c…
Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization
Wenjie Sun, Jinning Yang, Shuai Zhang +1
The guidance of scaling laws has increased the resource demands of modern large language models (LLMs), yet it remains questionable whether these models utilize resources effective…
Aligning Large Language Models and Geometric Deep Models for Protein Representation
Dong Shu, Bingbing Duan, Kai Guo +3
Latent representation alignment has become a foundational technique for constructing multimodal large language models (MLLM) by mapping embeddings from different modalities into a…