4 citations · 8 across the 31 of their papers we have counts for
22 papers · 1 filter
Where Steering Signals Come From: Activation Source Selection in Activation Steering
Jiaran Ye, Lingxu Ran, Zijun Yao +5
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as…
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Weihan Meng, Hongzhu Guo, Yi Jing +5
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on extern…
SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang +4
The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation…
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
Kejian Zhu, Zhuoran Jin, Shangqing Tu +5
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preli…
WildReward: Learning Reward Models from In-the-Wild Human Interactions
Hao Peng, Yunjia Qi, Xiaozhi Wang +3
Reward models (RMs) are crucial for the training of large language models (LLMs), yet they typically rely on large-scale human-annotated preference pairs. With the widespread deplo…
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
Zhidian Huang, Zijun Yao, Ji Qi +7
Recent advances in multimodal large language models (MLLMs) mark a shift from non-thinking models to post-trained reasoning models capable of solving complex problems through think…