5 papers · 1 filter
Enhancing Safety of Large Language Models via Embedding Space Separation
Xu Zhao, Xiting Wang, Weiran Shen
Large language models (LLMs) have achieved impressive capabilities, yet ensuring their safety against harmful prompts remains a critical challenge. Recent work has revealed that th…
Skywork-R1V3 Technical Report
Wei Shen, Jiangbo Pei, Yi Peng +8
We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning. Its key innovation lies in effectively transferrin…
A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
Jialun Zhong, Wei Shen, Yanzeng Li +7
Reward Model (RM) has demonstrated impressive potential for enhancing Large Language Models (LLM), as RM can serve as a proxy for human preferences, providing signals to guide LLMs…
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Enyu Zhou, Guodong Zheng, Binghai Wang +11
Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. Howe…
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
Ziyi Ye, Xiangsheng Li, Qiuchi Li +5
Learning from preference feedback is a common practice for aligning large language models~(LLMs) with human value. Conventionally, preference data is learned and encoded into a sca…