4 papers · 1 filter
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
Jiayun Wu, Peixu Hou, Shan Qu +3
Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior ac…
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
Yuzhuo Bai, Shitong Duan, Muhua Huang +7
Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values…
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Shitong Duan, Xiaoyuan Yi, Peng Zhang +5
Large language models (LLMs) have revolutionized the role of AI, yet pose potential social risks. To steer LLMs towards human preference, alignment technologies have been introduce…
Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks
Yuxuan Lu, Bingsheng Yao, Shao Zhang +5
Large Language Models (LLMs) have demonstrated considerable advances, and several claims have been made about their exceeding human performance. However, in real-world tasks, domai…