6 papers
What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility
Yuren Hao, Xiang Wan, ChengXiang Zhai
Despite the success of deep learning, training deep networks in biologically plausible and hardware-efficient ways remains an open challenge. Feedback alignment (FA) methods addres…
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
Zhuo Li, Pengyu Cheng, Zhechao Yu +7
Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values. However, RM training data is commonl…
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
Yuhao Du, Zhuo Li, Pengyu Cheng +4
Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning Large Language Models (LLMs) with human values. However, RLHF has been continuously challenged by its high…
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
Zhuo Li, Yuege Feng, Dandan Guo +3
The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective ha…
Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigm
Zhuo Li, Yuhao Du, Xiaoqi Jiao +5
Selecting high-quality and diverse training samples from extensive datasets plays a crucial role in reducing training overhead and enhancing the performance of Large Language Model…
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
Yuhao Du, Zhuo Li, Pengyu Cheng +2
Despite the substantial advancements in artificial intelligence, large language models (LLMs) remain being challenged by generation safety. With adversarial jailbreaking prompts, o…