1 citations · 1 across the 2 of their papers we have counts for
5 papers
Correcting Large Language Model Behavior via Influence Function
Han Zhang, Zhuo Zhang, Yi Zhang +8
Recent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature…
Online Self-Preferring Language Models
Yuanzhao Zhai, Zhuo Zhang, Kele Xu +6
Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward…
Multi-modal Stance Detection: New Datasets and Model
Bin Liang, Ang Li, Jingqian Zhao +5
Stance detection is a challenging task that aims to identify public opinion from social media platforms with respect to specific targets. Previous work on stance detection largely…
COPR: Continual Human Preference Learning via Optimal Policy Regularization
Han Zhang, Lin Gui, Yu Lei +8
Reinforcement Learning from Human Feedback (RLHF) is commonly utilized to improve the alignment of Large Language Models (LLMs) with human preferences. Given the evolving nature of…
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
Yuanzhao Zhai, Han Zhang, Yu Lei +5
Reinforcement learning from human feedback (RLHF) emerges as a promising paradigm for aligning large language models (LLMs). However, a notable challenge in RLHF is overoptimizatio…