1 citations · 1 across the 3 of their papers we have counts for
4 papers
Dishonesty in Helpful and Harmless Alignment
Youcheng Huang, Jingkun Tang, Duanyu Feng +4
People tell lies when seeking rewards. Large language models (LLMs) are aligned to human values with reinforcement learning where they get rewards if they satisfy human preference.…
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
Duanyu Feng, Bowen Qin, Chen Huang +3
The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture t…
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
Duanyu Feng, Bowen Qin, Chen Huang +2
Direct Preference Optimization (DPO), which derives reward signals directly from pairwise preference data, has shown its effectiveness on aligning Large Language Models (LLMs) with…
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
Youcheng Huang, Wenqiang Lei, Zheng Zhang +2
Knowledge-editing updates knowledge of large language models (LLMs) and contributes to the interpretability and application of LLMs. However, knowledge applying is context-consiste…