4 papers
Dishonesty in Helpful and Harmless Alignment
Youcheng Huang, Jingkun Tang, Duanyu Feng +4
People tell lies when seeking rewards. Large language models (LLMs) are aligned to human values with reinforcement learning where they get rewards if they satisfy human preference.…
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
Duanyu Feng, Bowen Qin, Chen Huang +3
The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture t…
Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers
Zihan Qiu, Zeyu Huang, Youcheng Huang +1
The feed-forward networks (FFNs) in transformers are recognized as a group of key-value neural memories to restore abstract high-level knowledge. In this work, we conduct an empiri…
See the Unseen: Better Context-Consistent Knowledge-Editing by Noises
Youcheng Huang, Wenqiang Lei, Zheng Zhang +2
Knowledge-editing updates knowledge of large language models (LLMs) and contributes to the interpretability and application of LLMs. However, knowledge applying is context-consiste…