3 papers
cs.CL2024
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
Yan Liu, Xiaoyuan Yi, Xiaokang Chen +6
The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models…
cs.CL2024
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
Yan Liu, Yu Liu, Xiaokang Chen +4
Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic…
cs.CL2024
Improving Long Text Understanding with Knowledge Distilled from Summarization Model
Yan Liu, Yazheng Yang, Xiaokang Chen
Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its…