1 citations · 1 across the 3 of their papers we have counts for
5 papers · 1 filter
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
Yan Liu, Xiaoyuan Yi, Xiaokang Chen +6
The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models…
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models
Yan Liu, Yu Liu, Xiaokang Chen +4
Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic…
Improving Long Text Understanding with Knowledge Distilled from Summarization Model
Yan Liu, Yazheng Yang, Xiaokang Chen
Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its…
Uncovering and Categorizing Social Biases in Text-to-SQL
Yan Liu, Yan Gao, Zhe Su +3
Content Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.}…
Uncovering and Quantifying Social Biases in Code Generation
Yan Liu, Xiaokang Chen, Yan Gao +6
With the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the socia…