most citedThe Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2024

Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment

Yan Liu, Xiaoyuan Yi, Xiaokang Chen +6

The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models…

cs.CL20241 cited

The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models

Yan Liu, Yu Liu, Xiaokang Chen +4

Pre-trained Language models (PLMs) have been acknowledged to contain harmful information, such as social biases, which may cause negative social impacts or even bring catastrophic…

cs.CL2024

Improving Long Text Understanding with Knowledge Distilled from Summarization Model

Yan Liu, Yazheng Yang, Xiaokang Chen

Long text understanding is important yet challenging for natural language processing. A long article or document usually contains many redundant words that are not pertinent to its…

cs.CL2023

Uncovering and Categorizing Social Biases in Text-to-SQL

Yan Liu, Yan Gao, Zhe Su +3

Content Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.}…

cs.CL20234 cited

Uncovering and Quantifying Social Biases in Code Generation

Yan Liu, Xiaokang Chen, Yan Gao +6

With the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the socia…