Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
Tingxu Han, Wei Song, Ziqi Ding +6
Large language models (LLMs) increasingly mediate decisions in domains where unfair treatment of demographic groups is unacceptable. Existing work probes when biased outputs appear…
cs.CL2025
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
Junchen Ding, Penghao Jiang, Zihao Xu +4
As large language models (LLMs) increasingly mediate ethically sensitive decisions, understanding their moral reasoning processes becomes imperative. This study presents a comprehe…