Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO
Naihao Deng, Yilun Zhu, Naichen Shi +2
Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and…
cs.CL2026
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Naihao Deng, Yilun Zhu, Joan Nwatu +2
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this w…