40 papers
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Naihao Deng, Yilun Zhu, Joan Nwatu +2
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this w…
The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference
Naihao Deng, Alissa Shen, Yiming Feng +5
Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. W…
The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs
Naihao Deng, Yiming Feng, Chimaobi Okite +4
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an…
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO
Naihao Deng, Yilun Zhu, Naichen Shi +2
Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and…
Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models
Angana Borah, Isabelle Augenstein, Rada Mihalcea
Large language models are increasingly used for social decision-making situations that require balancing cultural norms with personal preferences. For example, a user preferring ho…
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
Samee Arif, Angana Borah, Rada Mihalcea
Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidan…