activity
20242026
collaborators

40 papers

cs.CL2026

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

Naihao Deng, Yilun Zhu, Joan Nwatu +2

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this w…

cs.CL2026

The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference

Naihao Deng, Alissa Shen, Yiming Feng +5

Large language models (LLMs) are increasingly deployed in multilingual settings, yet the energy costs of serving these models across different languages remain poorly understood. W…

cs.CL2026

The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs

Naihao Deng, Yiming Feng, Chimaobi Okite +4

Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an…

cs.CL2026

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO

Naihao Deng, Yilun Zhu, Naichen Shi +2

Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and…

cs.CL2026

Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models

Angana Borah, Isabelle Augenstein, Rada Mihalcea

Large language models are increasingly used for social decision-making situations that require balancing cultural norms with personal preferences. For example, a user preferring ho…

cs.CL2026

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

Samee Arif, Angana Borah, Rada Mihalcea

Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidan…