3 papers
cs.CL2026
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
Jen-tse Huang, Jiantong Qin, Xueli Qiu +3
Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in…
cs.CL2025
Characterizing Selective Refusal Bias in Large Language Models
Adel Khorramrouz, Sharon Levy
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…
cs.CL2024
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats
Kuleen Sasse, Carlos Aguirre, Isabel Cachola +2
WARNING: This paper contains content that maybe upsetting or offensive to some readers. Dog whistles are coded expressions with dual meanings: one intended for the general public (…