3 papers
cs.CL2026
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
Sanjeeevan Selvaganapathy, Mehwish Nasim
We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare…
cs.CL2026
Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
Mehwish Nasim, Sanjeevan Selvaganapathy, Neel Ganapathi Sabhahit +6
Many benchmarks show that large language models can answer direct questions about culture. We study a different question: do they also change how they speak when culture is only im…
cs.CL2026
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
Pranav Bhandari, Nicolas Fay, Sanjeevan Selvaganapathy +3
Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The ne…