4 papers
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
Somnath Banerjee, Sayan Layek, Sayantan Adak +3
Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can a…
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar +4
While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluati…
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
Sayantan Adak, Somnath Banerjee, Rajarshi Mandal +4
Online memes are a powerful yet challenging medium for content moderation, often masking harmful intent behind humor, irony, or cultural symbolism. Conventional moderation systems…
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
Somnath Banerjee, Sayan Layek, Pratyush Chatterjee +2
Ensuring consistent safety across multiple languages remains a significant challenge for large language models (LLMs). We introduce Soteria, a lightweight yet powerful strategy tha…