10 papers
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
Rima Hazra, Bikram Ghuku, Ilona Marchenko +5
Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture…
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
Somnath Banerjee, Sayan Layek, Sayantan Adak +3
Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can a…
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar +4
While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluati…
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
Sayantan Adak, Somnath Banerjee, Rajarshi Mandal +4
Online memes are a powerful yet challenging medium for content moderation, often masking harmful intent behind humor, irony, or cultural symbolism. Conventional moderation systems…
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
Somnath Banerjee, Sayan Layek, Pratyush Chatterjee +2
Ensuring consistent safety across multiple languages remains a significant challenge for large language models (LLMs). We introduce Soteria, a lightweight yet powerful strategy tha…
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance
Somnath Banerjee, Avik Halder, Rajarshi Mandal +4
The integration of pretrained language models (PLMs) like BERT and GPT has revolutionized NLP, particularly for English, but it has also created linguistic imbalances. This paper s…