activity
20242026
collaborators

10 papers

cs.CL2026

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

Rima Hazra, Bikram Ghuku, Ilona Marchenko +5

Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture…

cs.CL2025

ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models

Somnath Banerjee, Sayan Layek, Sayantan Adak +3

Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can a…

cs.CL2025

Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations

Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar +4

While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluati…

cs.IR2025

MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation

Sayantan Adak, Somnath Banerjee, Rajarshi Mandal +4

Online memes are a powerful yet challenging medium for content moderation, often masking harmful intent behind humor, irony, or cultural symbolism. Conventional moderation systems…

cs.CL2025

Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment

Somnath Banerjee, Sayan Layek, Pratyush Chatterjee +2

Ensuring consistent safety across multiple languages remains a significant challenge for large language models (LLMs). We introduce Soteria, a lightweight yet powerful strategy tha…

cs.CL2025

Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance

Somnath Banerjee, Avik Halder, Rajarshi Mandal +4

The integration of pretrained language models (PLMs) like BERT and GPT has revolutionized NLP, particularly for English, but it has also created linguistic imbalances. This paper s…