collaborators

13 papers

cs.CL2026

Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations

Somnath Banerjee, Pranav Jha, Rima Hazra +1

LLMs deployed multilingually are often audited via English explanations for non-English inputs. We evaluate extractive explanations ''where the model identifies input token spans a…

cs.CL2026

SafeMath: Inference-time Safety improves Math Accuracy

Sagnik Basu, Subhrajit Mitra, Aman Juneja +3

Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we s…

cs.CL2026

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

Rima Hazra, Bikram Ghuku, Ilona Marchenko +5

Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture…

cs.CL2026

Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages

Somnath Banerjee, Rima Hazra, Animesh Mukherjee

Large language models (LLMs) are being deployed across the Global South, where everyday use involves low-resource languages, code-mixing, and culturally specific norms. Yet safety…

cs.AI2026

From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents

Razeen A Rasheed, Somnath Banerjee, Animesh Mukherjee +1

A deep research agent produces a fluent scientific report in minutes; a careful reader then tries to verify the main claims and discovers the real cost is not reading, but tracing:…

cs.CL2025

ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models

Somnath Banerjee, Sayan Layek, Sayantan Adak +3

Current language model safety paradigms often fall short in emotionally charged or high-stakes settings, where refusal-only approaches may alienate users and naive compliance can a…