Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
Avni Mittal, Shanu Kumar, Sandipan Dandapat +1
We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are missing. This problem is comm…
cs.CL2025
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar +4
While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluati…
cs.CL2024
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
Somnath Banerjee, Sayan Layek, Hari Shrawgi +7
As LLMs are increasingly deployed in global applications, the importance of cultural sensitivity becomes paramount, ensuring that users from diverse backgrounds feel respected and…