activity
20242026
collaborators

19 papers

cs.CL2026

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discar…

cs.CL2026

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Alicia Parrish, Rajat Shinde, Sanket Badhe +57

Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…

cs.CL2026

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

Sourabrata Mukherjee, Hamna Hamna, Kalika Bali +1

LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by mea…

cs.CL2026

DEPART: DEcomposing PARiTy across Multilingual LLMs

Manan Uppadhyay, Prashant Kodali, Pranjal Chitale +3

Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering pr…

cs.CL2026

Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model

Divyanshu Aggarwal, Sankarshan Damle, Navin Goyal +2

A common challenge towards the adaptability of Large Language Models (LLMs) is their ability to learn new languages over time without hampering the model's performance on languages…

cs.CL2026

Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings

Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5

Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…