19 papers
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discar…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
Sourabrata Mukherjee, Hamna Hamna, Kalika Bali +1
LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by mea…
DEPART: DEcomposing PARiTy across Multilingual LLMs
Manan Uppadhyay, Prashant Kodali, Pranjal Chitale +3
Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering pr…
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
Divyanshu Aggarwal, Sankarshan Damle, Navin Goyal +2
A common challenge towards the adaptability of Large Language Models (LLMs) is their ability to learn new languages over time without hampering the model's performance on languages…
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…