10 papers
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discar…
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
Sourabrata Mukherjee, Hamna Hamna, Kalika Bali +1
LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by mea…
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
Neha Joshi, Pamir Gogoi, Aasim Mirza +7
We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered…
What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets
Pamir Gogoi, Neha Joshi, Ayushi Pandey +6
This paper presents a multimodal dataset of 1,000 indigenous recipes from remote regions of India, collected through a participatory model involving first-time digital workers from…
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
Agrima Seth, Monojit Choudhary, Sunayana Sitaram +3
Representational bias in large language models (LLMs) has predominantly been measured through single-response interactions and has focused on Global North-centric identities like r…