10 papers
Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discar…
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
Sourabrata Mukherjee, Hamna Hamna, Kalika Bali +1
LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared signal or shared bias by mea…
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
Sourabrata Mukherjee, Atharva Mehta, Sougata Saha +2
The obligatory use of third-person honorifics is a distinctive feature of several South Asian languages, encoding nuanced socio-pragmatic cues such as power, age, gender, fame, and…
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
Sourabrata Mukherjee, Atul Kr. Ojha, John P. McCrae +1
Text style transfer (TST) is the task of transforming a text to reflect a particular style while preserving its original content. Evaluating TST outputs is a multidimensional chall…
Are Large Language Models Actually Good at Text Style Transfer?
Sourabrata Mukherjee, Atul Kr. Ojha, OndÅej DuÅ¡ek
We analyze the performance of large language models (LLMs) on Text Style Transfer (TST), specifically focusing on sentiment transfer and text detoxification across three languages:…