2 papers
cs.CL2026
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
Sourabrata Mukherjee, Hamna Hamna, Kalika Bali +1
LLM judges now score most open-ended NLP output, and their mutual agreement is routinely read as evidence that the scores can be trusted. That reading is unsafe: judges may agree b…
cs.CL2025
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…