2 papers
cs.CL2026
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee +5
Large Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end user…
stat.ME2026
The Global Representativeness Index: A Total Variation Distance Framework for Measuring Demographic Fidelity in Survey Research
Evan Hadfield, Andrew Konya
Global survey research increasingly informs high-stakes decisions in AI governance and cross-cultural policy, yet no standardized metric quantifies how well a sample's demographic…