2 papers
cs.LG2026
STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majorit…
cs.AI2026
The Validity Gap in Health AI Evaluation: A Cross-Sectional Analysis of Benchmark Composition
Alvin Rajkomar, Pavan Sudarshan, Angela Lai +1
Background: Clinical trials rely on transparent inclusion criteria to ensure generalizability. In contrast, benchmarks validating health-related large language models (LLMs) rarely…