13 citations · 35 across the 18 of their papers we have counts for
8 papers · 1 filter
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
Jean Feng, Vishal Patel, Patrick Heagerty +5
Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Adaptive auditing of AI systems with anytime-valid guarantees
Siyu Zhou, Patrick Vossler, Venkatesh Sivaraman +2
A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gai…
Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading
Nicholas Sadjoli, Tim Siefken, Atin Ghosh +2
Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice…
Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
Penny Chong, Harshavardhan Abichandani, Jiyuan Shen +4
Agent applications are increasingly adopted to automate workflows across diverse tasks. However, due to the heterogeneous domains they operate in, it is challenging to create a sca…
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…