2 papers
cs.AI2026
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
Minji Jung, Minjae Lee, Yejin Kim +2
LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rat…
cs.HC2026
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
Min Hun Lee
Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-A…