1 citations · 1 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang, Sunbeom Kwon, Jinwen Luo +2
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative example…
cs.AI2026★ 1 cited
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu +6
This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…